-
Notifications
You must be signed in to change notification settings - Fork 783
Pull requests: pytorch/FBGEMM
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Add FTRL-Proximal as a fused TBE optimizer (CUDA + CPU)
cla signed
#6295
opened Sep 11, 2026 by
tiankongdeguiji
Contributor
Loading…
Optimize bf16 atomicCAS
cla signed
meta-exported
#6294
opened Sep 11, 2026 by
yvonne-lab
Contributor
Loading…
Avoid FP32 scratch allocation in FP16 communication
cla signed
meta-exported
#6293
opened Sep 11, 2026 by
purnawirman
Loading…
Cap remaining block-bucketize launches
cla signed
meta-exported
#6292
opened Sep 11, 2026 by
q10
Contributor
Loading…
Modify TBE bounds check to support trailing indices
cla signed
meta-exported
#6291
opened Sep 11, 2026 by
JiaJiunn
Loading…
Account for X/Z blocks in grid-Y caps
cla signed
meta-exported
#6290
opened Sep 10, 2026 by
q10
Contributor
Loading…
Mark GenerateEmbeddingSpMDM* is_bf16 params [[maybe_unused]]; use to_float/from_float in QuantUtils (#6237)
cla signed
meta-exported
#6279
opened Sep 9, 2026 by
q10
Contributor
Loading…
2 of 3 tasks
[ROCm] use kWarpSizeHost() in TBE host-side launch configs
ciflow/rocm
cla signed
module: rocm
#6278
opened Sep 9, 2026 by
jeffdaily
Contributor
Loading…
Simplify the warning-flag comments
cla signed
meta-exported
#6265
opened Sep 4, 2026 by
q10
Contributor
Loading…
Add opt-in MI350 pack_segments fused kernel
cla signed
meta-exported
#6263
opened Sep 2, 2026 by
Ruishenl
Contributor
Loading…
[ROCm] Fix GPU memory fault in batched_unary_embeddings backward
ciflow/rocm
cla signed
module: rocm
#6261
opened Sep 2, 2026 by
avbokovoy
Contributor
Loading…
[ROCm] Make the Python NFP8 dtype selection architecture-aware
ciflow/rocm
cla signed
module: rocm
#6260
opened Sep 2, 2026 by
aryaman-gupta
Contributor
Loading…
[ROCm] Fix packed-bag pooling truncation in nbit inference forward
ciflow/rocm
cla signed
module: rocm
#6259
opened Sep 2, 2026 by
aryaman-gupta
Contributor
Loading…
Parallelise the NOBAG inference TBE over row ranges, not tables (#6256)
cla signed
meta-exported
#6256
opened Sep 2, 2026 by
zhaozhul
Contributor
Loading…
Guard optional CPU pragmas
cla signed
meta-exported
#6247
opened Sep 1, 2026 by
q10
Contributor
Loading…
Parallelize reorder_batched_ad_indices within a segment
cla signed
meta-exported
#6245
opened Aug 31, 2026 by
zhaozhul
Contributor
Loading…
TBE backward: consume precomputed index-preproc tensors
cla signed
meta-exported
#6235
opened Aug 27, 2026 by
gchalump
Contributor
Loading…
Add dual-wave TBE benchmark and test harness
cla signed
meta-exported
#6231
opened Aug 27, 2026 by
q10
Contributor
Loading…
Add tbe_bwd_indices_preproc op + reference-impl unit test (#6222)
cla signed
meta-exported
#6222
opened Aug 25, 2026 by
gchalump
Contributor
Loading…
[DO NOT MERGE][DEBUG] Probe GPU round-trip latency on the MI350 VF runner
cla signed
module: rocm
#6219
opened Aug 25, 2026 by
aryaman-gupta
Contributor
•
Draft
fbcode/deeplearning/fbgemm/fbgemm_gpu/test/tbe/inference/inference_converter_test.py
cla signed
meta-exported
#6216
opened Aug 25, 2026 by
meta-codesync
Bot
Loading…
Fix three -Wunused-variable sites in HIP-built sources
cla signed
meta-exported
#6202
opened Aug 23, 2026 by
q10
Contributor
Loading…
[fbgemm_gpu] Add opt-in device sync before permute_2D / jagged_to_pad…
cla signed
#6195
opened Aug 21, 2026 by
RupengWang
Loading…
Use distinct float16 and bfloat16 CPU template types
cla signed
#6194
opened Aug 21, 2026 by
cyyever
Contributor
Loading…
1 of 2 tasks
Previous Next
ProTip!
Filter pull requests by the default branch with base:main.