Skip to content

[blackwell] sm100: Router GEMM Triton kernel (fused_router_gemm.py) #229

Description

@Andrewxu313

Parent: #204 — Blackwell Phase 2 kernel port
Branch: tairan/blackwell-02-06-router-gemm-triton
Spec: blackwell-kernel-port-v1.md § Sub-task 6

What

Write batchgen_kernels/triton/fused_router_gemm.py:

  • tl.dot BF16->FP32 matmul [N, K_dim] x [K_dim, E]
  • Fused BF16 bias add in FP32 accumulator (HAS_BIAS: tl.constexpr)
  • Output: logits [N, E] FP32

torch.matmul is NOT acceptable — it outputs BF16 and cannot fuse the bias-to-FP32 accumulation.

Wire into batchgen/moe/routing/ dispatch for arch == 'sm100'.
Existing gate_topk_softmax.cu (Kernel B, generic CUDA) reused as-is.

Acceptance

  • logits dtype is FP32 (matching SM90a kernel)
  • max_abs_err < 2e-3 vs float32 reference
  • Router top-k indices identical to H20 reference

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions