Describe the bug
-
On SM120 (RTX PRO 6000), dense HSTU varlen fwd (head_dim=256, bf16, no paged KV) fails with CUDA error: invalid argument at hstu_fwd.h:830.
-
Actual tile is Arch80 {128,96,8} (160KB smem). SM120 per-block opt-in is 99KB.
-
Default --disable_kvcache with CUDA graphs on still works: capture/replay is paged (page_size=32 → {128,32,8}, 96KB).
Steps/Code to reproduce bug
Set use_cudagraph=False in inference_benchmark.py, then: cd examples/hstu && python3 ./inference/benchmark/inference_benchmark.py --disable_kvcache
Expected behavior
-
SM120 dense should use {64,64,4}.
-
Map arch 120 → 89 in ARCH_SWITCH.
Environment details (please complete the following information):
- RTX PRO 6000 Blackwell
- Docker: devel_latest
Describe the bug
On SM120 (RTX PRO 6000), dense HSTU varlen fwd (head_dim=256, bf16, no paged KV) fails with CUDA error: invalid argument at hstu_fwd.h:830.
Actual tile is Arch80 {128,96,8} (160KB smem). SM120 per-block opt-in is 99KB.
Default --disable_kvcache with CUDA graphs on still works: capture/replay is paged (page_size=32 → {128,32,8}, 96KB).
Steps/Code to reproduce bug
Set use_cudagraph=False in inference_benchmark.py, then: cd examples/hstu && python3 ./inference/benchmark/inference_benchmark.py --disable_kvcache
Expected behavior
SM120 dense should use {64,64,4}.
Map arch 120 → 89 in ARCH_SWITCH.
Environment details (please complete the following information):