Skip to content

[BUG] SM120 dense HSTU smem overflow #492

Description

@Clebrate

Describe the bug

  • On SM120 (RTX PRO 6000), dense HSTU varlen fwd (head_dim=256, bf16, no paged KV) fails with CUDA error: invalid argument at hstu_fwd.h:830.

  • Actual tile is Arch80 {128,96,8} (160KB smem). SM120 per-block opt-in is 99KB.

  • Default --disable_kvcache with CUDA graphs on still works: capture/replay is paged (page_size=32 → {128,32,8}, 96KB).

Steps/Code to reproduce bug
Set use_cudagraph=False in inference_benchmark.py, then: cd examples/hstu && python3 ./inference/benchmark/inference_benchmark.py --disable_kvcache

Expected behavior

  • SM120 dense should use {64,64,4}.

  • Map arch 120 → 89 in ARCH_SWITCH.

Environment details (please complete the following information):

  • RTX PRO 6000 Blackwell
  • Docker: devel_latest

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions