-
Notifications
You must be signed in to change notification settings - Fork 237
Pull requests: NVIDIA/cudnn-frontend
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Add native SM100 d192/d128 SDPA fprop kernel for DSv3 MLA
#488
opened Aug 5, 2026 by
adshen
Collaborator
Loading…
2 tasks done
Add HSTU attention CuTe DSL kernels for Blackwell
#487
opened Aug 5, 2026 by
jiayus-nvidia
Contributor
•
Draft
Add SM120 FROST SDPA backward engine (sdpa_bwd_sm120)
#486
opened Aug 5, 2026 by
Adnios
Loading…
2 tasks done
frost(sdpa): causal right-band widening + per-sequence THD bottom-right diagonal on SM100
cat-enhancements
mod-frontend
cuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.
orig-nv-eng
Reported or requested by NVIDIA engineering.
Fix default-stream race: resolve _get_default_stream(None) to torch's current stream
cat-enhancements
mod-frontend
cuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.
orig-nv-eng
Reported or requested by NVIDIA engineering.
Fix absolute paths in installed CMake configs.
cat-infra
Build, packaging, tooling, dependency, release, or repository maintenance work.
mod-frontend
cuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.
orig-nv-eng
Reported or requested by NVIDIA engineering.
BSA: add Sage FP8 forward support for Blackwell
#475
opened Aug 4, 2026 by
jiayus-nvidia
Contributor
Loading…
Fix ragged SDPA backward workspace under-allocation for non-token-major stats layouts
#462
opened Jul 31, 2026 by
YangXu1990uiuc
Collaborator
Loading…
Support FP32 output and dynamic M in row-scaled FP4 grouped GEMM
cat-enhancements
mod-cutedsl
CuTeDSL kernels, generated kernels, examples, or related integration work.
orig-nv-eng
Reported or requested by NVIDIA engineering.
#461
opened Jul 31, 2026 by
zianglih
Contributor
Loading…
CSA compressor: review-response fixups for the ratio=128 kernels (follow-up to #427)
#452
opened Jul 30, 2026 by
zkyue
Contributor
Loading…
Add MoE + expert-parallel (MoeEp) Python API with MegaMoE CuTe DSL backend
#448
opened Jul 29, 2026 by
mhoqueanik
Loading…
2 of 6 tasks
DSA indexer forward: add a lean SM100 fast path for the head_dim=128 / qhead_per_kv_head=64 regime (up to 1.46x)
#416
opened Jul 21, 2026 by
zkyue
Contributor
Loading…
Reject FP8/MXFP8 SDPA forward combinations exposed to a cuDNN 9.24 split-KV bug
cat-enhancements
mod-frontend
cuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.
orig-nv-eng
Reported or requested by NVIDIA engineering.
Add version-adaptive nvvm.atomicrmw wrapper for cutlass-dsl 4.5.x compat
#411
opened Jul 20, 2026 by
Anerudhan
Collaborator
Loading…
Support installing python bindings via cmake --install (fixes #149)
cat-ci
CI failures, test flakiness, workflow breakage, or automation issues.
cat-infra
Build, packaging, tooling, dependency, release, or repository maintenance work.
mod-infra
Infrastructure, CI/CD, build systems, packaging, releases, or repo maintenance.
orig-external
Reported or requested by an external user, customer, or community contributor.
Add large-tensor convolution fuzzer
cat-ci
CI failures, test flakiness, workflow breakage, or automation issues.
mod-backend
cuDNN backend API, graph execution, descriptors, engines, or backend integration.
orig-nv-eng
Reported or requested by NVIDIA engineering.
Add FLOOR_MOD (floored modulo) pointwise mode
cat-feature
Requests for new functionality, APIs, examples, or behavior improvements.
mod-backend
cuDNN backend API, graph execution, descriptors, engines, or backend integration.
orig-nv-eng
Reported or requested by NVIDIA engineering.
Add MoE + expert-parallel Python API surface, PyTorch reference, and tests
#389
opened Jul 14, 2026 by
Anerudhan
Collaborator
Loading…
[DRAFT] Add framework-agnostic operator APIs with JAX support
#363
opened Jul 8, 2026 by
mgoldfarb-nvidia
•
Draft
benchmark: compute causal SDPA FLOP counts without allocating masks
#347
opened Jul 5, 2026 by
fallintoplace
Contributor
Loading…
Previous Next
ProTip!
Follow long discussions with comments:>50.