forked from Dao-AILab/flash-attention
-
Notifications
You must be signed in to change notification settings - Fork 168
Pull requests: vllm-project/flash-attention
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[WIP/RFC][SM120] Add bounded BF16 D256 paged-decode specialization
#170
opened Jul 22, 2026 by
goodnight654
•
Draft
Title: Migrate FA3 hopper API to the PyTorch stable ABI (Retry)
#165
opened Jul 16, 2026 by
cleonard530
Loading…
feat(cute): add SM90 FP8 KV support with in-kernel dequantization
#164
opened Jul 15, 2026 by
jhaotingc
Loading…
Refactor FA3 onto the PyTorch stable API
#163
opened Jul 14, 2026 by
LucasWilkinson
Collaborator
•
Draft
[FA4] Support paged/decode compile specs in compile_flash_attn_varlen_func_from_specs
#158
opened Jul 9, 2026 by
sfc-gh-goliaro
Loading…
[Bugfix] Fix two SM80/SM120 forward kernel bugs: missing is_split_kv default, mDynamicCausal NameError
#156
opened Jun 30, 2026 by
tgmerritt
Loading…
feat(cute): add SM90 FP8 KV support with in-kernel dequantization
#147
opened Jun 12, 2026 by
qixiang-99
•
Draft
fix: handle suffix-less runtime arch (sm_103/GB300) in SM 10.x gate
#146
opened Jun 11, 2026 by
aoshen02
Loading…
Fix illegal memory access in FA2 varlen SplitKV early-exit LSE write
#139
opened May 18, 2026 by
wangyxbh
Loading…
[Perf] SM103 tcgen05.ld.red for fused TMEM load + row-max in softmax
#131
opened Apr 9, 2026 by
LopezCastroRoberto
Loading…
Combine kernel: increase pipeline depth from 4 to 8 stages
#124
opened Mar 4, 2026 by
jmkuebler
Loading…
Previous Next
ProTip!
Type g p on any issue or pull request to go back to the pull request listing page.