Skip to content

Pull requests: vllm-project/flash-attention

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

FA3 to FA4 (with MLA)
#174 opened Jul 30, 2026 by simon-veitner-redhat Draft
Rope dim
#172 opened Jul 23, 2026 by JaredforReal Loading…
Sync upstream
#171 opened Jul 22, 2026 by MatthewBonanni Member Loading…
Refactor FA3 onto the PyTorch stable API
#163 opened Jul 14, 2026 by LucasWilkinson Collaborator Draft
Fa4 fp8 foldscale opt
#159 opened Jul 9, 2026 by qixiang-99 Draft
feat(fa4): fuse per-group fp8 output
#151 opened Jun 18, 2026 by carlyou Loading…
[Fix] Mark kcache and vcache as mutable in fwd_kvcache
#148 opened Jun 15, 2026 by remi-or Loading…
SM100 dynamic causal
#144 opened Jun 10, 2026 by MatthewBonanni Member Loading…
Reapply #122
#137 opened May 6, 2026 by MatthewBonanni Member Loading…
SM100 tile size 64
#132 opened Apr 9, 2026 by MatthewBonanni Member Draft
add support for newer CUDA archs (Spark/Thor)
#121 opened Feb 13, 2026 by askliar Loading…
ProTip! Type g p on any issue or pull request to go back to the pull request listing page.