-
Notifications
You must be signed in to change notification settings - Fork 668
Pull requests: tile-ai/tilelang
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[CUDA][TMA] Separate atomic-add dtype support from layout encoding
#2846
opened Aug 2, 2026 by
LeiWang1999
Member
Loading…
[BugFix] Fix absmax and abssum for uint dtypes
#2845
opened Aug 2, 2026 by
jjppp
Contributor
Loading…
[BugFix] Handle vectorized SelectNode in codegen_cuda
#2843
opened Aug 2, 2026 by
jjppp
Contributor
Loading…
[BugFix][CUDA] Guard cp.async transfers by full source extent
#2842
opened Aug 1, 2026 by
morluto
Contributor
Loading…
[BugFix][Autotune] Drain timed-out benchmark calls
#2840
opened Aug 1, 2026 by
morluto
Contributor
Loading…
[CUDA] Resolve symlinked nvcc before deriving CUDA_HOME
#2839
opened Aug 1, 2026 by
morluto
Contributor
Loading…
[CUDA] Adopt multi-staged buffers in examples
#2836
opened Aug 1, 2026 by
Yongqi-Zhuo
Collaborator
Loading…
[BugFix] Include buffer offsets in TMA descriptor bases
#2833
opened Jul 31, 2026 by
morluto
Contributor
Loading…
[CUDA] Pack logical TMEM buffers into shared
tcgen05.alloc arenas
#2831
opened Jul 31, 2026 by
Rachmanino
Collaborator
Loading…
[Fix] Discover flat CUDA includes for NVRTC
#2829
opened Jul 31, 2026 by
morluto
Contributor
Loading…
[CUDA][Reduce] Restore exact thread image counting
#2825
opened Jul 31, 2026 by
KellyFrog
Contributor
Loading…
[Metal] Derive pointer address spaces from TIR types
#2824
opened Jul 31, 2026 by
GY-Bai
Contributor
Loading…
[BugFix][Transform] Close transitive pipeline dependencies
#2822
opened Jul 30, 2026 by
JayceSu98
Contributor
Loading…
[BugFix][Transform] Detect cross-thread WAW hazards
#2821
opened Jul 30, 2026 by
JayceSu98
Contributor
Loading…
[BugFix][Transform] Preserve nested cast-store guards
#2820
opened Jul 30, 2026 by
JayceSu98
Contributor
Loading…
[BugFix][Transform] Normalize vectorized loop domains
#2819
opened Jul 30, 2026 by
JayceSu98
Contributor
Loading…
[BugFix][CUDA] Fall back for non-MMA GEMM K shapes
#2818
opened Jul 30, 2026 by
JayceSu98
Contributor
Loading…
[BugFix] Support keep-dim destinations in fragment reduce
#2815
opened Jul 30, 2026 by
li-ruinan
Collaborator
Loading…
[Enhancement] Speed up cold parallel/AOT compilation up to ~4x
#2809
opened Jul 29, 2026 by
cklxx
Contributor
Loading…
[BugFix] Respect layouts in logical reductions
#2807
opened Jul 29, 2026 by
lijinpei
Contributor
Loading…
[Feature] Add SM100 MQA logits kernels and native lowering
#2774
opened Jul 27, 2026 by
Rachmanino
Collaborator
Loading…
feat: allow T.const() variables used only in grid dims or computations
#2762
opened Jul 24, 2026 by
NolanHo
Loading…
Previous Next
ProTip!
no:milestone will show everything without a milestone.