-
Notifications
You must be signed in to change notification settings - Fork 2.1k
Pull requests: NVIDIA/cutlass
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
CuTeDSL: preserve outer comprehension generator_targets across nested comprehensions (#3499)
#3641
opened Sep 17, 2026 by
SIDDARTHAREDDY8
Loading…
[CuTeDSL] Use vectorized codegen for CUDA 12.9 2CTA grouped GEMM
#3639
opened Sep 15, 2026 by
myu-guo
Contributor
Loading…
Fix maximum_active_blocks wrappers passing int to CudaHostAdapter*
#3638
opened Sep 14, 2026 by
chrisXYZhang
Loading…
2 tasks
Fix: enable CUDA_CTA_RECONFIG_ACTIVATED for SM110a (Jetson AGX Thor)
#3637
opened Sep 14, 2026 by
roma5087
Loading…
[CuTeDSL] Compare missing runtime arguments by identity
#3634
opened Sep 13, 2026 by
voipmonitor
Loading…
Fix misaligned local memory access in 02_dump_reg_shmem
#3633
opened Sep 13, 2026 by
random25160765-collab
Loading…
[GEMM] Fix EllGemm column-major split-K workspace sizing (issue #3539)
#3624
opened Sep 13, 2026 by
XFDG
Loading…
[Conv] Fix filter iterator pointer offset scaling (issue #3504)
#3623
opened Sep 13, 2026 by
XFDG
Loading…
[CuTe] Validate dynamic tiler divisibility for static layouts (issue #3513)
#3622
opened Sep 13, 2026 by
XFDG
Loading…
[Core] Add lowest() to platform numeric_limits<float> (issue #3509)
#3621
opened Sep 13, 2026 by
XFDG
Loading…
[Core] Fix numeric_limits constants for low-precision types (issue #3508)
#3620
opened Sep 13, 2026 by
XFDG
Loading…
[CuTe DSL] Reject direct kernel compilation with a user error (issue #3429)
#3619
opened Sep 13, 2026 by
XFDG
Loading…
Expose resolved GEMM schedules in library descriptions
#3616
opened Sep 12, 2026 by
zupengwang
Loading…
Fix blockwise scale indexing under split-K (#3537)
#3615
opened Sep 12, 2026 by
amacharla15
Loading…
[Bugfix][Blackwell] Avoid fmha_bwd workspace int32 overflow
#3614
opened Sep 11, 2026 by
XFDG
Loading…
Previous Next
ProTip!
Type g p on any issue or pull request to go back to the pull request listing page.