Skip to content

Pull requests: NVIDIA/TransformerEngine

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

surface comm overlap rank errors
#3384 opened Aug 14, 2026 by sudhakarsingh27 Member Loading…
13 tasks
Produce BOLT-compatible libraries
#3383 opened Aug 14, 2026 by fheinecke Collaborator Draft
2 of 13 tasks
[PyTorch] Fuse eager EP prepare and dispatch into a single op
#3380 opened Aug 13, 2026 by phu0ngng Collaborator Draft
13 tasks
fix: duplicate JUnit XML filename overwrites encoder test results community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3376 opened Aug 13, 2026 by andrewwhitecdw Contributor Loading…
fix: SBHD reorder skip uses original shape instead of swapped tensor community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3373 opened Aug 13, 2026 by andrewwhitecdw Contributor Loading…
fix: comment typo 'mas' -> 'mask' in two TODOs community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3372 opened Aug 13, 2026 by andrewwhitecdw Contributor Loading…
fix: make_mask uses deprecated jnp.bool instead of jnp.bool_ community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3371 opened Aug 13, 2026 by andrewwhitecdw Contributor Loading…
fix: Optional segment position annotations in mask helpers community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3369 opened Aug 13, 2026 by andrewwhitecdw Contributor Loading…
fix: Unreachable backend check after earlier backend skip community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3368 opened Aug 13, 2026 by andrewwhitecdw Contributor Loading…
fix: remove redundant self-assignment out_ = out_ community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3367 opened Aug 13, 2026 by andrewwhitecdw Contributor Loading…
fix: simplify qkv_format checks using a membership tuple community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3365 opened Aug 13, 2026 by andrewwhitecdw Contributor Loading…
fix: simplify redundant softmax version condition community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3363 opened Aug 13, 2026 by andrewwhitecdw Contributor Loading…
fix: add missing NVTE_BHSD to qkv format to_string community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3362 opened Aug 13, 2026 by andrewwhitecdw Contributor Loading…
fix: add missing NVTE_BHSD_BHSD_BHSD to layout to_string community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3361 opened Aug 13, 2026 by andrewwhitecdw Contributor Loading…
[PyTorch] Restore FlashAttention 2 head dim support on sm103 community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3360 opened Aug 12, 2026 by kalectory Loading…
8 of 13 tasks
[Common/PyTorch] Fused grouped MXFP8 requantization
#3359 opened Aug 12, 2026 by YangFei1990 Collaborator Loading…
1 of 13 tasks
[Common] Use wide instructions in SR to reduce issue-bound bottleneck community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3357 opened Aug 12, 2026 by janekb04 Collaborator Loading…
6 of 13 tasks
[Pytorch][Do not merge] Let MXFP8 dispatch return plain torch tensor
#3355 opened Aug 12, 2026 by YangFei1990 Collaborator Loading…
13 tasks
[JAX] Optimize MoE block 2.19
#3354 opened Aug 12, 2026 by jberchtold-nvidia Collaborator Loading…
7 of 13 tasks
[PyTorch] Share CUDA graph memory across dynamic CP variants community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3353 opened Aug 12, 2026 by xiaoyao0115 Draft
[PyTorch] Decline fused grouped MLP when the backward format is not E4M3 community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3352 opened Aug 12, 2026 by wilyan09007 Loading…
5 of 13 tasks
[PyTorch] Fused GDN attention 2.19
#3351 opened Aug 12, 2026 by ksivaman Member Draft
9 of 15 tasks
[Pytorch] MOE Sequential Block
#3350 opened Aug 11, 2026 by vthumbe1503 Collaborator Draft
13 tasks
[PyTorch] Fix MXFP8 master weight cast when a rank owns no shard. community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3348 opened Aug 11, 2026 by rapatel Loading…
8 of 13 tasks
[PyTorch] Make the Triton mask-map permute respect num_out_tokens community-contribution PRs from external contributor outside the core maintainers, representing community-driven work.
#3347 opened Aug 11, 2026 by truong-v Loading…
2 tasks done
ProTip! Filter pull requests by the default branch with base:main.