The official repo of Qwen (通义千问) chat & pretrained large language model proposed by Alibaba Cloud.
-
Updated
Mar 5, 2026 - Python
The official repo of Qwen (通义千问) chat & pretrained large language model proposed by Alibaba Cloud.
LeetCUDA: Modern CUDA Learn Notes with PyTorch for Beginners, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.
Official release of InternLM series (InternLM, InternLM2, InternLM2.5, InternLM3).
📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉
MoBA: Mixture of Block Attention for Long-Context LLMs
cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.
rCM & Causal-rCM: Leading and Unified Algorithms/Infrastructures for Bidirectional/Autoregressive Video Diffusion Distillation at Scale
Trainable fast and memory-efficient sparse attention
From teacher to tiles — a from-scratch LLM distillation & serving engine: custom Triton/CUDA kernels, FSDP distillation, paged-KV continuous batching, speculative decoding, a Rust gateway, a JAX oracle, and interpretability tooling.
InternEvo is an open-sourced lightweight training framework aims to support model pre-training without the need for extensive dependencies.
FFPA: Kernel Library for Large Headdim Attention - 1.5x~6x speedup over PyTorch SDPA.
[CVPR 2025 Highlight] The official CLIP training codebase of Inf-CL: "Breaking the Memory Barrier: Near Infinite Batch Size Scaling for Contrastive Loss". A super memory-efficiency CLIP training scheme.
Windows-only version of ComfyUI which uses AMD's official ROCm and PyTorch libraries to get better performance with AMD GPUs. [auto-installation and popular performance enhancing packages like triton * sage-attention * flash-attention * bitsandbytes included ]
Implementation of FlashAttention (FA1-FA4) in PyTorch for educational and algorithmic clarity
The official repository of FlashSinkhorn [ICML 2026 Oral]
Implementation of FlashAttention-2 for Nvidia Tesla V100 / Titan V
Triton implementation of FlashAttention2 that adds Custom Masks.
Pre-built wheels that erase Flash Attention 3 installation headaches.
Add a description, image, and links to the flash-attention topic page so that developers can more easily learn about it.
To associate your repository with the flash-attention topic, visit your repo's landing page and select "manage topics."