[ExecuTorch][WebGPU] Replace broad QKV fusion with BK64 kernel - #21131
Open
JCNTH wants to merge 10 commits into
Open
[ExecuTorch][WebGPU] Replace broad QKV fusion with BK64 kernel#21131JCNTH wants to merge 10 commits into
JCNTH wants to merge 10 commits into
Conversation
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21131
Note: Links to docs will display an error until the docs builds have been completed. ❌ 1 Cancelled JobAs of commit b14b60d with merge base ad3a71f ( CANCELLED JOB - The following job was cancelled. Please retry:
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
This was referenced Jul 22, 2026
This PR needs a
|
This was referenced Jul 30, 2026
psiddh
approved these changes
Jul 30, 2026
Contributor
|
Same as other PRs, linter issues |
This was referenced Aug 5, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stack from ghstack (oldest at bottom):
The previously landed broad QKV fusion applied too widely and did not match the
BK64 schedule now used for the ordinary projections. This corrective diff
replaces it with a capability- and geometry-qualified BK64 kernel that fuses the
exact Llama Q/K/V projection triple at live M128, M508, and M512, packing the
constant weights and scales once and scattering the result into three distinct
planner-safe outputs. Outside the accepted shapes it switches atomically back
to the ordinary Steel and bicol routes, and it deletes the obsolete broad
shader and header so only one QKV path remains. Mirrors Vulkan
xplat/executorch/backends/vulkan/runtime/graph/ops/glsl/q4gsw_linear_gemm__w_4x8.glsl
for the per-projection GEMM; the three-output fusion itself is WebGPU-specific.
Key changes:
removes the obsolete q4gsw_linear_gemm_qkv_fused shader and header.
packing, distinct Q/K/V outputs, and atomic fallback to Steel/bicol.
@exported-using-ghexport
Differential Revision: D113171749
Differential Revision: D113171749