Skip to content

[Fix][Relax][Metal] Constrain wide-head prefill tiling - #20235

Open
akaashrp wants to merge 2 commits into
apache:mainfrom
akaashrp:upstream/metal-wide-head-prefill-tiling
Open

[Fix][Relax][Metal] Constrain wide-head prefill tiling#20235
akaashrp wants to merge 2 commits into
apache:mainfrom
akaashrp:upstream/metal-wide-head-prefill-tiling

Conversation

@akaashrp

Copy link
Copy Markdown
Contributor

Apply the existing low-storage WebGPU prefill configuration to Metal for wide attention heads. This keeps generated threadgroup allocations below Metal device limits, including Gemma 4 global attention heads.

Apply the existing low-storage WebGPU prefill configuration to Metal for wide attention heads. This keeps generated threadgroup allocations below Metal device limits, including Gemma 4 global attention heads.
Choose Metal prefill tiling from the target shared-memory limit and the kernel's exact Q/K/V layout. Account for unequal ragged-attention dimensions and MLA's merged KV buffer, and reject configurations that do not satisfy the scheduler's factorization constraints.\n\nShare configuration and factorization helpers with both tree-attention paths while preserving existing WebGPU behavior. Add allocation-level coverage for standard, ragged, MLA, and tree-attention kernels, custom Metal limits, and impossible limits.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant