Skip to content

dasLLAMA: Pocket's text prompt runs on the device on both GPU homes, and thirteen TTS kernel pairs become shared templates both homes stamp - #4216

Merged
borisbat merged 14 commits into
masterfrom
bbatkin/tts-shared-kernels
Oct 6, 2026
Merged

borisbat merged 14 commits into
masterfrom
bbatkin/tts-shared-kernels

Conversation

@borisbat

@borisbat borisbat commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

Behavior change

Pocket's text prompt runs on the device on both GPU homes. The prompt stage - the voice's text through the backbone into every layer's K/V rows - ran on the CPU under Metal and under Vulkan; only the codec and frame seats were on the device. The PocketGpuDriver gains a prompt seat (vulkan_pocket_prompt / metal_pocket_prompt), served by the same rows transformer the codec seat runs (ts_pk_rows_tf / pk_rows_tf under a rows form), and the K/V rows read back to the host cache (48 KB a position, 1-4 MB a chunk). The device rows stay resident on the voice slot (TtsPkPromptDev, keyed by the embedding rows, spent by the frames seat at entry), so the frames seat reads them without a second upload. The CPU fallback under kv_only_last skips the last layer's output projection and MLP, which nothing after the stage reads.

Thirteen TTS kernel pairs fold into dasllama_gpu_kernels_common.das templates both homes stamp: GkRopeTab, GkPkRows, GkSigSum, the harmonic source family (GkSrcLaw/GkSrcLow/GkSrcCumsum/GkSrcNoise/GkSrcSines), GkStft, GkIstft, GkAddScale, GkAdain, GkDwConvRows, GkPkAttn over a shared GkWgReduce (the workgroup reduce both homes derive, its reductions as statement methods because the MSL emitter refuses a multi-statement template method in value position; the per-home lane primitive gk_subgroup_add/max is subgroupAdd on Vulkan and simd_sum on Metal). tts_div moves to dasllama_gpu_math.das. Metal gains MetalPkRopeTab (the prompt's rope from the voice tables) and a one-call st2_add.

Validation

Request wall, Pocket q8, one sentence (ms, master -> branch; cv under 3% on every kept row):

box request prompt backbone
M5 Max Metal 67.5 -> 49.5 15.5 -> 2.1 38.4 -> 34.2
RunPod Vulkan 114.7 -> 68.1 50.0 -> 3.8 50.2 -> 49.2

Kokoro and Kitten mini unchanged on both boxes (73.5/64.5 ms M5; 48/47 ms pod, the pod's kitten rows void on cv). Whisper turbo q8 served A/B on the pod 48.8 -> 48.5 ms (the shared reduce sits under its tower stamps).

Quality rig (200 sentences, WER % / UTMOS, before -> after): Pocket f32 4.18/4.373 -> 3.95/4.368, q8 4.41/4.363 -> 4.23/4.354, kq 3.91/4.325 -> 3.77/4.314, stuart-kq 3.32/4.137 -> 3.27/4.139; kitten-nano, kitten-mini, kokoro identical. K/V parity of the prompt seat against the CPU chain: f32 5e-7 / 1.6e-6, q8 3e-3 / 1.2e-2, kq 7e-3 / 2.8e-2 (rel-l2 K / V) on both boxes.

Stamp diffs by kind (never source): SPIR-V - 24 TTS stamps moved (+52 bytes for the template's call, dropped element offsets, +208 bytes for the WgReduceBase wrappers on 5 derivers), 8 tower stamps moved on the pod; MSL - 18 moved, pk_rope_tab and st2_add appeared, MetalSt2Add_metal_ew2 vanished. Kernel census on the M5: metal_pk_rope_tab 34, metal_st2_add 157, metal_pk_attn 400, no TTS zeros.

Tests: Metal pocket 23/24 (1 skip), kokoro 17/17, kitten 18/18, the kernels suite 7 files 0 failed; Vulkan on the pod - the five TTS cell files 20/15/12/11/28, pocket 23/24, kokoro 17/17, kitten 18/18, and at the tip test_vulkan_kernels 205/206 (1 skip), test_vulkan_dec_tail 5/5, test_vulkan_kv_codec_kernels 12/12, test_vulkan_moe_cm2 8/8 (the shared reduce's derivers), kitten 18/18, kokoro 17/17. New cells: test_pocket_prompt_gpu (every layer's K and V rows against the CPU chain on three files, the text rotated by one token as the control), test_pocket_prompt_rules (model-free admission and residency rules), pk_rope_tab_gate, the Vulkan cells over the renamed Gk args.

Claims the tests do not prove

  • The CPU fallback's last-layer cut is output-neutral by construction (nothing reads the dropped rows); no cell asserts its absence.
  • A device or gpu_error decline on the prompt seat has no seam to force; the knob-off decline is the covered leg.
  • pk_rows_lin's bias arm is unreachable with the stocked files (no Pocket file carries a biased backbone GEMM).
  • The bxh half path runs under the sidecar's crown only.
  • Two existing cells were retuned: the prompt admit fixture (5+15 rows at a 20 cap read as fitting; the cap is now 16) and the Vulkan cell planes at base 0 for the Axpy-style args.
  • The stocked runs were scoped (tts, kernels, the Vulkan files above), not the whole --changed set.

Not in this PR

  • The work-split pairs (Elem/PkRowScale/GeluTanh, Reflect1, ColStats, PkGemvT) and the different-algorithm pairs (LSTM, ALBERT attention) stay per home - followup_vulkan.md row 115 carries them with the profile each fold owes.
  • RowGather (the LLM embed) and Im2col (the whisper stem) fold in the LLM and audio PRs of this arc.
  • No comment harvest or style-hygiene round this PR; no nightly run asked.
  • Lint ideas: a REVIEW.das gate for a push-constant *off field no site sets, and one for the integer-division guard the Vulkan rule describes.

Copilot AI balanced review requested due to automatic review settings October 6, 2026 03:13

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

It rewrites hand-written GPU kernel math into shared templates stamped on two backends and adds device cache-residency/readback logic whose cross-backend parity and correctness can only be confirmed by running the GPU test suites on real Metal/Vulkan hardware.

Review effort: Balanced
Findings: None

What changed in this PR

This PR extends the dasLLAMA Pocket TTS path so the text-prompt prefill stage runs on-GPU on both backends (Metal and Vulkan), where previously only the codec and frame stages were device-resident while the prompt ran on the CPU. It adds a prompt seat to PocketGpuDriver (served by the same rows-transformer as the codec seat under a "rows form"), reads the computed K/V rows back into the host cache, and leaves them resident on the voice slot (TtsPkPromptDev, keyed by embedding rows and gen) so the frames seat consumes them without a second upload. In parallel, it folds thirteen per-backend TTS kernel pairs into shared class templates in dasllama_gpu_kernels_common.das that both homes stamp (over a shared GkWgReduce), moves tts_div to dasllama_gpu_math.das, and adds Metal's MetalPkRopeTab/st2_add.

Changes:

  • New GPU prompt seat + host-cache readback + device residency record (TtsPkPromptDev) spent by the frames seat on entry; CPU fallback skips the last prompt layer's out-proj/MLP via kv_only_last.
  • Thirteen TTS kernels consolidated into shared Gk* templates (GkPkRows, GkPkAttn, GkAdain, GkDwConvRows, GkStft/GkIstft, the GkSrc* family, GkRopeTab, GkSigSum, GkAddScale) with per-home lane primitives.
  • New tests (test_pocket_prompt_gpu, test_pocket_prompt_rules, pk_rope_tab_gate), retuned fixtures (removed per-plane element offsets), and extensive architecture/review doc updates.
File Description
dasllama/​dasllama_gpu_kernels_common.das Core of the PR — the new shared Gk* kernel class templates both backends stamp, plus GkWgReduce.
dasllama/​dasllama_vulkan_tts.das New vulkan_pocket_prompt seat, prompt readback, generalized ts_pk_rows_tf/TsPkRowsForm, frames seat takes residency on entry.
dasllama/​dasllama_metal_kernels.das Metal classes now derive Gk* bases; adds MetalPkRopeTab/st2_add, lane primitives gk_subgroup_add/max.
dasllama/​dasllama_pocket.das CPU kv_only/kv_only_last cut, PocketGpuDriver.prompt seat, pocket_prompt_embed/pocket_prompt_rows.
dasllama/​dasllama_tts_slab.das Prompt admission + residency (tts_pk_prompt_admit, TtsPkPromptDev, _take/_set, tts_pk_prompt_emb_key).
dasllama/​dasllama_tts_blocks.das kv_cache_append delegates to new kv_cache_append_rows (raw-pointer rows).
dasllama/​dasllama_gpu_math.das tts_div moved here (Newton-refined IEEE-exact divide for cross-backend parity).
dasllama/​dasllama_vulkan_classes.das /​ dasllama_metal_*.das Vulkan/Metal stamp leaves retargeted onto shared Gk* bases and args.
tests/​test_tts_pocket.das New prompt-seat tests + updated encode-count expectations.
tests/​test_metal_prefill_kernels.das pk_rope_tab_gate, Metal gates rewritten to Gk*Args + joined weight buffers.
tests/​test_vulkan_tts_*.das, test_vulkan_tower_kernels.das Vulkan cells migrated to renamed Gk*Args, fixtures retuned (offsets removed).
ARCHITECTURE_*.md, REVIEW_*.md, PERF_LEDGER.md, followup_*.md, tests/​* Doc/ledger/review updates reflecting the prompt seat and shared kernel classes.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Copilot AI balanced review requested due to automatic review settings October 6, 2026 15:17

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

It is a large, correctness-sensitive GPU refactor adding a new on-device code path across Metal, Vulkan, and CPU that cannot be built or validated in this environment and warrants human review.

Review effort: Balanced
Findings: 1 High severity

Open (1)

Comment thread modules/dasLLAMA/dasllama/dasllama_gpu_kernels_common.das Outdated
Copilot AI balanced review requested due to automatic review settings October 6, 2026 15:30

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The change folds thirteen GPU kernels into shared cross-backend templates and adds a new device prompt-seat/residency path whose correctness rests on on-device numerical parity that cannot be verified in this environment, so it needs final human review.

Review effort: Balanced
Findings: None

Resolved since last review (1)

borisbat and others added 13 commits October 6, 2026 12:36
… a third seat beside codec and frames, the frames slab's layers in the rows form at the voice's positions, the keys and values into the voice slot and back to the host caches, the last layer ending at them since only the caches are read after a prompt (the CPU chain too); the rope that takes a position base is one class template both homes stamp

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ER_POCKET.md (the parent past the line cap), the prompt seat described on both homes, the ledger rows it answers closed, the prompt cell's lint

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…te - the Pocket row copy, the duration sigmoid sums, the source cumsum and noise draw, the STFT on both pad laws, the inverse STFT and the residual add-and-scale; the one source argument struct carries the slab offsets both homes index, the dead Metal totals and the Vulkan add-scale offsets no site passed are gone

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… one template both GPU homes stamp on either resample law; the IEEE-correcting divide moves to the shared math, where it lands the same quotient over Metal's divide

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ion row are templates both GPU homes stamp - the attention's workgroup reductions as statement methods over each home's subgroup primitive (gk_subgroup_add / gk_subgroup_max), which the MSL emitter can splice; the access classifier resolves a helper a body hands a buffer to in the shared module as well; the ledger row names the pairs that stay two and why

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ry, so a decline cannot leave it for a later chunk; the dupe audit's folds - the workgroup reductions as one GkWgReduce both chains derive, the Metal codec transformer and prompt as one rows-form loop over frame-slab layers, the CPU prompt cut as a flag of transformer_rows, the two cache admissions over one

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…templates declare weight, the AdaIN eps is the body's literal again, the prompt's residency record names the slot build and the embedding rows it came from, a Metal cell for the position-based rope, the family-move lines, the test verb cold, the stale seat counts and ledger citations, the checklist rules that could not decide a shared kernel

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…l-free cell; the access classifier's lookup in the shared module goes, nothing hands a buffer to a helper there since the reductions became methods

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… both boxes, the whisper row the shared reduce base owes, the quality rig before and after

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…dits - the shared reduce as its own rule, the home rules back to their obligations, the paths to the common module, the division rule's guard decided lexically and under the cap, the encode-stage rule by what later code reads

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… a const rope-table argument record, a float increment

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… spawn under a 480 s cap - the windows nightly runner compiles the facade in about two minutes, so two spawns at a 120 s cap sat on the edge and tipped over

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…y it - the guard compared the product of frames and bins, which the folder's division rule rules out as a guard

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@borisbat
borisbat force-pushed the bbatkin/tts-shared-kernels branch from 9fd89fd to 23458ea Compare October 6, 2026 19:41
Copilot AI balanced review requested due to automatic review settings October 6, 2026 19:41

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

It makes large cross-backend GPU kernel and on-device prompt-seat changes whose correctness rests on GPU-hardware parity validation that cannot be reproduced in this environment, warranting final human review.

Review effort: Balanced
Findings: 1 Low severity

Open (1)

Comment thread modules/dasLLAMA/tests/test_tts_pocket.das Outdated
…xt rotated by one token - and the rotated ids carry that name

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Copilot AI balanced review requested due to automatic review settings October 6, 2026 19:51

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

It spans numerically parity-critical GPU kernels on two backends plus a device-residency lifecycle and CPU fast-path cut whose correctness rests on hardware A/B runs that cannot be verified here, warranting final human review.

Review effort: Balanced
Findings: None

Resolved since last review (1)

@borisbat
borisbat merged commit 3730bcd into master Oct 6, 2026
39 checks passed
@borisbat
borisbat deleted the bbatkin/tts-shared-kernels branch October 6, 2026 20:44
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants