You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
dasLLAMA: the Vulkan dedup pass (followup_vulkan rows 117 / 133 / 138 and the full sweep) - #4210
The Vulkan arc (#4152 / #4153 / #4163 / #4166 / #4187 / #4200) left three dedup rows in followup_vulkan.md and a tree where the same shape lived in two to fifteen places: kernel cells copied per format, resident-driver helpers copied per family, decode and prefill recorders copied per rail. Every later format or family paid the copy count. This pass closes the arc's dedup debt in one PR, as ruled: all tier A rows, all tier B rows but the prefill GEMM set caches (row 139), the probe-first kernel rows last and gated on a probe row within noise, the sweep's defects fixed in-PR each with a cell.
What
15 sweep shards + one merge over 3586 functions read (daslib / utils / modules corpus 38,742): 127 duplicates, 157 templatable sets, 123 local blocks -> 348 ledger rows; A 247 folded, B 76 of 77 folded (row 139 ledgered), C 24 kept by ruling.
One commit per fold: 304 commits over 131 files, +10054 / -15353 lines against origin/master (net -5299).
Defects the sweep found, fixed with a cell each: the ungated DaAttnBH128T, the would_serve gap, a NaN passing a bar, the gemm probe's k6 / iq scales and deltanet layout, seven TTS seats logging nothing, ... (F1-F21).
The review round's fix batch: the trim-lane regression (a run-mode trim decline writes no lane, on both the streaming and the RAM-bake arms), target_logp reading a NaN as NaN, the cm2 ladder gate reading the engine's ladder, [boot_restore] refusing a function global below it, the f16 wide-head attention cell, mad on the bit-for-bit paths, the lens refusing a * prefix naming two methods, [hot_path] on the TTS ensures, the workgroup-cap decline memo, the _vkd_oracles area map, the kernel-cells checklist's four violations.
Two folds taken back on the served row: the mm prefill tiles stay two bodies (one coopmat accumulator array cost 2.5% of Qwen3-4B Q8_0 pp512 where its 4096x4096 probe had read flat) and the Q8 cm2 decode keeps its 16-bit-lane unpack8 select (the shift form cost 1.5%); both are measured forks in row 93 and the architecture doc.
Ledger: rows 117 / 133 / 138 / 77 deleted, 90 / 92 / 93 / 97 / 103 trimmed, rows 139-143 added or renumbered; rows renamed to current symbols.
Observable
Behaviour: under DASLLAMA_PIN_PREFILL, a foreign override, the route off, or a plan past the card, DASLLAMA_TRIM mints nothing where it minted an untrimmed lane under the trimmed identity.
Everything else is meant to be invisible: SPIR-V and bench rows below.
SPIR-V (harness/vk_spv_diff.das over DASLLAMA_VK_SPV_DUMP dumps of the model-free suite), master f3a03fa (352 stamps) against the tip (349): 145 identical, 202 moved, 2 appeared (da_attn_bw_f16 - a cell builds it now; q8_gemv_gu_n1 - the N-column family replaces q8_gemv_gu), 5 vanished (the five classes the folds removed: cls_ar_rq, f16cvt, q8_gemv_gu, topk, tower_rms). The 202 moved by opcode delta: 84 unrolled copies folded into helper functions (OpFunctionParameter up, call plumbing only); 12 x * sigmoid(g) -> x / (1 + exp(-g)); 3 the gated q panel on the h128 / wide decode attention (a sweep defect, F1: a feature); 2 fused multiply-adds; 101 kernel-helper moves (29 cm2 tiles lose a dynamic vector extract, 18 khr tiles lose six packHalf2x16 - five also the sign fold's twelve negations, 9 clamp + fma, 7 zero-init as OpConstantNull, 7 opcode-multiset identical, 7 norm stamps norm by construction, the rest listed in the classification file). The rename commits alone: 241 identical, 0 moved.
Kernel cells: test_vulkan_kernels 206 / 205 passed / 1 skipped (RTX 5060 Ti), bit for bit where tests/REVIEW_KERNEL_CELLS.md rules it; the model-free suite 85 files / 0 failed / 5 cells skipped at the tip.
Local -jit chain at the fix batch: vulkan_tier 36/36, gpu_tier 25/25, facade_docs 7/7, gpu_resident_hybrid 69/69, vulkan_mint 8/8, mtp 18/18, gpu_resident_gemma3_1b 12/12.
Pod (RTX PRO 4500 Blackwell, tests/run.das, DASLLAMA_GPU=1 DASLLAMA_PARITY_FULL=1) at ecf18ce: model-free 85 files (the one red was the facade stub, fixed in 09273ed); stocked 87 files, 1 red = test_vision_chat.das killed at exit 137 by the container's 62 GB cgroup, on master's binary the same (followup_general.md row 181: the process reaches 57 GB resident on both trees); coverage (--arm coverage-vk) 1 file, 6 cells / 4 passed / 2 skipped; audio+vision 20 files, the same vision_chat red, 82 cells skipped; tts 17 files / 0 failed / 47 skipped; model_image 90 / 77 / 13 skipped and model_image_tts 90 / 68 / 22 skipped; the per-file cells with DASLLAMA_PARITY_FULL=1 (cells / passed / skipped, 0 failed each): gpu_resident_hybrid 69 / 65 / 4, gpu_resident_moe 17 / 6 / 11, gpu_resident_hc 6 / 3 / 3, mtp 18 / 15 / 3, gpu_tier 25 / 25, gpu_serving_declines 56 / 56, gpu_slot_swap 2 / 1 / 1, gpu_model_swap 6 / 5 / 1, vulkan_mint 8 / 7 / 1, model_image_vulkan 4 / 2 / 2, regions_tq4kv 26 / 26, regions_q8kv 26 / 26, regions_qwen3 12 / 12, regions_hybrid_k 4 / 4, gpu_resident_llama 14 / 12 / 2.
Bench rows (lcpp_bench.das --ngl 0 -p 512 -n 128 -r 5 -t 16, one process a row, master 1a0729a then the tip on the same card and lane), tg128 master / tip: Llama 1B Q4_K_M 554.87 / 551.73, Qwen3-4B Q8_0 146.54 ± 12.62 / 152.21 ± 0.22, gemma-4 E4B Q8_0 98.33 ± 15.27 / 112.00 ± 0.05, gemma-4 12B Q4_K_M 82.12 / 82.18, Qwen3.5-9B MTP 105.20 / 105.07, Qwen3.8-27B tq4 K/V 36.97 / 36.96, Qwen3.8-27B q8_0 K/V 36.98 / 37.01 (the ±12–15 spreads land on either tree at random: a stall, not a tree). pp512 read 3.7% and 2.8% under on the two Q8_0 models with tight spreads, held by a 10-rep recheck; the per-stage profile and an A/B named two folds (the mm tile's accumulator array, the Q8 cm2 decode's shift form), both taken back in d922920. At the tip, 10 reps, master-tip-master-tip: Qwen3-4B Q8_0 pp512 11765 / 11753 / 11508 / 11754, gemma-4 E4B Q8_0 9610 / 9584 / 9575 / 9587. The entry is in PERF_LEDGER.md.
make-pr chain (utils/internal/make-pr/main.das --, the full preflight once): sync, stamp-reach, jit-smoke, untracked, format, hash-refs, review-md, review-md-tests, md-ascii, ast-verify (112 files), ci-das (27 files), compile-sweep (768 roots) green; lint red once on a 302-line architecture doc, re-wrapped (36af567) and the lint lane re-run alone: 130 files clean on both rails. The six lanes the red fast tier skipped, each run once: docs (sphinx-html) green, tests-cpp green, tests-aot green; tests-interp 14836 / 1 failed and tests-jit 14752 / 2 failed = tests/watchdog/test_watchdog.das (80 s) and tests/jit_tests/jit_lib.das (90 s) past dastest's 60 s per-file cap under 46 workers - alone they run 76 s and 59 s and pass 52 / 2 skipped and 7 / 7; this box's standing per-file-budget reds on files the branch does not touch. utils-tests: all 14 suites green (0 failed), the lane red is msbuild reading the daspkg and mcp-setup suites' expected error: refusal lines (Windows only; CI's lane is Linux make).
Claims
The folds keep every kernel's SPIR-V or change it only by the classified moves above.
Ledger rows the .md stop kept from the deleted dedup rows: 140 (the trim's seven base plane names spelled twice), 141 (the Metal prefill's tq4 un-rotate site, owed an Apple run with the M5 Metal cells), and row 93's four probe-first candidates plus the Q8 byte-store pair.
test_vision_chat.das on the pod: killed by the container's 62 GB memory cap on master and on this branch alike (followup_general.md row 181).
~160 checklist self-review proposals, and tests/CLAUDE.md's shape (it is the tests folder's architecture doc at 2690 lines with no anchors; the cell catalogue belongs in tests/ARCHITECTURE.md companions, and its history passages go): the owed rule-doc PR.
…ayers`. The f32 beta/alpha sizing, the verify rollback's sizing and the resident upload each counted the trunk's recurrent layers with their own loop (the upload counting the attention layers and subtracting, beside a first-attention-layer index nothing read); `resident_recurrent_layers(c)` counts them, the upload's no-attention decline and its attention count read it, and the unread index goes. Ledger row 157 (E3 L6). Proof, the tier armed: test_vulkan_tier.das 36 passed, test_gpu_serving_declines.das 56 passed, test_mtp.das 18 passed (the hybrid's rollback sizing), no skips; the resident batch runs after this group of folds.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…of`. `kvdt_name` spelled each `KVDtype`'s name by ordinal beside `kv_dtype_of` (dasllama_kv_dtype), whose value interpolates to the same name; the codec pass's text interpolates `kv_dtype_of(int(a))` and the armed codec itself, and `kvdt_name` goes. Its "codec N" fallback for an unknown ordinal goes with it: every caller passes a session's own codec, and `kv_dtype_of` panics by name on any other. Ledger row 175 (E3 D5). Proof, the tier armed: test_vulkan_tier.das 36 passed, test_gpu_serving_declines.das 56 passed, no skips.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…on maxima. `resident_head_decline` walked the trunk's attention layers for the widest head and K/V width, the walk `resident_attn_maxima` already is; it reads `resident_attn_maxima(c)`, the decline text unchanged. Ledger row 180 (E3 D4). Proof, the tier armed: test_vulkan_tier.das 36 passed, test_gpu_serving_declines.das 56 passed, test_mtp.das 18 passed (DASLLAMA_PARITY_FULL=1, the head's admission), no skips.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… per-op walk's streamed-slot reserve skipped a layer by spelling `layer_is_moe`'s test (past the dense lead, inside the expert offsets, all three stacks present) inline; it calls `layer_is_moe(t, sl)`, whose expert-count guard the enclosing `n_expert > 0` already holds. Ledger row 226 (E3 D8). Proof, the tier armed: test_vulkan_tier.das 36 passed, test_gpu_moe_shexp.das 3 passed (DASLLAMA_PARITY_FULL=1, the per-op MoE walk), no skips.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…se_plane_ok`. The dense rail's q/o check and the attention-quad rail's per-plane check each spelled "Q8_0 or a superblock format the box serves" inline beside `dense_plane_ok`, which is that predicate; both call it. Ledger row 233 (E3 D6). Proof, the tier armed: test_vulkan_tier.das 36 passed, test_gpu_serving_declines.das 56 passed, no skips.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ost_experts_step` spelled the pool test (the layer inside the pool table, its pool on) inline beside `rdec_hot_on(l)`; it calls it. Ledger row 244 (E3 D7). Proof, the tier armed: test_vulkan_tier.das 36 passed; test_gpu_resident_hc.das (the hot pool's step) runs in the resident batch after this group of folds.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… plan's pool fit and the pool's arming each spelled the asked slot count (`DASLLAMA_GPU_HEAT`, else the want's count, else the auto arm); `resident_hot_ask()` answers it for both. Ledger row 238 (E3 L9). Proof, the tier armed: test_vulkan_tier.das 36 passed, test_gpu_tier.das 25 passed (the pool's slot rules); test_gpu_resident_hc.das runs in the resident batch after this group of folds.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…walk_stops`. The resident-expert walk and the streamed-expert walk each read a layer's three expert formats, tested gate and up for one activation image and down for a tier form, and logged the same stop line but for the walk's name; `tier_expert_walk_stops(t, l, walk)` reads, tests and logs it, the line word for word as before, and each walk breaks on it beside its dense-lead and missing-plane tests. Ledger row 205 (E3 L4). Proof, the tier armed: test_vulkan_tier.das 36 passed, test_gpu_moe_shexp.das 3 passed (DASLLAMA_PARITY_FULL=1, the per-op expert walk), no skips.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…_place_experts_or_host`. The resident upload's recurrent-layer and attention-layer arms each computed whether the router slot falls among the host layers, passed it in and tested the result against it again; the helper takes the router slot and the host-layer count and answers `ok`, each arm keeping its own decline text. Ledger row 181 (E3 L11). Proof, the tier armed: test_vulkan_tier.das 36 passed; test_gpu_resident_moe.das (the host-layer form) runs in the resident batch after this group of folds.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ice_off`. The slice's two exits - the walk past the plan's end, and a gather the plan's entry does not match - each took the tier's refusal for a gather it would refuse anyway, else warned, turned the slice off and regathered live; `vk_bake_slice_off(n, rows, fmt, wq, ws, why)` does it, each exit passing the words of its warning, the lines as before. Ledger row 179 (E3 L1). Proof, the tier armed: test_vulkan_mint.das 6 passed, test_model_image_vulkan.das 4 passed (the mapped lanes slice the plan), no skips.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…The one-row step and the verify rows' step each walked their picks, placing every expert the pool did not hold into the slot `rdec_hot_slot_for` offered until the step's fill or swap budget ran out (scaled by the rows on the verify); `rdec_hot_fill(t, l, H, idx, n, nrows)` is that loop. Ledger row 154 (E3 T12). Proof, the tier armed: test_vulkan_tier.das 36 passed; test_gpu_resident_hc.das (the pool's hits, picks and uploads) runs in the resident batch after this group of folds.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…, the page pad through `w_page`. `pad_to_page` and the vulkan twin's two appends (the mint sink and the RAM blob) each spelled `(x + a - 1) / a * a` beside daslib/math_boost's `round_up`, which they now call; the image writers that padded to a page and then read the cursor (the plane writer, the meta tail, the vulkan mint walk, the streamed section, the metal blob plane) call `w_page(w)`, which pads and answers the offset. The bytes do not move, so IMAGE_VERSION stays; the layout closure's text moved, so modules/dasLLAMA/REVIEW.das re-stamps IMAGE_LAYOUT_STAMP_HASH (0x4e64f128eb806584 -> 0x52e01c586b1c4c6f) and lists `w_page` among the layout helpers, since it decides where a section starts. dasllama_whisper.das:625 spells the same pad pair and is not this cluster's file. Ledger row 123 (X45: E3 D3, E4 D4; the test members are the test clusters'). Proof: REVIEW.das clean after the re-stamp; the seven-lane mint of Qwen2.5-0.5B-Instruct-Q8_0 sha256 identical; under DASLLAMA_GPU=1 test_vulkan_mint.das 6 passed, test_model_image_vulkan.das 4 passed, test_model_image.das (arms mechanics, smol, untied) 66 passed 24 skipped.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…The bake collected an entry through one of two appends - the streaming mint's sink (`vk_mint_sink_append`, zero pads written through the sink) or the RAM blob's own aligned append - and `vk_bake_collect` branched between them; the RAM blob is now `vk_bake_blob_sink`, a `VkMintSinkFn` that `vulkan_bake_begin` installs (its 2 GB-step growth policy kept), every entry goes through `vk_mint_sink_append`, and the branch and the RAM append go. The blob's bytes are the same: the pads the RAM append left to `resize`'s zero fill now arrive as written zeros. Ledger row 153 (E3 T9, row 117f). Proof: REVIEW.das clean (the layout closure unchanged); the seven-lane mint of Qwen2.5-0.5B-Instruct-Q8_0 sha256 identical (the converter's `-f vulkan` lanes bake through the RAM blob); under DASLLAMA_GPU=1 test_vulkan_mint.das 6 passed, test_model_image_vulkan.das 4 passed (the eager dry bake against the minted plan, entry for entry and byte for byte).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…_u32` and `store_u64` differed only in the stored type; `store_at(b, at, v)` is the generic over the value's type, and the header writer's nine stores call it with the same values at the same offsets. The bytes do not move, so IMAGE_VERSION stays; the layout closure's text moved, so modules/dasLLAMA/REVIEW.das lists `store_at` among the layout helpers in their place and re-stamps IMAGE_LAYOUT_STAMP_HASH (0x52e01c586b1c4c6f -> 0x62c76ef2581a3a96), and ARCHITECTURE_IMAGE.md's closure list names `store_at` - and `w_page`, which row 123 added to the closure and the list left out. Ledger row 161 (E4 T5). Proof: REVIEW.das clean; the seven-lane mint of Qwen2.5-0.5B-Instruct-Q8_0 sha256 identical; under DASLLAMA_GPU=1 test_vulkan_mint.das 6 passed, test_model_image_vulkan.das 4 passed, test_model_image.das (arms mechanics, smol, untied) 66 passed 24 skipped.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…write_plane` and `write_metal_blob_plane` differed only in where the bytes come from (a field appended whole, the metal blob derived band by band); `write_section(w, sections, name, nbytes, rate_from, how, fill)` owns the page pad, the timing, the failed-write line, the rate line and the section push, and each writer passes its fill. The bytes are unchanged; two log lines move: a plane whose write failed no longer logs a rate line before its failure line, and the rate line reads "MB in" / "MB derived+written in" as before. Ledger row 124 (E4 T6). Proof: REVIEW.das clean; the seven-lane mint of Qwen2.5-0.5B-Instruct-Q8_0 sha256 identical (the plane writer); test_model_image.das (arms mechanics, smol, untied) 66 passed 24 skipped. The derived metal blob (`metal_derive_plane`, the load path's metal flavor) runs on an Apple box only - needs Apple for its own proof.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…. `image_peek` and `parse_image`'s identity decline each copied the meta blob's leading 64 KiB out of the mapping, wrapped it in a reading archive, read the strings it leads with and checked the stream stayed inside it; `read_meta_head(bp, meta_off, meta_bytes) $(harch) { ... }` does it, each caller reading the strings it wants. `parse_image` sits in the layout closure, so modules/dasLLAMA/REVIEW.das re-stamps IMAGE_LAYOUT_STAMP_HASH (0x62c76ef2581a3a96 -> 0xefb41da77dd2d6a1); no byte moves and IMAGE_VERSION stays. Ledger row 131 (E4 L1). Proof: REVIEW.das clean; under DASLLAMA_GPU=1 test_vulkan_mint.das 6 passed, test_model_image_vulkan.das 4 passed, test_model_image.das (arms mechanics, smol, untied - the peek and the identity decline) 66 passed 24 skipped.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…section`. `borrow_plane` and `bind_plane_section` each looked a section up by name and declined one that was not page-aligned inside the mapping (the borrow also in whole elements); `image_section(by_name, msize, path, name, esz, detail, ok)` does it, the borrow passing its element size and its "(esz N)" detail, the bind one-byte elements and none, so both decline lines read as before. Ledger row 183 (E4 L2). Proof: REVIEW.das clean; under DASLLAMA_GPU=1 test_vulkan_mint.das 6 passed, test_model_image_vulkan.das 4 passed, test_model_image.das (arms mechanics, smol, untied) 66 passed 24 skipped.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…. The staged cache rail and the plain cache rail each logged the write time, released the carrier the image was built from, and mapped the file back or panicked; `map_written_image(src, out, img, tag, quant, ts)` does it, the line and the panic text as before. Ledger row 176 (E4 L3). Proof: test_model_image.das (arms mechanics, smol, untied, whisper, kitten, canary - the staged carriers' mints) 70 passed 20 skipped.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… `fill`. `reset_gpu_cpu_passes` zeroed its count array with a for loop; it calls `fill(g_gpu_cpu_pass_counts, 0l)`, and the tier requires daslib/algorithm. Ledger row 67 (X46: TL1 D4, the tier member; row 67's common and decode members landed with C1a, the test members are the test clusters'). Proof, the tier armed: test_vulkan_tier.das 36 passed, test_gpu_tier.das 25 passed, test_gpu_serving_declines.das 56 passed, no skips.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… bases are named once. The tier's head-block base and plane length and the resident upload's final-norm copy and q/k rows each spelled the plane's layout - past every layer's `RDEC_NORM_ROWS` rows the final norm, one row on the q/k rows - as arithmetic; `rdec_norms_final_base(n_layers, dim)` and `rdec_norms_qk_base(n_layers, dim)` in the tier answer it, and both files read them. The plane's layout is unchanged. Ledger row 215 (E3 L10). Proof, the tier armed: test_vulkan_tier.das 36 passed, test_gpu_tier.das 25 passed (the norms plane's head block), test_gpu_resident_qwen2.das 15 passed (q/k-free norms), no skips; the q/k-norm and head carriers run in the next resident batch.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… go through `rdec_dn_row_own`. The verify and the one-row token each reset the deltanet state at position 0, advanced its position and handed every recurrent layer's state to the driver, the spelled-out `rdec_dn_row_own` (the verify by `n` rows); the prefill handed the states over through the window chain's seat with `rdec_dn_own_all`'s loop written again. `rdec_dn_row_own` takes the row count, `rdec_dn_own_all` takes the prefill seat as a flag, and the three sites call them - the one-row token keeping its rewind pass ahead of the call. Ledger row 130 (E3 L5). Proof, the tier armed: test_vulkan_tier.das 36 passed; test_gpu_resident_hybrid.das, test_mtp.das and test_gpu_resident_regions_hybrid.das (the recurrent carriers) run in the next resident batch.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…d once. The plan pushed a mixer site's q8 down and up planes four times (the layer's two sites, the head mixer, the NextN head's own mixer) and the hc upload placed a down/up pair three times (every site, the head mixer, the NextN head's mixer), each with its own failed-placement test; `push_hc_mixer_planes(c, planes)` and `resident_place_hc_mixer(t, down_off, up_off, mr, kg, wb)` do them, the decline texts unchanged. Ledger row 158 (E3 L8). Proof, the tier armed: test_vulkan_tier.das 36 passed; test_gpu_resident_hc.das (the hyper-connection carrier, its head mixers) runs in the next resident batch.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ent_fit_ctx` picks. BEHAVIOUR (the message only): the decline's "would fit at ctx N" divided the room past the weights, scratch and pool by one region's K/V row and the shadow's bytes a position, so with more than one mirror region, and always by the re-plan's one-part-in-eight slack, it named a context the re-plan would not pick; it now reads `resident_fit_ctx(p, seq_cap, shadow_pos)`, the arithmetic the re-plan shortens by (regions and slack included), and its own copy of that arithmetic goes. Before, one region: N = room / (row + shadow); after: N = (room x 7/8) / (regions x row + shadow) - the context the shortened plan arms. The plan's fit, its numbers and every other decline are unchanged. Ledger row 155 (E3 T15, finding F16). Proof, the tier armed: test_vulkan_tier.das 36 passed, test_gpu_tier.das 25 passed (test_resident_fit_ctx), test_gpu_serving_declines.das 56 passed, no skips.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ough `rdec_dn_own_all` again. Row 130 routed the verify's reset, advance and hand-over through `rdec_dn_row_own(s, c, pos, n)`, and modules/dasLLAMA/REVIEW.das's verify-seat check (every decline of the seat above the state's move to the device, ARCHITECTURE_GPU_VULKAN_MTP.md#resident-verify-rollback) finds that move by its `rdec_dn_own_all(` call, so it reported the seat as having none. The verify spells its three lines as before row 130 again, and `rdec_dn_row_own` loses the row count only that site passed; the one-row token and the prefill keep the fold. Ledger row 130 (E3 L5) fix. Proof: REVIEW.das clean; the verify's text is its text before row 130; test_mtp.das and test_gpu_resident_hybrid.das ran green at row 130 with the same semantics (batch g2).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ormat's plane pair. The trim deleted the kq table's pairs and spelled the per-32 and mx4 formats' pairs (q51, iq4nl32, mx4) field by field, the trim's dropped-section test named those pairs' sections by hand beside a kq loop, and the mint's zero view resized each format's pair through a four-way branch; `kq_plane_pair(t, f) $(q, s)` answers a format's quant and scale planes (the kq slot of a superblock format, the per-32 and mx4 formats' own pair) and `kq_plane_pair_names(f)` their section names, and the three sites walk every `KqFmt` through them. The trim's single planar families (qblob, qscales, qscales16, wblob, bf16blob, q4blob, q4scales) stay named in both its delete list and its name test - folding those needs a field walk over `Model`. dasllama_common's `kq_planes_of` (the kq slots alone) is not this cluster's file. Ledger row 292 (X48a: E3 T10, E4 T8; row 117g). Proof: REVIEW.das clean; the seven-lane mint of Qwen2.5-0.5B-Instruct-Q8_0 sha256 identical (both trimmed lanes); under DASLLAMA_GPU=1 test_vulkan_mint.das 6 passed, test_model_image_vulkan.das 4 passed (the trim cells).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… q/k/v form. `tq4_rotate_batch`, `tq4_unrotate_batch` and the raw-pointer `tq4_unrotate_rows` each walked every head segment in one direction; `tq4_basis_rows(p, nrows, stride, head_size, signs, inverse)` is the walk and the two array forms forward to it. `tq4_rotate_for_store` / `tq4_unrotate_from_store` and the Metal prefill continuation's gathered panel rows each rotated K under a tq4 K side and V under a tq4 V side; `tq4_sides(s, kp, vp, nrows, kv_dim, head_size, inverse)` does it, and `tq4_unrotate_rows` goes (its one caller, dasllama_metal_prefill.das's continuation, now calls `tq4_sides`). The q/k/v rotate block at the RoPE seam, copied in dasllama_batch.das's batch decode attention and dasllama_blocks.das's prefill attention, is `tq4_rotate_qkv`; the batch decode's STYLE037 nolint, which the fold left suppressing nothing, goes. Ledger row 50 (X12: E4 T3, row 138c). Proof: test_kv_codec.das 41 passed 3 skipped (stories15M.bin absent, unchanged), test_vulkan_kv_codec_kernels.das 12 passed, under DASLLAMA_GPU=1 test_gpu_perop_kv_codec.das 5 passed, _k 4 passed, _iq 4 passed (tq4 sessions through the per-op rails), test_batch_decode.das 17 passed 9 skipped (stories15M and the large tier absent). The Metal prefill site does not compile on this box - needs Apple (arm7b-tq4kv).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ugh `tq4_rotate_for_store`. The window's K/V readback, before the codec store, rotated the K rows under a tq4 K side and the V rows under a tq4 V side by hand, the body of `tq4_rotate_for_store(s, s.k_b, s.v_b, npos, kvd_l, hs_l)` (dasllama_common), which the CPU prefill already calls at the same seam; it calls it. Ledger row 159 (X12: E4 D1, row 138c). Proof: needs Apple - dasllama_metal_prefill.das does not compile on this box; the Metal parity arm arm7b-tq4kv owes it. The helper's own carriers are row 50's (test_kv_codec.das, test_gpu_perop_kv_codec*.das).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…uf_read`. The tower's private `vt_readback_copy` (a mapped host buffer's bytes from an offset into an array) is `host_buf_read` (dasllama_vulkan_common), which C1a added for the decode, prefill and seam landings; its six callers (the residual rows of every chain, the deepstack taps, the qwen3a and canary fronts, the stem's rows) call `host_buf_read` and the private copy goes, and the whisper decoder's logits landing, the same memcpy, calls it too. The bytes landed are unchanged. Ledger row 56 (X21: E6 D2 and asr_dec:770-772, the tower and ASR members; the decode, prefill and seams members landed with C1a). Proof, under DASLLAMA_GPU=1: test_whisper.das (DASLLAMA_WHISPER_DIR set, the test_whisper_vulkan_* cells) 14 passed, test_audio.das (test_gemma4a_vulkan_twin, test_canary_vulkan_twin, test_qwen3a_vulkan_front, test_encoder_blocks_vulkan) 10 passed, no skips.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…r's private `vt_cached` is the `cached(tab, key) { build }` C1a moved into dasllama_vulkan_common.das for the prefill; its six callers call `cached` and `vt_cached` goes, and the four sites that open-coded the key-exists-then-build (the split-k planes' fc2 set, the whisper-class stem's conv2 set, the qwen3a rel-projection set, the canary front's set) go through it too. The tables, keys and sets are unchanged. Ledger row 112 (X5a: E6 D1, the tower members; the prefill's six landed with C1a). Proof, under DASLLAMA_GPU=1: test_whisper.das test_whisper_vulkan_* 14 passed (DASLLAMA_WHISPER_DIR set) and again under DASLLAMA_CM2_SPLITK=4 (the fc2 split-k set, its "splits K" line logged) 3 passed, test_audio.das's Vulkan twins 10 passed, test_vulkan_tower_kernels.das 28 passed, no skips.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…plitk_reduce_enc`. The chain built the reduce's push (`SkRedArgs`) and its one-workgroup-per-1024-floats grid by hand, the encode C1a named `splitk_reduce_enc(raw, h, rset, nelem, k)` beside the reduce class for the prefill; the chain calls it. The push and the grid are unchanged. Ledger row 195 (X6a: E6 T4's encode, the tower member; the prefill's members landed with C1a). Proof, under DASLLAMA_GPU=1 with DASLLAMA_WHISPER_DIR set: test_whisper.das test_whisper_vulkan_twin 3 passed under DASLLAMA_CM2_SPLITK=4 (the reduce encoded) and 3 passed on the wave model's pick, no skips.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… admission chain (resident_decline_reason -> resident_trim_decline) reads the prefill pin, a foreign override, the route, the tier's seams and the card's room beside the model's shape, and the mint answered every decline alike - an untrimmed lane saved under the trimmed identity, which the load-time recheck (trimmed lanes alone) never asked again, so once the pin lifted the untrimmed lane kept serving, and a pinned run against a good trimmed lane deleted it and re-minted it untrimmed. `resident_run_mode_decline` is the pin/override check the chain shares; `resident_trim_decline_is_run_mode` tells this run's declines (route, pin, override, seams, plan fit) from the model's (experts, residual, features, layers, head). The mint under a run-mode decline writes NO lane and serves the streamed build planar from memory, as master did; a trimmed lane on disk a run-mode decline cannot serve stays on disk (planar from memory this run); a model-shape decline deletes and re-mints; frozen artifacts are never removed, the panic names the kept lane. The .dlim arm's panic tells a knob from a shape. BEHAVIOUR: under DASLLAMA_PIN_PREFILL / a foreign override / the route off / a plan past the card, DASLLAMA_TRIM mints nothing where it minted an untrimmed lane. test_vulkan_mint_trim_follows_the_driver holds the lane's size unchanged under the override, no lane on a cold trim ask, the override put back from its getter and the next load minting the trimmed lane; test_gpu_serving_declines gains the run-mode / model-shape classification over the chain's texts and a selected foreign override, and the codec ask's put-back round trip (`restore_gpu_kv_dtype` on a pinned and an unpinned ask). Review round E3 / REVIEW_IMAGE V1.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…target_logp` moved onto `hmax` / `hlse` in the pass, whose folds take `max`, which drops a NaN operand - a device row with a NaN anywhere but the target id (and the scalar tail) passed every `ppl_compare` cell where the hand loop gave the row a -90 log-prob; the row is scanned for a NaN first and returns it. Review round T1 (tests/REVIEW.md: a moved bar computation reads a NaN as outside the bar); probed interpreted on a 37-element row - old -94.9, new -7.97 before the fix, NaN after it.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…oot_restore] refuses a global below it. `check_cm2_ladder_set` read the `khr_cls_*` and `cm2e_cls_*` trios, which the engine stopped calling when `cm2_cls_*` grew its own `kq_tile_stamp` over `CM2_TILE_NONE` - a format added to that `let` compiled, passed the gate and panicked at its first dispatch of the tile. The gate now walks `cm2_cls_*` (stamped over `tile_<verb>()` with the let) and reads the let's `<fmt>/<tail>` pairs against the cm2 templates per tile, keeps the kernel cells' `khr_cls_*` ladder under the stamp rule, and `let_words` is the one roster reader (the kqformat alias roster reads through it); the dead `cm2e_cls_*` trio goes; the walker fixture plants the live shape (red: the khr ladder hand-written, the let skipping k5's e tile). `[boot_restore]`'s `apply` sees the globals above it alone (the parser has not reached the rest), so its `lint` pass walks the whole module and refuses a `= @@fn` global its guards missed; the gate matches the annotation line itself (comments stripped) and reds a function global below it where it skipped the whole file on a substring. test_dasllama_lint_contracts spawns the two fixtures (a global below refused by name, every global above compiles); test_walkers 9/9; GEMM.md and VULKAN.md say which ladder the engine and the cells each run. Review round E1 / E5 / REVIEW_GATES F1.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…on the paths a cell holds bit for bit (the gate-up GEMV's accg/accu, tts_gelu_tanh's cubic, DaAttnT's scalar-mirror accumulations); the plain norm stamps norm by construction (`ln_on` reads the push block under ADD alone) and the one-column gate-up GEMV has one column by NC; AtAttnTileT.run, AtAttn.run and DaAttnBWT.run spell TQ / KS / LD where they meant the tile; the Vulkan-only decode helpers (grid_sv4, grid_sv1, k4_scale_min, q40_d_of, k6_sc, half_sat) and the flash head gates (fa_head_ok, fa_cm2_serves, fa_khr_serves) live in the classes file, gpu_math keeps what both homes splice (and drops shader_lingua_franca); the probe's `k4_dmin` was `k4_scale_min` and calls it; `dev_upload` is one public upload in common (the decode's row uploads and the cell harness read it); the workgroup-cap decline remembers the classes it declined (the TTS stages re-ask every chunk and built two strings each time) and keys its counter on `VkdBuildDecline`; the Vulkan lens refuses a `*` kernel prefix naming two methods; the tq4 store helpers check the codec before taking an address; the K/V seat asserts name the prepare and the trunk apart; the four PL-BERT / predictor / decoder ensures carry `[hot_path]` (the perf lint cannot follow the `@@` the stages now pass). Proof: test_vulkan_kernels 204 cells / 203 passed / 1 skipped before and after; vk_spv_diff over the two dumps 184 identical, 57 moved, 1 appeared (da_attn_bw_f16, now built by a cell): 20 reorder their helper functions only (sorted id-normalized disassembly identical), 13 inline the moved helpers (OpFunctionParameter -> OpVariable), 12 fuse multiply-adds (FMul+FAdd -> fma), 6 da_attn stamps fuse the scalar-mirror accumulations, 4 at_attn / da_attn_b tiles carry one OpISub from the `TQ - 1u` spelling, since re-parenthesized. Review round E6 / E7 / E10 / E11 / E14 / E16 / E17 / E23 / E25 / E26 / E29, dupe T11.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…wide-head attention stamp (gemma-4's global heads on the f16 mirror) runs its cell beside every wide arm and the gated arm - it served with no oracle cell while the census said it had one; `gemv_ar_planes` / `qkn_cell_enc` take `cp` untyped (a typed CellPlanes behind the `?vulkan` require broke the model-free file's compile on a build without the module, the Apple default); the row bars gain the controls their loosening owed (a device row at 0.8 of the one-off distance holds, 0.9 reds at TQ4_CTRL_SHARE; a NaN reference with no control row reds) and `rows_bar_fails` is `row_within_ctrl`; `half_pair_word` is `packHalf2x16`; the MTP token compares log both decoded streams and the reject cadence goes back to what its getter read; the per-32 block literals the folds carried read `KQ_B32_ELEMS` / `KQ_SB_ELEMS`; the kernel file's header and tests/CLAUDE.md name the cells' skip facts; `rows_within_control`'s doc reads the live max; the probe's KHR CPU check takes its floor over the finite outputs. Docs: followup_vulkan rows 4 / 32 / 78 and two narrative lines name the current seats, stamp and gate (`rdec_kv_sync`, `topk_n_cls`, `fa_cm2_serves`, `cls_ar_rq_b`, the clamp-convert); lcpp_bench's two "test_parity" comments point at `tests/_model_tier.das`; HOW_TO_ADD_A_FORMAT's ladder steps say the ladders are stamps; ENGINE_FORMATS carries `stamp_ladder` and the KV block byte constants where they live; GPU.md names `sig_gate`, the seats' setter and the plane-pair accessors. Review round T3 / T4 / T6 / T7 / T8 / T9 / T10 / T11 / E2 / E17 / E21 / E28, TDD F11.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…l log reads a tokenizer-less model as ids. The streaming closure answered a run-mode trim decline with `return false`, which handed the load to the eager rail - it cached a planar image, and the RAM bake (`vulkan_bake_flavor`) then wrote an untrimmed lane under the trimmed identity, the very lane the decline exists to refuse; the closure now raises `mint_held` and skips the save while the streamed build serves from memory, and the RAM arm runs the same run-mode check and writes no lane under it (it never trims). test_vulkan_mint_trim_follows_the_driver gains the control that the held mint cached no planar image. `eyeball_text` on a `load_gguf` model (no tokenizer, `TokKind.unset`) indexed past the SPM vocabulary; it logs the ids themselves. Found by the fix batch's -jit run (test_vulkan_mint 6/8, test_mtp 12/18 before, 8/8 and 18/18 after).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…entry describes the lane the run writes and the one it keeps; the codec attention arms read as one list; the da_attn_rqk entry claims the cells it has; the gated-cell list carries the f16 twin; act_gelu, draft_suppressed ("suppression floor"), the model-free cell list, the llama Q4_0 file and the tq4 share property are named as the code names them; rewraps. tts_tanh's doc says where the device tanh fails (about +-44) apart from where the clamp sits (+-15).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ells. `areas_for_path` sent every `_vk*` fixture to `llm`, while `_vkd_oracles.das`'s CellPlanes is the harness of the four `test_vulkan_tts_*` files and of `test_vulkan_tower_kernels.das` too - an edit to it under `--changed` ran none of the TTS or tower Vulkan cells, where tests/CLAUDE.md said the fixture reaches the areas it serves. The fixture maps to `llm, tts, audio, vision`; test_run_suites pins the row. The preflight lane's doc says the stocked lane is the whole model-gated suite and that a PR's run on a box with models is `--changed`, which has no lane. Dragon pass 2 over tests/CLAUDE.md, S3 / S8.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…plit three-kernel path of the k4 gate-up cell reused the fused run's gate plane without the sentinel back, so a skipped gate GEMV read as a byte-for-byte match - the plane is refilled before the record; the gated h128 attention arm barred a stamp that stages q/K/V through halves at a bare 2e-2 - the arm takes its bar off the oracle rows (`cm_bar` for the h128 stamp, the fixed bars for the f32 stamps); the fused gate-up GEMV's third activation branch (gpt-oss's clamped swiglu, `ACT_OAI`) had no cell on either the one-column or the N-column stamp - both runners take the act code and run it; the count-only compares log what they measured: `check_qbytes` the first word off, `gu_blocks_hold` the largest distance in quant steps at its element (`Q8BlocksOff.worst` / `worst_at`), `mirror_rows_check` the largest difference at its mirror element (`MirrorOff`), the combine-folded residual rows through `check` with their row shape, the fused-vs-split quant words and scales through `check_exact`. test_vulkan_kernels 206 cells / 205 passed / 1 skipped before and after. Review round REVIEW_KERNEL_CELLS logging / every-path / sentinel / narrower-precision bar.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…e (three paragraphs re-wrapped; the opening's anchor-cited routing had grown it to 302). LINT027 from the make-pr chain's lint lane.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ree. ARCHITECTURE_GPU_VULKAN_HC.md named `RdecPass.hc_prefill`, which the folds removed - the served prompt's witness is the served-prefill counter; HOW_TO_ADD_A_FORMAT.md's two ladder sentences named a `cm2e_cls_*` ladder where the e tiles are tails of the one `cm2_cls_*` ladder and `khr_cls_*` is the kernel cells'; tests/CLAUDE.md's family-filter line listed two files whose `qwen35` is a label, not a `family_on` tag.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…0 keeps the one fold of the deleted row 117 that landed by half (the trim's seven base plane names spelled twice); row 141 keeps the deleted row 138's Metal half (the prefill's tq4 un-rotate site beside `tq4_unrotate_from_store`, owed an Apple run); row 93 names the four probe-first candidates and the Q8 byte-store pair still open; the row the pass had written under the recycled number 90 is row 142 at the tail; the second row numbered 93 (MoltenVK) is 143; the two dated ar-fusion records read `f16cvt` again - history keeps the names of its day.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…hem. The bench rows against master on the RTX PRO 4500 (one process a row, master then the tip, the same lane) held tg128 on every model and read pp512 3.7% under on Qwen3-4B Q8_0 and 2.8% under on gemma-4 E4B Q8_0 with tight spreads, twice; the per-stage GPU profile put every Q8 weight GEMM stage 3-5% longer and the rest flat. An A/B on the pod named two folds: `MmBatchT`'s `l_tile` / `m_tile` as one coopmat accumulator array walked under `[unroll_full]` (2.5% - its own probe at 4096x4096 had read flat per dispatch; the served shape did not agree, and the old and new tiles load the same fragments, so the array itself is what the driver stops keeping in registers), and `Q8Cm2T.decode`'s shift form from the F8 rule change (1.5% against the 16-bit lane's `unpack8` byte select). With both back the row reads 11894 +- 331 against master's 11688-11772. The two mm tiles stay two bodies (followup_vulkan.md row 93 names the fork with the readings); the Q8 decode keeps the select and ARCHITECTURE_GPU_VULKAN_GEMM.md#cm2-decode-16bit-lanes carries it as the measured exception to the shift form. test_vulkan_kernels 206 / 205 / 1 skipped; vk_spv_diff against master: q8_batch_mm, _mm_a, _mm_m and q8_batch_cm2l/m/s byte-identical again.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ench row master against the tip on the RTX PRO 4500 (the 1B, Qwen3-4B Q8_0, gemma-4 E4B Q8_0, gemma-4 12B, Qwen3.5-9B MTP, Qwen3.8-27B under both K/V codecs); the two Q8_0 prefill folds the served row took back, with the stage profile and the A/B that named them; the tq4 control share on this card; the llama Q4_0 carrier under the file's logit bar; `rows_within_control` as the rigs' one bar statistic.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… kq superblock formats. The rule's readings came from the expert-schedule shape on the superblock formats; the Q8_0 tile's 16-bit-lane select read 1.5% ahead of the shift form on the served dense prefill and is the measured exception, carried by the architecture section the rule cites.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The reason will be displayed to describe this comment to others. Learn more.
Copilot review overview
🔵 Needs a closer look
The change spans 131 files and 304 commits with deliberate bit-level SPIR-V/FP changes, measured performance forks, and an intentional observable DASLLAMA_TRIM behavior change that require human judgment and hardware validation beyond automated review.
Review effort: Balanced Findings: None
What changed in this PR
This PR closes the deduplication debt accrued across the dasLLAMA Vulkan arc (#4152/#4153/#4163/#4166/#4187/#4200). It folds duplicated shapes — kernel cells copied per format, resident-driver helpers copied per family, decode/prefill recorders copied per rail — into single shared implementations, and clears the three dedup rows (117/133/138) plus the full sweep from followup_vulkan.md. It is a very large mechanical refactor (reported as 304 commits, 131 files, net −5299 lines) with a handful of deliberate, measured exceptions and a small set of in-PR defect fixes.
Changes:
Consolidates scalar math (sigmoid_f32/silu_f32/sig_gate/softplus/dn_decay) into dasllama_gpu_math.das, tq4 rotation into shared tq4_* helpers in dasllama_common.das, and find_kernel_fn into dasllama_kernel_access.das shared by the Metal and SPIR-V lenses; adds a new [boot_restore] function macro to restore = @@fn function globals after deserialization.
Deletes five redundant kernel classes (cls_ar_rq, f16cvt, q8_gemv_gu, topk, tower_rms) in favor of generalized families, keeping two measured forks (the mm prefill tiles and the Q8 cm2 decode unpack8 select) as documented perf exceptions.
One intentional behavior change: under DASLLAMA_PIN_PREFILL, a foreign override, the route off, or an over-card plan, DASLLAMA_TRIM no longer mints an untrimmed lane; plus doc/ledger/test-harness updates.
File
Description
dasllama/dasllama_gpu_math.das
New home for shared scalar math; softplus/dn_decay moved in, sig_gate added, tts_gelu_tanh uses mad.
dasllama/dasllama_metal_kernels.das
Removed local dn_decay; still calls it/softplus via the gpu_math require.
dasllama/dasllama_arch_qwen35.das
Added require dasllama_gpu_math for the moved math helpers.
…ulkan again. The darwin core lane's changed-set lint failed the file: `tok_buf`, a helper the CellPlanes fold orphaned, typed its parameter `array<TokMeta>` at top level (the type lives behind the `?vulkan` require) - deleted; `ar_comb_check` typed its half-row words and so resolved `half_at` (a `_vkd_oracles` function) outside any vulkan arm - the parameter is untyped, as every other top-level helper the vulkan cells share spells its device-side arguments, so the call binds only where a guarded cell instantiates it. test_vulkan_kernels 206 / 205 / 1 skipped under -jit.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The reason will be displayed to describe this comment to others. Learn more.
Copilot review overview
🔵 Needs a closer look
The change spans 304 commits across ~131 files with deep GPU/kernel semantics and behavior-bearing folds that cannot be fully verified from the visible diffs or exercised without the Vulkan hardware test suite, so it needs final human review.
…On a build without dasVulkan the `static_if` arms that use `_vulkan_lane` (test_model_image_vulkan), `math` (test_vulkan_kv_codec_kernels) and `dasllama_gpu_math` (test_vulkan_tower_kernels) are dropped before STYLE030 walks the program, so the rule's verdict follows the build - the standing `nolint:STYLE030,LINT019` marking until the compiler records the dropped arm's names.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The reason will be displayed to describe this comment to others. Learn more.
Copilot review overview
🔵 Needs a closer look
The change spans 131 files and 304 commits whose correctness rests on device-level SPIR-V/parity and bench validation that cannot be reproduced in this environment, so it needs final human review.
Review effort: Balanced Findings: None
Previously missed (1)
In code that hasn't changed since last review
Missing blank line between adjacent top-level functions
modules/dasLLAMA/dasllama/dasllama.das:744
Minor/optional: the blank line between reset_gpu_kv_dtype and gpu_ctx_pins was dropped, so these two top-level functions now sit directly adjacent. Every other adjacent top-level def in this file is separated by exactly one blank line (e.g. 737→740, 747→750). Restoring the blank line keeps the file's spacing consistent.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
The Vulkan arc (#4152 / #4153 / #4163 / #4166 / #4187 / #4200) left three dedup rows in
followup_vulkan.mdand a tree where the same shape lived in two to fifteen places: kernel cells copied per format, resident-driver helpers copied per family, decode and prefill recorders copied per rail. Every later format or family paid the copy count. This pass closes the arc's dedup debt in one PR, as ruled: all tier A rows, all tier B rows but the prefill GEMM set caches (row 139), the probe-first kernel rows last and gated on a probe row within noise, the sweep's defects fixed in-PR each with a cell.What
would_servegap, a NaN passing a bar, the gemm probe's k6 / iq scales and deltanet layout, seven TTS seats logging nothing, ... (F1-F21).target_logpreading a NaN as NaN, the cm2 ladder gate reading the engine's ladder,[boot_restore]refusing a function global below it, the f16 wide-head attention cell,madon the bit-for-bit paths, the lens refusing a*prefix naming two methods,[hot_path]on the TTS ensures, the workgroup-cap decline memo, the_vkd_oraclesarea map, the kernel-cells checklist's four violations.mmprefill tiles stay two bodies (one coopmat accumulator array cost 2.5% of Qwen3-4B Q8_0 pp512 where its 4096x4096 probe had read flat) and the Q8 cm2 decode keeps its 16-bit-laneunpack8select (the shift form cost 1.5%); both are measured forks in row 93 and the architecture doc.Observable
DASLLAMA_PIN_PREFILL, a foreign override, the route off, or a plan past the card,DASLLAMA_TRIMmints nothing where it minted an untrimmed lane under the trimmed identity.Where
modules/dasLLAMA/dasllama/*(decode, prefill, resident, tier, common, codec, image, kernel classes, tower / TTS / ASR drivers),modules/dasLLAMA/tests/*,modules/dasLLAMA/harness/*, the module's ARCHITECTURE / REVIEW / HOW_TO / ledger documents,tests/CLAUDE.md,utils/internal/review-md/test_walkers.das,utils/internal/preflight/main.das(one doc string).Validation
harness/vk_spv_diff.dasoverDASLLAMA_VK_SPV_DUMPdumps of the model-free suite), master f3a03fa (352 stamps) against the tip (349): 145 identical, 202 moved, 2 appeared (da_attn_bw_f16- a cell builds it now;q8_gemv_gu_n1- the N-column family replacesq8_gemv_gu), 5 vanished (the five classes the folds removed:cls_ar_rq,f16cvt,q8_gemv_gu,topk,tower_rms). The 202 moved by opcode delta: 84 unrolled copies folded into helper functions (OpFunctionParameter up, call plumbing only); 12x * sigmoid(g)->x / (1 + exp(-g)); 3 the gated q panel on the h128 / wide decode attention (a sweep defect, F1: a feature); 2 fused multiply-adds; 101 kernel-helper moves (29 cm2 tiles lose a dynamic vector extract, 18 khr tiles lose sixpackHalf2x16- five also the sign fold's twelve negations, 9 clamp + fma, 7 zero-init asOpConstantNull, 7 opcode-multiset identical, 7 norm stamps norm by construction, the rest listed in the classification file). The rename commits alone: 241 identical, 0 moved.test_vulkan_kernels206 / 205 passed / 1 skipped (RTX 5060 Ti), bit for bit wheretests/REVIEW_KERNEL_CELLS.mdrules it; the model-free suite 85 files / 0 failed / 5 cells skipped at the tip.tests/run.das,DASLLAMA_GPU=1 DASLLAMA_PARITY_FULL=1) at ecf18ce: model-free 85 files (the one red was the facade stub, fixed in 09273ed); stocked 87 files, 1 red =test_vision_chat.daskilled at exit 137 by the container's 62 GB cgroup, on master's binary the same (followup_general.mdrow 181: the process reaches 57 GB resident on both trees); coverage (--arm coverage-vk) 1 file, 6 cells / 4 passed / 2 skipped; audio+vision 20 files, the same vision_chat red, 82 cells skipped; tts 17 files / 0 failed / 47 skipped; model_image 90 / 77 / 13 skipped and model_image_tts 90 / 68 / 22 skipped; the per-file cells withDASLLAMA_PARITY_FULL=1(cells / passed / skipped, 0 failed each): gpu_resident_hybrid 69 / 65 / 4, gpu_resident_moe 17 / 6 / 11, gpu_resident_hc 6 / 3 / 3, mtp 18 / 15 / 3, gpu_tier 25 / 25, gpu_serving_declines 56 / 56, gpu_slot_swap 2 / 1 / 1, gpu_model_swap 6 / 5 / 1, vulkan_mint 8 / 7 / 1, model_image_vulkan 4 / 2 / 2, regions_tq4kv 26 / 26, regions_q8kv 26 / 26, regions_qwen3 12 / 12, regions_hybrid_k 4 / 4, gpu_resident_llama 14 / 12 / 2.lcpp_bench.das --ngl 0 -p 512 -n 128 -r 5 -t 16, one process a row, master 1a0729a then the tip on the same card and lane), tg128 master / tip: Llama 1B Q4_K_M 554.87 / 551.73, Qwen3-4B Q8_0 146.54 ± 12.62 / 152.21 ± 0.22, gemma-4 E4B Q8_0 98.33 ± 15.27 / 112.00 ± 0.05, gemma-4 12B Q4_K_M 82.12 / 82.18, Qwen3.5-9B MTP 105.20 / 105.07, Qwen3.8-27B tq4 K/V 36.97 / 36.96, Qwen3.8-27B q8_0 K/V 36.98 / 37.01 (the ±12–15 spreads land on either tree at random: a stall, not a tree). pp512 read 3.7% and 2.8% under on the two Q8_0 models with tight spreads, held by a 10-rep recheck; the per-stage profile and an A/B named two folds (the mm tile's accumulator array, the Q8 cm2 decode's shift form), both taken back in d922920. At the tip, 10 reps, master-tip-master-tip: Qwen3-4B Q8_0 pp512 11765 / 11753 / 11508 / 11754, gemma-4 E4B Q8_0 9610 / 9584 / 9575 / 9587. The entry is in PERF_LEDGER.md.utils/internal/make-pr/main.das --, the full preflight once): sync, stamp-reach, jit-smoke, untracked, format, hash-refs, review-md, review-md-tests, md-ascii, ast-verify (112 files), ci-das (27 files), compile-sweep (768 roots) green; lint red once on a 302-line architecture doc, re-wrapped (36af567) and the lint lane re-run alone: 130 files clean on both rails. The six lanes the red fast tier skipped, each run once: docs (sphinx-html) green, tests-cpp green, tests-aot green; tests-interp 14836 / 1 failed and tests-jit 14752 / 2 failed =tests/watchdog/test_watchdog.das(80 s) andtests/jit_tests/jit_lib.das(90 s) past dastest's 60 s per-file cap under 46 workers - alone they run 76 s and 59 s and pass 52 / 2 skipped and 7 / 7; this box's standing per-file-budget reds on files the branch does not touch. utils-tests: all 14 suites green (0 failed), the lane red is msbuild reading the daspkg and mcp-setup suites' expectederror:refusal lines (Windows only; CI's lane is Linux make).Claims
Not done
vkd_cached_set(ruled out of this pass).*_termsmethods, E18 spirv fixtures, E24alloc_cmd, hot-path F5 / F1, the lane-pins hook uninstall, T10's remainder (vk_gemv_probe 174/179, _kq_fixtures 421/623), T12 c7_probe figures, E4 converter narrowing,DECVEC=1rows,gemma4a_test2.wav..mdstop kept from the deleted dedup rows: 140 (the trim's seven base plane names spelled twice), 141 (the Metal prefill's tq4 un-rotate site, owed an Apple run with the M5 Metal cells), and row 93's four probe-first candidates plus the Q8 byte-store pair.test_vision_chat.dason the pod: killed by the container's 62 GB memory cap on master and on this branch alike (followup_general.mdrow 181).tests/ARCHITECTURE.mdcompanions, and its history passages go): the owed rule-doc PR.🤖 Generated with Claude Code