Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
47 commits
Select commit Hold shift + click to select a range
7032f51
dasllama-server: an ASR model whose decoder is an LLM decodes on the …
borisbat Oct 4, 2026
90ce366
dasLLAMA: the vision towers' refusal of a projector past the staged-m…
borisbat Oct 4, 2026
f44bb5a
dasLLAMA: the Metal tower reads a whisper encoder's q8 planes - the b…
borisbat Oct 4, 2026
f69887a
dasLLAMA: the Metal whisper decoder leaves a window's cross-KV layout…
borisbat Oct 4, 2026
cfb5713
dasLLAMA: gemma-4 E4B audio loads - its audio embedder widens the enc…
borisbat Oct 4, 2026
a441d16
dasLLAMA: the Metal tower fills the whisper blocks-with-post-norm sea…
borisbat Oct 4, 2026
20c1a39
dasLLAMA: the Metal whisper decode step serves from a text width of 5…
borisbat Oct 4, 2026
66f7da4
dasLLAMA: a whisper.cpp q8_0 bin loads - the loader reads Q8_0 tensor…
borisbat Oct 4, 2026
172a3a8
dasLLAMA: the Metal conv stem serves the 80-mel whisper models - the …
borisbat Oct 4, 2026
4c7bb5c
dasLLAMA: a parakeet q8_0 bin loads - the whisper and parakeet bin re…
borisbat Oct 4, 2026
5cd84a0
dasLLAMA: the canary encode clocks its stages - the mel, the front, t…
borisbat Oct 4, 2026
42c8e6f
dasLLAMA: canary's mel and subsample front run on the Metal tower - t…
borisbat Oct 4, 2026
bc1b378
dasLLAMA: the audio embedder carries the whisper-class tower - a Qwen…
borisbat Oct 4, 2026
10e736e
gc: the mark walk sizes a range in 64 bits - a borrowed array whose b…
borisbat Oct 4, 2026
8356edc
dasLLAMA: the chat towers' projector tail and mel run on the Metal to…
borisbat Oct 4, 2026
351447e
dasllama-server: --audio-mmproj arms a slot's audio from any audio pr…
borisbat Oct 4, 2026
fe88e3f
dasLLAMA: a media span is part of the prompt the prefix cache matches…
borisbat Oct 4, 2026
64f422d
dasllama-server: a chat pays for an image or a clip once - media stay…
borisbat Oct 4, 2026
5aaa035
dasllama-server: a known media part is encoded only where the prefix …
borisbat Oct 4, 2026
85863de
dasLLAMA: the documents state the media cache - a prompt is one posit…
borisbat Oct 4, 2026
02b9fa5
dasLLAMA: gemma-3's vision projector tail runs on the Metal tower - t…
borisbat Oct 4, 2026
7c6c4e3
dasLLAMA: gemma-4's vision stem and projector tail run on the Metal t…
borisbat Oct 4, 2026
4f953ab
dasLLAMA: a gemma-4 E-series turn with thinking off opens bare, as it…
borisbat Oct 4, 2026
e5f2149
dasLLAMA: a Llama-3.1+ system turn opens with the header its template…
borisbat Oct 4, 2026
2fae352
dasLLAMA: a Q8_0 audio projector file bakes the halfword twin - its b…
borisbat Oct 4, 2026
65f5b4a
served_bench: the request whose reply a media row prints asks for thi…
borisbat Oct 4, 2026
f936439
dasLLAMA: the paged KV pool reserves its address room at scheduler cr…
borisbat Oct 4, 2026
ca5d265
dasLLAMA: a served media turn prefills its span and the text around i…
borisbat Oct 4, 2026
c660d9a
served_bench: a media row's first rep opens like the others - an unti…
borisbat Oct 4, 2026
0b083db
dasLLAMA: a gemma-4 E-series media body carries each row's token id -…
borisbat Oct 4, 2026
bb6820c
dasHV: get_body_bytes sizes a body by the body - a chunked answer car…
borisbat Oct 4, 2026
fb19811
Merge remote-tracking branch 'origin/master' into bbatkin/served-asr
borisbat Oct 5, 2026
b531d43
dasllama-server: the media queue's limit counts every new part a requ…
borisbat Oct 5, 2026
231b5dc
dasLLAMA: the new Metal tails zero the row pad of their pooled buffer…
borisbat Oct 5, 2026
68cc791
dasLLAMA: the gemma-4 vision ends' GEMMs ride the tower's one wrapper…
borisbat Oct 5, 2026
507e6a4
dasllama-server tests: the lost-span retry, a known clip's re-encode,…
borisbat Oct 5, 2026
ae98c04
dasLLAMA tutorials: the inline media turn built and evaled as one bod…
borisbat Oct 5, 2026
c0dd7e6
dasLLAMA: the whisper-class tower is a carrier family of its own file…
borisbat Oct 5, 2026
97e35af
dasLLAMA: the four ASR turns build and eval their head, audio and tai…
borisbat Oct 5, 2026
fb05289
dasLLAMA: the device tails' landing buffers and the scheduler's body …
borisbat Oct 5, 2026
957fce0
docs: the media body pair, the prefix probe, the span turn, the pinne…
borisbat Oct 5, 2026
1d3205c
tests/gc: the 4 GiB collect test stays interpreted - it collects with…
borisbat Oct 5, 2026
debf69a
dasLLAMA ledger: the served media turn's first token and the speech r…
borisbat Oct 5, 2026
f66f23c
dasLLAMA ledger: the arc's unfolded sibling sets, the unraced tail ti…
borisbat Oct 5, 2026
778a7b1
dasllama-server tests: the rig reads the client's file reader without…
borisbat Oct 5, 2026
3a8bafe
dasllama-server tests: the flags test drops the JSON require its remo…
borisbat Oct 5, 2026
ba4ea1c
dasllama-server tests: the prefix cache cell holds two cached runs to…
borisbat Oct 5, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 5 additions & 5 deletions doc/reflections/das2rst.das
Original file line number Diff line number Diff line change
Expand Up @@ -325,7 +325,7 @@ def document_module_openai(_root : string) {
group_by_regex("Audio (TTS and STT)", mod, %regex~(speech|speak|transcribe|translate)$%%),
group_by_regex("Moderations", mod, %regex~(moderations)$%%),
group_by_regex("Image generation", mod, %regex~(generate_image)$%%),
group_by_regex("Vision", mod, %regex~(chat_vision|vision_request_body|image_file_data_uri)$%%)
group_by_regex("Vision", mod, %regex~(chat_vision|chat_audio|vision_request_body|audio_request_body|image_file_data_uri|file_base64)$%%)
)
documents("OpenAI-compatible API client (chat, embeddings, audio, vision, ...)", mod, "openai.rst", groups)
}
Expand All @@ -343,13 +343,13 @@ def document_module_dasllama(_root : string) {
var mod = find_module("dasllama")
var groups <- array<DocGroup>(
group_by_regex("Model loading and sessions", mod, %regex~(load_model|create_session|create_kv_pool|release_kv_pages|create_batch_workspace|setup_dasllama_jobque|with_dasllama_jobque|caps|mtp_drafter_sidecar|attach_mtp_drafter|mtp_capable)$%%),
group_by_regex("Prefix cache", mod, %regex~(create_prefix_cache|prefix_attach|prefix_insert|prefix_checkpoint_at|prefix_release|prefix_held_groups|prefix_chain_list)$%%),
group_by_regex("Prefix cache", mod, %regex~(create_prefix_cache|prefix_attach|prefix_insert|prefix_checkpoint_at|prefix_match_len|prefix_release|prefix_held_groups|prefix_chain_list)$%%),
group_by_regex("Tokenizer", mod, %regex~(encode|decode|piece)$%%),
group_by_regex("Evaluation and sampling", mod, %regex~(eval|eval_embd|eval_embd_span|eval_embd_span_mrope|eval_batch|sample|set_seed|stats)$%%),
group_by_regex("Evaluation and sampling", mod, %regex~(eval|eval_embd|eval_embd_span|eval_embd_span_mrope|eval_embd_body|media_body_rows|eval_batch|sample|set_seed|stats)$%%),
group_by_regex("Generation", mod, %regex~(generate|generate_embd)$%%),
group_by_regex("Embeddings", mod, %regex~(embed)$%%),
group_by_regex("Vision and audio encoders", mod, %regex~(encode_image|encode_audio)$%%),
group_by_regex("Chat", mod, %regex~(create_chat|create_chat_renderer|add_user|add_user_audio|add_user_image|add_user_image_rows|add_user_audio_rows|add_assistant|render_turn|render_turn_marked|render_turn_image|render_turn_audio|render_assistant|render_close|respond|set_thinking)$%%),
group_by_regex("Chat", mod, %regex~(create_chat|create_chat_renderer|add_user|add_user_audio|add_user_image|add_user_image_rows|add_user_audio_rows|add_user_span|set_chat_date|add_assistant|render_turn|render_turn_marked|render_turn_image|render_turn_audio|render_assistant|render_close|respond|set_thinking)$%%),
group_by_regex("Tool calling", mod, %regex~(set_tools|add_tool_results|render_assistant_calls|parse_calls)$%%),
group_by_regex("Reasoning (thinking models)", mod, %regex~(split_reasoning|make_think_stream|think_feed|think_finish|think_drain|effective_stop_ids|turn_stop_ids|make_nothink_guard|nothink_stop_here)$%%),
group_by_regex("Operations: prepared images and dispatch", mod, %regex~(dlim_inventory|dlim_clean|set_dispatch_worker_limit|get_dispatch_worker_limit|set_jobque_spin_us|get_jobque_spin_us|set_jobque_spin_gpu_us|get_jobque_spin_gpu_us|get_jobque_spin_in_force|set_single_thread|get_single_thread|select_matmul_backend_for_load|kernel_backend_available)$%%),
Expand Down Expand Up @@ -402,7 +402,7 @@ def document_module_dasllama(_root : string) {
dasllama_type_stanza(f, "struct-dasllama_vision_embedder-VisionState", "VisionState",
"Caller-owned scratch for the embedder forward of whichever family is carried: the buffers ``encode_image`` reuses across calls. One per embedder user; holds no image state between calls.")
dasllama_type_stanza(f, "struct-dasllama_audio_embedder-AudioEmbedder", "AudioEmbedder",
"A loaded audio encoder of whatever family the mmproj GGUF turned out to be (gemma4a — the gemma-4 E-series Conformer), as produced by ``load_audio_embedder``, which probes the file; ``audio_probe_proj_dim`` answers 0 where absence is an answer. ``audio_proj_dim`` must match the decoder's embedding width. The server's media worker owns one per armed slot.")
"A loaded audio encoder of whatever family the mmproj GGUF turned out to be (gemma4a — the gemma-4 E-series Conformer; or a whisper-class tower that ``load_audio_tower`` also serves — qwen2-audio, qwen2.5-omni, ultravox, voxtral), as produced by ``load_audio_embedder``, which probes the file; ``audio_probe_proj_dim`` answers 0 where absence is an answer. ``audio_proj_dim`` must match the decoder's embedding width. The server's media worker owns one per armed slot.")
dasllama_type_stanza(f, "struct-dasllama_audio_embedder-AudioState", "AudioState",
"Caller-owned scratch for the audio encoder forward of whichever family is carried: the buffers ``encode_audio`` reuses across calls. One per embedder user; holds no clip state between calls.")
dasllama_type_stanza(f, "struct-dasllama_image-DlimImageInfo", "DlimImageInfo",
Expand Down
20 changes: 20 additions & 0 deletions doc/source/reference/tutorials/dasLLAMA_02_chat.rst
Original file line number Diff line number Diff line change
Expand Up @@ -116,6 +116,26 @@ memory spent only when the stream actually runs:
var toks : array<int64>
render_assistant(m, rchat, "Paris.", toks) // the exchange, as tokens

A dated system turn
===================

The Llama-3.1+ template writes two lines into every system turn: the model's
knowledge cutoff and today's date. Today's date changes the prompt every
midnight. A test that compares token streams, or a benchmark with a fixed
prompt, pins the date with ``set_chat_date``; ``""`` gives the clock back:

.. code-block:: das

set_chat_date("26 Jul 2024")
var dated = create_chat_renderer(m, SYSTEM)
add_user(dated, "What day is it?")
print(decode(m, render_turn(m, dated)))
set_chat_date("")

On Llama-3.2-1B-Instruct the system turn then reads
``Cutting Knowledge Date: December 2023`` and ``Today Date: 26 Jul 2024``.
A template that states no date - ChatML on SmolLM2, gemma - ignores the pin.

.. seealso::

Full source: :download:`tutorials/dasLLAMA/02_chat.das <../../../../tutorials/dasLLAMA/02_chat.das>`
Expand Down
81 changes: 66 additions & 15 deletions doc/source/reference/tutorials/dasLLAMA_08_audio_chat.rst
Original file line number Diff line number Diff line change
Expand Up @@ -15,9 +15,10 @@ decoder reads inline with text. Supported pairs (decoder + mmproj GGUF):
Qwen2-Audio, Qwen2.5-Omni (audio side), Ultravox v0.5 (over *stock* Llama-3
decoders), and Voxtral-Mini. The chat template picks the audio framing
automatically — the code below is identical for every pair. Qwen3-Omni and
Gemma-4 E-series audio are served too, but through the ASR surface
(:ref:`tutorial 07 <tutorial_dasLLAMA_speech_to_text>`'s two-path
``load_asr_model``), not ``load_audio_tower``.
Gemma-4 E-series audio are served too, but not by ``load_audio_tower``:
:ref:`tutorial 07 <tutorial_dasLLAMA_speech_to_text>`'s two-path
``load_asr_model`` transcribes them, and a Gemma-4 E-series pair chats on this
page through the ``AudioEmbedder`` carrier rail at the end.

Run::

Expand Down Expand Up @@ -105,34 +106,84 @@ encoder's rows between them. Unlike an image span, audio rows stay *causal*:
sound has a left-to-right order.

The rows themselves come from the ``AudioEmbedder`` carrier — the
family-neutral encoder a scheduler owns. Probe the mmproj with
``audio_probe_proj_dim`` (0 means no carrier-served audio tower), load it with
``load_audio_embedder``, and ``encode_audio`` turns 16 kHz PCM into the
soft-token rows that splice between the two spans — ``encode_image``'s audio
twin, and exactly what the server's media worker does per clip. The tutorial
probes the mmproj first: a carrier-served file (the gemma-4 E-series) takes the
carrier rail — ``section_render_spans`` plus ``section_carrier_encode`` — while
every ``load_audio_tower`` pair takes the chat rail above.
family-neutral encoder a scheduler owns. It serves every audio mmproj on this
page: the ``load_audio_tower`` families and the gemma-4 E-series alike. Probe
the mmproj with ``audio_probe_proj_dim`` (0 means no audio family serves the
file), load it with ``load_audio_embedder``, and ``encode_audio`` turns 16 kHz
PCM into the soft-token rows that splice between the two spans —
``encode_image``'s audio twin, and exactly what the server's media worker does
per clip. The tutorial asks a second question first:
``audio_tower_probe_proj_dim`` answers non-zero for a file ``load_audio_tower``
serves. Such a pair runs every section, the carrier ones last. A gemma-4
E-series file answers 0 there, so it runs only the carrier rail —
``section_render_spans`` plus ``section_carrier_encode``.

The carrier rail closes the loop at the chat layer with the pre-encoded-rows
seam, ``add_user_image_rows``'s audio twin: ``add_user_audio_rows`` moves the
encoder's rows onto a *plain* chat — no tower attached — and ``respond`` runs
the spliced turn, the audio span rendered around the rows. That is how a
carrier-served family hears in a conversation at all, and the path for a
gemma-4 E-series pair hears in a conversation at all, and the path for a
scheduler that owns its own encoder; the rows are ``dim``-wide on every family
and the call length-checks them:
and the call length-checks them. ``audio_span_bare(e)`` says whether the span
takes markers: an ultravox span sits bare in a stock Llama template, which
knows no audio marker, so pass it as ``bare``:

.. das-doc: given var rows : array<float>; let n = 0l
.. das-doc: given var rows : array<float>; let n = 0l; var e = AudioEmbedder()
.. code-block:: das

var chat <- create_chat(m, "", 96l)
add_user_audio_rows(m, chat, rows, n) // moves the rows in
add_user_audio_rows(m, chat, rows, n, audio_span_bare(e)) // moves the rows in
add_user(chat, "What did you hear?")
respond(m, chat, SamplingParams()) $(piece) {
print("{piece}")
return true
}

The audio in its place: one body by hand
========================================

So far the audio led the turn. A user message often carries its media in the
middle: "Here is a recording. <audio> What is being said?". ``add_user_span``
puts the audio where it sits: ``text_before`` bytes into the turn's text. The
span carries the media's content key — the same clip gives the same key.
``render_turn`` then writes ``n_rows`` *media position ids* where the rows go.
They are negative numbers, which no vocab id uses, so a prefix cache matches
across the span the same way it matches across text. ``render_turn_audio``
and ``render_turn_image`` take the same ``text_before`` for the two-span shape.

A program that owns its encoder and scheduler prefills that turn as *one body*
instead of three evals (text, rows, text). ``media_body_rows`` builds the
body's rows — the head text, the media rows, the tail text — and
``eval_embd_body`` prefills them. The span ``(0, 0)`` is empty, so every row
stays causal, as audio wants; an image passes its non-causal span and its
grid. On a gemma-4 E-series decoder the body does one more thing: each text
row's token id rides on the session, so text rows get their own per-layer
input, and ``eval_embd_body`` spends the ids:

.. das-doc: given let bare = false; let key = 0ul
.. code-block:: das

let lead = "Here is a recording. "
var rchat <- create_chat_renderer(m, "", 64l)
add_user(rchat, "{lead}What is being said?")
add_user_span(m, rchat, key, n, length(lead), false, bare) // audio, not image
var turn <- render_turn(m, rchat)
var lo = 0l
while (!is_media_position(turn[lo])) {
lo ++
}
var head <- [for (k in range64(lo)); turn[k]]
var tail <- [for (k in range64(lo + n, long_length(turn))); turn[k]]
var s <- create_session(m)
var body : array<float>
media_body_rows(m, s, head, rows, n, m.config.dim, tail, body)
eval_embd_body(m, s, body, long_length(turn), 0l, 0l, int2(0))
// sample(s, ...) now answers the turn

On Llama-3.2-1B with the ultravox mmproj and the JFK clip the turn renders as
``36 text tokens | 187 media ids | 13 text tokens``; on gemma-4 E2B it is
``17 | 100 | 14``. Both decoders then describe the clip from the one body.

.. seealso::

Full source: :download:`tutorials/dasLLAMA/08_audio_chat.das <../../../../tutorials/dasLLAMA/08_audio_chat.das>`
Expand Down
12 changes: 11 additions & 1 deletion doc/source/reference/tutorials/dasLLAMA_15_prefix_cache.rst
Original file line number Diff line number Diff line change
Expand Up @@ -45,14 +45,24 @@ every full page of it. The preview string is only a label for dashboards:
prefix_insert(cache, pool, s1, p1, "system prompt")
release_kv_pages(s1)

Before B takes anything, we ask the cache how much of B's prompt it holds.
``prefix_match_len`` answers with the count ``prefix_attach`` would attach
right now, and it attaches nothing and touches no entry. A server uses it to
decide what a request must still bring - the rows of a picture the cache
already holds need no encode:

.. code-block:: das

let p2 <- encode(m, "{SYSTEM}{Q2}")
print("the cache holds {prefix_match_len(cache, pool, p2)} of B's leading tokens\n")

Stream B starts fresh from the same pool. ``prefix_attach`` walks B's prompt
against the cached chains: matched pages join B's page table, B's ``n_past``
jumps past them, and we prefill only the tail. The match is capped one token
short of the prompt — the model still needs one eval to make logits:

.. code-block:: das

let p2 <- encode(m, "{SYSTEM}{Q2}")
var s2 = create_session(m, pool)
let matched = prefix_attach(cache, pool, s2, p2)
// eval() only p2[matched..] — the matched pages are already KV
Expand Down
2 changes: 1 addition & 1 deletion doc/source/reference/utils/dasllama_cli.rst
Original file line number Diff line number Diff line change
Expand Up @@ -344,7 +344,7 @@ where the server's catalog downloads land; ``DASLLAMA_MODELS_DIR`` overrides).
The ``dasllama-server.toml`` in the current directory, else in
``~/.dasllama`` (where the server's control page saves it), else beside the
program - the server's own lookup - fills whatever the flags leave empty: the model (a ``[[models]]``
roster's default entry included), its ``image_mmproj``, the ``asr`` and
roster's default entry included), its ``image_mmproj`` and ``audio_mmproj``, the ``asr`` and
``mmproj`` pair, the ``tts`` model and its lane (an ``[[asr]]`` / ``[[tts]]``
roster's first table), ``gpu`` and the Vulkan detail
keys, ``threads``, ``ctx``, ``models_dir``. On a box the server's setup page
Expand Down
Loading
Loading