Qwen3.8-27B native MTP crashes when Q8 batched GEMM exceeds MAX_BATCH=64
Summary
Native Qwen3.8-27B MTP on gfx1200 works for short single-prompt inference, but the daemon crashes as soon as the Q8 MTP path receives a batch larger than 64:
gemm_q8_0_batched: batch_size N exceeds kernel MAX_BATCH=64
I can reproduce this in three independent workloads:
- an ordinary OpenAI-compatible API request with system + user messages —
batch_size 70;
- a short multi-request battery —
batch_size 76;
- long-context NIAH prefill —
batch_size 256.
The same model is stable in plain AR mode. Reducing MTP K does not solve the problem: K1 and K7 both reproduce the API crash.
This looks like a hard kernel/runtime batch-size limit that the Qwen3.8 native-MTP path does not currently chunk or guard against.
Environment
OS: CachyOS / Linux
GPU: AMD Radeon RX 9060 XT 16 GB
arch: gfx1200
CPU: Ryzen 7 9800X3D
ROCm: 7.2.4
hipfire: 0.3.0
commit: f4e1412e81b84e35934ba51fb453d329f4c2f041
No HSA_OVERRIDE_GFX_VERSION.
Model
Target:
repo: hipfire-models/qwen3.8-27b
file: qwen3.8-27b.mq3
size: 12,618,796,032 bytes
sha256:
09c3544690aceca29e1822d79adab6ffcc8fd9e4b58359fe8dfb185ef49811c9
The MQ3 trunk itself contains no mtp.* tensors.
Native MTP sidecar was extracted from the official Qwen/Qwen3.8-27B checkpoint using current-master mtp_extract.
Only the official shard containing all 15 MTP tensors was required.
source revision:
Qwen/Qwen3.8-27B@1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0
MTP sidecar:
qwen3.8-27b.mtp
quant:
q8
size:
451,323,904 bytes
sha256:
401fe4b90297f2939a7df9ff3d4e54d95a1662b22e5dd5c2fe11e82963caab58
round-trip:
PASS, 15 tensors, arch_id=21
Baseline: AR is healthy
With:
max_seq=8192
Q8 KV
greedy
thinking=off
max_think_tokens=1
single GPU gfx1200
AR is coherent and stable.
Representative results:
short AR: 10.7 tok/s
steady decode: 22.8–22.9 tok/s
512-token AR: 19.1–19.2 tok/s
OpenAI-compatible AR serving also works:
basic chat: PASS
required tool calls: 3/3 PASS
tool_choice=none: PASS
Long-context AR also succeeds:
21,553-token prompt:
NIAH recall: PASS
prefill: 158.8 tok/s
decode: 3.4 tok/s
43,369-token prompt:
NIAH recall: PASS
prefill: 119.2 tok/s
decode: 3.4 tok/s
So this does not appear to be a general Qwen3.8/MQ3/gfx1200 failure.
Native MTP short-prompt behavior
Native MTP loads successfully and produces the same coherent answer on a short prompt.
Example sweep:
K1: 8.3 tok/s, tau=0.79
K2: 8.6 tok/s, tau=1.27
K3: 8.6 tok/s, tau=1.47
K5: 8.5 tok/s, tau=1.71
K7: 8.1 tok/s, tau=1.71
For a longer 512-token generation, K2 is beneficial:
AR: 19.1–19.2 tok/s
MTP K2: 22.0 tok/s
delta: approximately +14.6%
tau: 1.35
So the native MTP implementation itself is functional for compatible small batches.
Reproduction 1 — ordinary API request
Start Qwen3.8-27B MQ3 with the Q8 native-MTP sidecar and native MTP enabled.
An ordinary system + user OpenAI-compatible request crashes the daemon:
gemm_q8_0_batched: batch_size 70 exceeds kernel MAX_BATCH=64
This reproduces with both K7 and K1.
Therefore the failure does not appear to depend on a large MTP proposal depth.
Plain AR serving with the same model and API request is healthy.
Reproduction 2 — multiple short requests
A five-prompt MTP battery starts normally.
The first two requests complete successfully.
On the third request the daemon crashes:
gemm_q8_0_batched: batch_size 76 exceeds kernel MAX_BATCH=64
Subsequent requests fail with empty responses / broken pipe because the daemon has exited.
Again, the corresponding AR battery completes successfully.
Reproduction 3 — long-context prefill
With native MTP K3, Asym3 KV and VMM, the committed niah_32k fixture fails after the MTP sidecar has loaded:
gemm_q8_0_batched: batch_size 256 exceeds kernel MAX_BATCH=64
The exact same workload succeeds in AR mode.
The fixture contains 21,553 actual prompt tokens with max_seq=32768.
Expected behavior
One of the following would seem appropriate:
gemm_q8_0_batched supports the batch shapes generated by native Qwen3.8 MTP;
or
- the caller chunks batches larger than 64 before dispatch;
or at minimum
- native MTP detects the unsupported shape and cleanly falls back to another path instead of terminating the daemon.
A chunked fallback would probably be preferable because batches >64 occur during completely ordinary API usage, not just synthetic long-context tests.
Actual behavior
Native MTP dispatch reaches:
with batch sizes such as:
and the kernel path hard-fails because:
This makes native Qwen3.8 MTP unsuitable for persistent API/tool-serving workloads on this configuration even though short-prompt MTP inference itself works.
Impact
This currently blocks using native Qwen3.8 MTP for:
- normal system+user OpenAI requests;
- persistent/multi-request serving;
- native tool-call workloads;
- long-context prefill.
AR remains usable.
Possible direction
The symptom looks like a caller/kernel contract mismatch rather than an MTP acceptance issue.
It may be worth checking whether the native Qwen3.8 MTP path should use the same kind of prefill chunking already used elsewhere when the logical batch exceeds the kernel's supported maximum.
I have not attempted to modify the kernel or bypass the guard because I wanted to preserve a clean reproducer first.
I can provide the full benchmark report, daemon logs, exact API probe, and additional K/batch-size sweeps if useful.
Qwen3.8-27B native MTP crashes when Q8 batched GEMM exceeds MAX_BATCH=64
Summary
Native Qwen3.8-27B MTP on
gfx1200works for short single-prompt inference, but the daemon crashes as soon as the Q8 MTP path receives a batch larger than 64:I can reproduce this in three independent workloads:
batch_size 70;batch_size 76;batch_size 256.The same model is stable in plain AR mode. Reducing MTP K does not solve the problem: K1 and K7 both reproduce the API crash.
This looks like a hard kernel/runtime batch-size limit that the Qwen3.8 native-MTP path does not currently chunk or guard against.
Environment
No
HSA_OVERRIDE_GFX_VERSION.Model
Target:
The MQ3 trunk itself contains no
mtp.*tensors.Native MTP sidecar was extracted from the official
Qwen/Qwen3.8-27Bcheckpoint using current-mastermtp_extract.Only the official shard containing all 15 MTP tensors was required.
Baseline: AR is healthy
With:
AR is coherent and stable.
Representative results:
OpenAI-compatible AR serving also works:
Long-context AR also succeeds:
So this does not appear to be a general Qwen3.8/MQ3/gfx1200 failure.
Native MTP short-prompt behavior
Native MTP loads successfully and produces the same coherent answer on a short prompt.
Example sweep:
For a longer 512-token generation, K2 is beneficial:
So the native MTP implementation itself is functional for compatible small batches.
Reproduction 1 — ordinary API request
Start Qwen3.8-27B MQ3 with the Q8 native-MTP sidecar and native MTP enabled.
An ordinary system + user OpenAI-compatible request crashes the daemon:
This reproduces with both K7 and K1.
Therefore the failure does not appear to depend on a large MTP proposal depth.
Plain AR serving with the same model and API request is healthy.
Reproduction 2 — multiple short requests
A five-prompt MTP battery starts normally.
The first two requests complete successfully.
On the third request the daemon crashes:
Subsequent requests fail with empty responses / broken pipe because the daemon has exited.
Again, the corresponding AR battery completes successfully.
Reproduction 3 — long-context prefill
With native MTP K3, Asym3 KV and VMM, the committed
niah_32kfixture fails after the MTP sidecar has loaded:The exact same workload succeeds in AR mode.
The fixture contains 21,553 actual prompt tokens with
max_seq=32768.Expected behavior
One of the following would seem appropriate:
gemm_q8_0_batchedsupports the batch shapes generated by native Qwen3.8 MTP;or
or at minimum
A chunked fallback would probably be preferable because batches >64 occur during completely ordinary API usage, not just synthetic long-context tests.
Actual behavior
Native MTP dispatch reaches:
with batch sizes such as:
and the kernel path hard-fails because:
This makes native Qwen3.8 MTP unsuitable for persistent API/tool-serving workloads on this configuration even though short-prompt MTP inference itself works.
Impact
This currently blocks using native Qwen3.8 MTP for:
AR remains usable.
Possible direction
The symptom looks like a caller/kernel contract mismatch rather than an MTP acceptance issue.
It may be worth checking whether the native Qwen3.8 MTP path should use the same kind of prefill chunking already used elsewhere when the logical batch exceeds the kernel's supported maximum.
I have not attempted to modify the kernel or bypass the guard because I wanted to preserve a clean reproducer first.
I can provide the full benchmark report, daemon logs, exact API probe, and additional K/batch-size sweeps if useful.