Skip to content

Qwen3.8-27B native MTP crashes when Q8 batched GEMM exceeds MAX_BATCH=64. #732

Description

@Sergey7210

Qwen3.8-27B native MTP crashes when Q8 batched GEMM exceeds MAX_BATCH=64

Summary

Native Qwen3.8-27B MTP on gfx1200 works for short single-prompt inference, but the daemon crashes as soon as the Q8 MTP path receives a batch larger than 64:

gemm_q8_0_batched: batch_size N exceeds kernel MAX_BATCH=64

I can reproduce this in three independent workloads:

  1. an ordinary OpenAI-compatible API request with system + user messages — batch_size 70;
  2. a short multi-request battery — batch_size 76;
  3. long-context NIAH prefill — batch_size 256.

The same model is stable in plain AR mode. Reducing MTP K does not solve the problem: K1 and K7 both reproduce the API crash.

This looks like a hard kernel/runtime batch-size limit that the Qwen3.8 native-MTP path does not currently chunk or guard against.

Environment

OS:      CachyOS / Linux
GPU:     AMD Radeon RX 9060 XT 16 GB
arch:    gfx1200
CPU:     Ryzen 7 9800X3D
ROCm:    7.2.4
hipfire: 0.3.0
commit:  f4e1412e81b84e35934ba51fb453d329f4c2f041

No HSA_OVERRIDE_GFX_VERSION.

Model

Target:

repo: hipfire-models/qwen3.8-27b
file: qwen3.8-27b.mq3
size: 12,618,796,032 bytes
sha256:
09c3544690aceca29e1822d79adab6ffcc8fd9e4b58359fe8dfb185ef49811c9

The MQ3 trunk itself contains no mtp.* tensors.

Native MTP sidecar was extracted from the official Qwen/Qwen3.8-27B checkpoint using current-master mtp_extract.

Only the official shard containing all 15 MTP tensors was required.

source revision:
Qwen/Qwen3.8-27B@1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0

MTP sidecar:
qwen3.8-27b.mtp

quant:
q8

size:
451,323,904 bytes

sha256:
401fe4b90297f2939a7df9ff3d4e54d95a1662b22e5dd5c2fe11e82963caab58

round-trip:
PASS, 15 tensors, arch_id=21

Baseline: AR is healthy

With:

max_seq=8192
Q8 KV
greedy
thinking=off
max_think_tokens=1
single GPU gfx1200

AR is coherent and stable.

Representative results:

short AR:       10.7 tok/s
steady decode:  22.8–22.9 tok/s
512-token AR:   19.1–19.2 tok/s

OpenAI-compatible AR serving also works:

basic chat:          PASS
required tool calls: 3/3 PASS
tool_choice=none:    PASS

Long-context AR also succeeds:

21,553-token prompt:
  NIAH recall: PASS
  prefill:     158.8 tok/s
  decode:      3.4 tok/s

43,369-token prompt:
  NIAH recall: PASS
  prefill:     119.2 tok/s
  decode:      3.4 tok/s

So this does not appear to be a general Qwen3.8/MQ3/gfx1200 failure.

Native MTP short-prompt behavior

Native MTP loads successfully and produces the same coherent answer on a short prompt.

Example sweep:

K1: 8.3 tok/s, tau=0.79
K2: 8.6 tok/s, tau=1.27
K3: 8.6 tok/s, tau=1.47
K5: 8.5 tok/s, tau=1.71
K7: 8.1 tok/s, tau=1.71

For a longer 512-token generation, K2 is beneficial:

AR:      19.1–19.2 tok/s
MTP K2:  22.0 tok/s
delta:   approximately +14.6%
tau:     1.35

So the native MTP implementation itself is functional for compatible small batches.

Reproduction 1 — ordinary API request

Start Qwen3.8-27B MQ3 with the Q8 native-MTP sidecar and native MTP enabled.

An ordinary system + user OpenAI-compatible request crashes the daemon:

gemm_q8_0_batched: batch_size 70 exceeds kernel MAX_BATCH=64

This reproduces with both K7 and K1.

Therefore the failure does not appear to depend on a large MTP proposal depth.

Plain AR serving with the same model and API request is healthy.

Reproduction 2 — multiple short requests

A five-prompt MTP battery starts normally.

The first two requests complete successfully.

On the third request the daemon crashes:

gemm_q8_0_batched: batch_size 76 exceeds kernel MAX_BATCH=64

Subsequent requests fail with empty responses / broken pipe because the daemon has exited.

Again, the corresponding AR battery completes successfully.

Reproduction 3 — long-context prefill

With native MTP K3, Asym3 KV and VMM, the committed niah_32k fixture fails after the MTP sidecar has loaded:

gemm_q8_0_batched: batch_size 256 exceeds kernel MAX_BATCH=64

The exact same workload succeeds in AR mode.

The fixture contains 21,553 actual prompt tokens with max_seq=32768.

Expected behavior

One of the following would seem appropriate:

  1. gemm_q8_0_batched supports the batch shapes generated by native Qwen3.8 MTP;

or

  1. the caller chunks batches larger than 64 before dispatch;

or at minimum

  1. native MTP detects the unsupported shape and cleanly falls back to another path instead of terminating the daemon.

A chunked fallback would probably be preferable because batches >64 occur during completely ordinary API usage, not just synthetic long-context tests.

Actual behavior

Native MTP dispatch reaches:

gemm_q8_0_batched

with batch sizes such as:

70
76
256

and the kernel path hard-fails because:

MAX_BATCH=64

This makes native Qwen3.8 MTP unsuitable for persistent API/tool-serving workloads on this configuration even though short-prompt MTP inference itself works.

Impact

This currently blocks using native Qwen3.8 MTP for:

  • normal system+user OpenAI requests;
  • persistent/multi-request serving;
  • native tool-call workloads;
  • long-context prefill.

AR remains usable.

Possible direction

The symptom looks like a caller/kernel contract mismatch rather than an MTP acceptance issue.

It may be worth checking whether the native Qwen3.8 MTP path should use the same kind of prefill chunking already used elsewhere when the logical batch exceeds the kernel's supported maximum.

I have not attempted to modify the kernel or bypass the guard because I wanted to preserve a clean reproducer first.

I can provide the full benchmark report, daemon logs, exact API probe, and additional K/batch-size sweeps if useful.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions