Skip to content

Latest commit

 

History

History
66 lines (51 loc) · 2.46 KB

File metadata and controls

66 lines (51 loc) · 2.46 KB

Testing and Development

README

Read CONTRIBUTING.md before proposing an inference change. The release QA guide defines the hardware matrix, checkpoint-matched quality tests, long-context agent tasks, and speed checks. It also records which checks were not completed. A smoke test is not full QA.

Focused checks

Model-free checks include:

make ds4_test ds4_agent_test test-session-state
./ds4_test --server
./ds4_agent_test

On Metal, small GPU tensor tests are available without loading a full GGUF:

make tests/test_session_state_gpu tests/test_glm53_kda tests/test_mxfp4_metal
./tests/test_session_state_gpu
./tests/test_glm53_kda
./tests/test_mxfp4_metal

make test also includes model-backed tests. Select the right GGUF and ensure that it fits before running it; do not accidentally load a large model on a single device during multi-GPU QA. ROCm has make test-rocm.

Official-vector tests must use continuations from the same checkpoint as the GGUF. Flash 0731 and Vision Experimental are not interchangeable fixtures. See test vectors and quality scoring.

Investigating output

./ds4 --dump-tokens -p "..."
./ds4 --dump-logprobs /tmp/out.json --logprobs-top-k 20 --temp 0 -p "..."
./ds4 --dump-logits /tmp/logits.json --nothink --prompt-file prompt.txt
./ds4-server --trace /tmp/ds4-trace.txt

Token dumps catch template differences without inference. Logits and continuations help distinguish sampling changes from graph errors. Server traces include cache and tool-parser decisions. Keep traces private when they contain real conversations.

Tools and source references

For changes to state handling, cover rewind/replay, save/load, images, and multiple sessions as well as a fresh prompt. For distributed changes, exercise both ranks and failures; a local command-parser test is not physical TP QA.