Skip to content

fix: refuse an over-long state before tokenizing every prompt; sturdier tests and release - #23

Merged
feder-cr merged 1 commit into
mainfrom
robustness
Oct 1, 2026
Merged

feder-cr merged 1 commit into
mainfrom
robustness

Conversation

@feder-cr

@feder-cr feder-cr commented Oct 1, 2026

Copy link
Copy Markdown
Owner

Fixes from a review of the code, the tests and the docs (three parallel reviews, each finding checked against the code before fixing).

Server

  • Over-long state: a state longer than the context made every prompt too long, but all prompts were tokenized first, each holding the whole state. A 250 KB state with 1,024 questions (or 40 choices of 26 options) took every core and gigabytes of memory, and an allocation failure on a tokenizing thread ended the process. The state is now measured first and the request refused with the same error as before (the first question's prompt length): 422 in 0.22 s, memory unchanged (1.03 GB peak).
  • Token buffers: tokens no longer keep a buffer sized to their text.
  • Thread errors: an error on a tokenizing thread fails its request instead of std::terminate; an error during idle warm-up stops the warm-up instead of the server.
  • Score probabilities: a score level's probability is never below 0 (pooled means could round to −1e-17).
  • Duplicate snapshots: two batched requests with the same new state keep one snapshot, not two.

Tests, CI, release

  • check.py:
    • refuses to run if its port already answers (it used to test that server instead of the build);
    • stops a server that did not start, and waits by time rather than by tries;
    • puts a timeout on every request and resolves paths from where it was started;
    • does not pass JEV_API_KEY on;
    • makes the choice and score checks independent of step order.
  • score-test: 20,000 random threshold sequences, no level below 0, levels sum to 1.
  • Workflows: jobs have timeouts. release.yml removes SHA256SUMS.txt before replacing the archives and attaches it again once all three are in, so a failed or partial run never leaves a stale one.

Docs: about 40 README and wiki pages were stale:

  • they still said score is refused;
  • they described the 8,192-token context as per request (it is per question's prompt);
  • they said every response carries Server-Timing (only successful ones do);
  • the offline page left out the build's pip install and its network step.

The options table gains --name, --ctx and --warmup.

Verified locally: full check.py --exact --threads 8 green, with zero-tolerance parity 215/215 identical (no numerical change).

🤖 Generated with Claude Code

…er tests and release

Found in a review of the code, the tests and the docs.

Server
- A state longer than the context made every prompt too long, but each one was tokenized first, all
  holding the whole state: a 250 KB state with 1,024 questions, or 40 choices of 26 options, took
  every core and gigabytes, and an allocation failure on a tokenizing thread ended the process. The
  state is now measured first and the request refused with the same error the Python loop stopped
  at (the first question's prompt length): 0.22 s, memory unchanged.
- Tokens no longer keep a buffer sized to their text (4 bytes per byte of each prompt).
- An error on a tokenizing thread fails its request instead of calling std::terminate; an error
  while warming up when idle stops the warm-up instead of the server.
- A score level's probability is never below 0 (pooled means could round to -1e-17).
- Two requests read together with the same new state keep one snapshot of it, not two.

Tests, CI, release
- check.py refuses to run if its port already answers (it tested that server instead of the
  build), stops a server that did not start, waits by time rather than by tries, puts a timeout on
  every request, resolves paths from where it was started, does not pass JEV_API_KEY on, and has
  the choice and score checks resume from a snapshot so they do not depend on the order of steps.
- score-test: 20,000 random threshold sequences, no level below 0 and levels summing to 1.
- Workflows have timeouts; release.yml removes SHA256SUMS.txt before replacing the archives and
  attaches it again once all three are in, so a failed or partial run never leaves a stale one.

Docs: about 40 README and wiki pages still said score is refused, that the 8,192-token context is
per request (it is per question's prompt), that every response carries Server-Timing (successful
ones do), or left out the build's pip install and network step; the options table gains --name,
--ctx and --warmup.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@feder-cr
feder-cr merged commit d110f6d into main Oct 1, 2026
6 checks passed
@feder-cr
feder-cr deleted the robustness branch October 1, 2026 04:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant