Skip to content

Improve the eval runner: surfaces, repeats, error handling, eval location - #87

Open
mattpodwysocki wants to merge 1 commit into
mainfrom
eval-runner-improvements
Open

mattpodwysocki wants to merge 1 commit into
mainfrom
eval-runner-improvements

Conversation

@mattpodwysocki

Copy link
Copy Markdown
Contributor

Split out of #80 so each change can be reviewed on its own. Four changes to how evals are run and scored.

Eval definitions move out of skills/

Both plugin manifests ship everything under skills/ — .claude-plugin/plugin.json declares "skills": "./skills/", and the Codex package copies the tree — so evals.json living there handed every installed agent the expected answers for the suite that grades these skills.

They now live in evals/<skill-name>/evals.json. Filtering the packagers would have been the narrower fix, but there are two of them and nothing stops a third; keeping answers out of skills/ holds for any packager. validate-skills.js and CONTRIBUTING.md follow the move.

AGENTS.md becomes a second eval surface

AGENTS.md is maintained by hand, ships to users as a drop-in, and no eval ever read it — so the two copies of a skill could disagree with nothing noticing.

--surface=skill|agents|both selects what to grade. Default is skill: grading both doubles the cost of every run and makes the headline total incomparable with earlier baselines, so the copy gets checked deliberately rather than on every invocation.

--repeats=N measures variance

Every eval ran once, so a score change could not be told apart from judge jitter — one eval scored 15/15 and 11/15 on identical content. Repeats report a mean, a 95% interval, and a VARIANCE section naming evals whose spread across identical runs exceeds 0.05, since an eval that unstable cannot support a decision whatever its mean is.

Failed runs are no longer scored as zero

max_tokens was 4096 — low enough that a thinking-heavy response could spend the whole budget before emitting any text. The judge then scored nothing against every expectation and the eval read as a 0% skill failure.

Default is now 8192 (EVAL_MAX_TOKENS to override), an empty response throws, and — the part that was missing on the first attempt — a failed run is excluded from the mean, the per-skill and per-surface totals, and the saved results. Failures get their own ERRORS section:

Evals scored:     0
Evals failed:     3 (excluded from all scores below)
Total score:      n/a — no eval produced a usable response

══════════════════════════════════════════════════
ERRORS (not scored)
══════════════════════════════════════════════════
  mapbox-cartography #1 [skill] — Empty response from claude-sonnet-5
    (stop_reason: max_tokens, output_tokens: 1). Raise EVAL_MAX_TOKENS …

A request that produced no response is missing data, not evidence about the skill.

Test plan

  • npm run check passes, including validate:skills against the new eval location
  • Discovery verified from evals/<skill>/ — a normal run scores 47/48
  • Error path verified by forcing failures with EVAL_MAX_TOKENS=1: 3 failed, 0 scored, totals report n/a rather than 0%
  • Confirmed no evals.json ships in either package after the move

Follow-up

evals/baseline.json is not regenerated here — it would need re-running after this merges, and bundling a 6,000-line regenerated baseline into a review of the runner itself would bury the diff.

🤖 Generated with Claude Code

…tion

Four changes to how evals are run and scored, split out of the caching PR
so each can be reviewed on its own.

**Eval definitions move out of skills/.** Both plugin manifests ship
everything under skills/ — .claude-plugin/plugin.json declares
"skills": "./skills/" and the Codex package copies the tree — so an
evals.json there handed every installed agent the expected answers for the
suite that grades these skills. They now live in evals/<skill>/evals.json.
Filtering the packagers would have been the narrower fix, but there are two
of them and nothing stops a third; keeping answers out of skills/ holds for
any packager. validate-skills.js and CONTRIBUTING.md follow the move.

**AGENTS.md becomes a second eval surface.** AGENTS.md is maintained by
hand, ships to users as a drop-in, and no eval ever read it, so the two
copies of a skill could disagree with nothing noticing.
--surface=skill|agents|both selects what to grade. The default is skill:
grading both doubles the cost of every run and makes the headline total
incomparable with earlier baselines, so the copy is checked deliberately
rather than on every invocation.

**--repeats=N measures variance.** Every eval ran once, so a score change
could not be told apart from judge jitter — one eval scored 15/15 and
11/15 on identical content. Repeats report a mean, a 95% interval, and a
VARIANCE section naming evals whose spread across identical runs exceeds
0.05, since an eval that unstable cannot support a decision whatever its
mean is.

**Failed runs are no longer scored as zero.** max_tokens was 4096, low
enough that a thinking-heavy response could spend the whole budget before
emitting any text; the judge then scored nothing against every expectation
and the eval read as a 0% skill failure. The default is now 8192
(EVAL_MAX_TOKENS to override), an empty response throws, and — the part
that was missing — a failed run is excluded from the mean, the per-skill
and per-surface totals, and the saved results. Failures are listed in
their own ERRORS section instead, because a request that produced no
response is missing data, not evidence about the skill.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant