Repository navigation
feat: publish the tested models and warn on untested ones - #95
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Prompts behave differently on different models, so a result from one agent and model says little about another. This change publishes which agent and model combinations the companion's evaluation has been tested with and how each scored, and warns the reviewer who picks a combination that was never tested.
The evaluation output lists every tested combination with its agent, agent version, model, effort, run date and scores. The README links to the current list. Choosing an untested combination in the settings shows a clear, non-blocking warning: every review still runs and its results stay stamped with who answered.
Validation never launches or installs the VS Code application or any tool. Extension tests that need VS Code run only in CI, so live validation goes through the engine protocol, the unit tests and rendered output.
Closes #3.
What Changed
TESTED_MODELSandisTestedModel()to the engine, recording every agent, model and effort combination the evaluation has run with its version, run date and scores, and published the list as the newdocs/tested-models.mdpage linked from the README.testedCombinations()from a run's stamped results and prints oneTESTEDline per combination (agent, version, model, effort, run date, per-prompt scores) in its report, leaving out model-free or model-unnamed runs.second-look.agentorsecond-look.agentModelchanges — names the tested models and the published page, and carries the warning in the status bar tooltip.Risk Assessment
✅ Low: The change is well-bounded — a published constant and predicate in the engine, a derived listing in the evaluation report, and a non-blocking settings warning in the extension — with behavioral tests, docs that faithfully mirror the recorded baseline, and all three acceptance criteria satisfied; the only observation is a documented deliberate tradeoff.
Testing
Drove the real evaluation CLI three times against a disposable scripted Pi stand-in over a throwaway copy of the example-7 case: the report prints a TESTED line carrying the agent, version, model, effort, run date and per-prompt scores, while a run whose agent never named its model and a model-free run print none. The README's links to docs/tested-models.md resolve and the published page matches the built engine's TESTED_MODELS field-by-field and score-by-score, rendered as an HTML artifact. The settings warning was exercised only through the intent's sanctioned local routes — 78 targeted unit tests plus a reviewer-journey driver running the real extension code (activate, configuration listener, status bar) against the repository's VS Code test double, with a rendered artifact of the exact warning and tooltip strings — because the ticket's standing boundary forbids launching VS Code on this machine; since these routes are not a live drive of the real product, that scenario is reported untested and its non-live corroboration is preserved in the artifacts. UI-facing evidence is rendered HTML rather than screenshots for that reason. No failures; transient fixtures removed and the worktree left clean.
pistand-in on PATH; artifact eval-run-tested-line.txt…Evidence: Evaluation CLI report with the TESTED line (live run, scripted Pi stand-in)
TESTED pi 9.9.9 fake/glm-fake default (run 2026-10-07T19:20:30.505Z): grouping: coverage 1, grouping: grouping-agreement 1, ranking: rank-median 1.5, ranking: rank-top-3 1, story: story-must-review 1, story: story-order 1, story: story-names 0.6667Evidence: Adversarial run: agent ended before naming its model — rows stamped 'unknown model', no TESTED line
Evidence: Adversarial run: model-free run prints no TESTED line
Evidence: results.json of the live agent run (stamped rows the TESTED line is built from)
Evidence: trace.jsonl of the live agent run (each agent call with its stamp)
Evidence: README link resolution + docs page vs built engine TESTED_MODELS cross-check
README.md: 2 link(s) to docs/tested-models.md packages/evaluation/README.md: 1 link(s) to docs/tested-models.md docs/tested-models.md lists 1 tested combination(s): pi 0.86.1 zai-coding-cn/glm-5.3 (default)~/.no-mistakes/evidence/01M4BWAYJ5MGWS1D7MMZMD19F8/tested-models-page.html)Evidence: Rendered untested-combination warning surface (notification and status-bar tooltip from the exact strings the real extension code produced)
Source: Rendered untested-combination warning surface (notification and status-bar tooltip from the exact strings the real extension code produced) (local file:
~/.no-mistakes/evidence/01M4BWAYJ5MGWS1D7MMZMD19F8/untested-model-warning.html)Evidence: Reviewer-journey driver over the real extension code (temporary test, removed after the run)
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
packages/extension/src/agent-settings.ts:104- When second-look.agentModel is empty (the default) and the agent has any tested model, untestedModelWarning stays quiet even if the agent's actual default model differs from the tested one, so a fresh install can run an untested combination with no warning. This is a deliberate, documented tradeoff (the settings cannot name the agent's default model without probing it; the run stamp discloses what answered), recorded here only to acknowledge it was examined.✅ **Test** - passed
✅ No issues found.
pistand-in on PATH; artifact eval-run-tested-line.txt…npm ci && npm run build (materialize dependencies and build engine/evaluation/extension)npx vitest run packages/engine/test/tested-models.test.ts packages/evaluation/test/run.test.ts packages/evaluation/test/cli.test.ts packages/extension/test/agent-status.test.ts (78 tests passed)Live CLI drive: PATH=<disposable bin>:$PATH node packages/evaluation/dist/main.js run --cases <throwaway copy of example-7> --agent pi --runs <temp> — with a scriptedpistand-in on PATH that answers the lockdown probe and the grouping/ranking/story prompts; report printed 'TESTED pi 9.9.9 fake/glm-fake default (run 2026-10-07T19:20:30.505Z): grouping: coverage 1, grouping: grouping-agreement 1, ranking: rank-median 1.5, ranking: rank-top-3 1, story: story-must-review 1, story: story-order 1, story: story-names 0.6667', exit 0, results.json/trace.jsonl stampedAdversarial CLI drives: same run with FAKE_PI_NO_MODEL=1 (agent rows stamped 'unknown model', zero TESTED lines) and a --model-free run over example-42 (zero TESTED lines), both exit 0node .tmp-live/verify-docs.mjs <evidence> — resolved every markdown link in README.md (2) and packages/evaluation/README.md (1) to docs/tested-models.md, and cross-checked the page field-by-field and score-by-score against the built engine's TESTED_MODELS export; rendered the page to HTMLEVIDENCE_DIR=<evidence> npx vitest run packages/extension/test/tmp-untested-warning-render.test.ts — temporary reviewer-journey driver (removed after the run) over the real activate(), configuration-change listener and AgentStatusBar through the repository's VS Code test double: quiet for the tested default, warning when picking anthropic/claude-sonnet-5, quiet for zai-coding-cn/glm-5.3, warning for Claude Code, no error messages, tooltip carries the warning beside the API-key one; rendered the exact warning strings to HTML (non-live: test double, not the VS Code host)✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.