Skip to content

feat: verify a claim, and ask what covers a part - #93

Merged
lbildzinkas merged 7 commits into
masterfrom
fm/second-look-ask-verify-cover
Oct 7, 2026
Merged

lbildzinkas merged 7 commits into
masterfrom
fm/second-look-ask-verify-cover

Conversation

@lbildzinkas

@lbildzinkas lbildzinkas commented Oct 7, 2026 •

Copy link
Copy Markdown
Owner

Intent

Closes #48.

This change adds two asks to a part's context menu, following the explain ask from #47 and the judging pass from #34.

Verify this claim runs the judging pass on one claim: text the reviewer selects on the head side of the part's diff, or one of the part's claims the reviewer picks. When the claim needs a library's source, the verdict offers the library fetch, which the reviewer still starts.

What covers this? lists the automated tests, in the change or in the head copy, that exercise the part, and the manual checks the pull request reports for it. When nothing covers the part, it says that none were found.

Both prompts land with their evaluation cases and scores, and "none found" is a valid, scored answer. Pull request text reaches the agent only as fenced, untrusted data.

Validation never launches, downloads or installs VS Code or any other tool. Extension tests that need the editor run only in CI; local validation goes through the engine protocol, the unit tests and the rendered output.

What Changed

  • Added the Verify this claim ask to the engine's ASKS registry (verify.ts): it runs the verdicts judging pass on the text the reviewer selected on the head side of a part's diff or on a claim they pick, re-reads every citation against the head copy and CI logs, offers the library fetch when needed, and keeps the judged claim — marked as judged by the ask, so its verdict stands even when the claims pass fell back — in the review's claims (review result protocol v17).
  • Added the What covers this? ask with a new versioned cover prompt (cover.ts): it lists the automated tests in the change or head copy that exercise the part and the manual checks the pull request's description reports, every cited test line re-read in the head copy and every manual quote checked against the description, with "none found" a valid answer shown as such.
  • Wired both asks through the extension (context-menu commands, diff-selection-to-claim mapping with a claim quick-pick fallback, judged claim joining the shown review as a finding) and the evaluation package (cover prompt registered at version 1 with verdicts bumped to 5, new verify/cover cases and scores including none-found, refreshed baseline).

Risk Assessment

✅ Low: The change is well-bounded and follows the repo's established ask/prompt/evaluation patterns; every prior fix-round decision (asked mark, fell-back/unjudged protocol acceptance, verdictsNote and tree accounting, fetchLibrary preservation of the asked mark) is correctly implemented with behavioral tests, and I found no reachable wrong-result, disclosure, or regression path in the intended usage.

Testing

Drove the real engine process over its JSON-RPC protocol against a disposable GitHub/PyPI fixture and a stand-in agent: the verify ask judged a picked claim and a diff selection (marked asked, cited head lines re-read), its library fetch offer was pressed and the fetched verdict landed while the claims pass stayed fell back, a fell-back empty listing and a judged listing both accepted asked claims with every updated review passing the extension protocol guard, five malformed or misplaced asks were refused with plain messages, and the cover ask returned a covering test line plus the description's manual check and a valid scored 'None found'; the reviewer-facing overview and tree rendered the asked-claim accounting from those live results, and the evaluation's own loaders confirmed both prompts registered at the engine's versions with verify/cover cases including none-found. Everything passed; transient harness files were removed and the worktree is clean.

  • Live validation: ✅ go - 9 of 9 scenarios driven live against the product
Scenario Result Live Evidence
Verify this claim by picking a listed claim: the judging pass runs on it alone and the answer shows the claim, its verdict with evidence source, and is stamped by the run that answered ✅ pass live live-checks.txt lines 'verify (picked claim)…' against the real engine process; full wire transcript in live-transcript.jsonrpc.txt
Verify this claim by selecting text on a head-side diff line: the selection becomes a reviewer-sourced claim judged singly (marked asked) and joined to the engine's latest review after the others, wit… ✅ pass live live-checks.txt 'verify (selection)…' checks; the judged claim with asked:true appears in live-ask-answers.json and live-review-after-fetch.json
The verdict offers a library fetch only when the claim needs one; nothing downloads until the reviewer presses it from the finding, and pressing it downloads the pinned wheel, judges the claim in the… ✅ pass live live-checks.txt 'the pressed fetch…' and 'the fetched verdict lands…'; the pinned httpx 0.27.2 wheel was served by the disposable PyPI fixture and the fetchLibrary answers are in live-transcript.jsonr…
An asked claim stands in every listing and judging state: a fell-back judging pass, a fell-back empty listing with no verdicts pass, and a judged listing all accept it, and each updated review parses… ✅ pass live live-checks.txt checks for reviews 1–3 (isReviewResult run on every engine answer inside the live drive); persisted states in live-review-after-fetch.json, live-review-2-after-fetch.json, live-review-…
The reviewer-facing overview and tree render the asked-claim accounting in every judging state: the verdicts note attributes checked verdicts to the ask, the claim detail says 'judged singly by the Ve… ✅ pass live rendered-overview-*.html pages and the tree assertions rendered from the live-captured results by the extension's real overviewHtml/buildTree code (VS Code itself is never launched on this machine, pe…
Adversarial: the ask refuses a verify with no claim, a claim on an ask that takes none, a claim belonging to another part, a selection whose text is not on those head lines, and a selection over the 5… ✅ pass live live-checks.txt 'verify without a claim…', 'a claim on an ask that takes none…', "a claim of another part is refused plainly…", 'a selection not on those lines…', 'a selection over the cap…'
What covers this? on a part its tests exercise: the answer lists the covering test line re-read in the head copy as a cited line, and reports the manual check the description carries with its descript… ✅ pass live live-checks.txt 'cover (fresh part)…' checks citing tests/test_fresh.py:4 and the description's manual check at line 3; rendered in rendered-overview-asks.html
What covers this? on a part nothing covers: 'None found' is a valid answer with its own section heading, where the agent looked, and no citations ✅ pass live live-checks.txt 'cover (uncovered part) says none found…'; rendered in rendered-overview-asks.html
Both prompts land with their cases and scores: prompts.json registers cover v1 and verdicts v5 at the engine's own versions, and the evaluation's real case loaders read verify selections (library fetc… ✅ pass live eval-cases-checks.txt, produced by driving the evaluation package's loadRegistry/loadCase over prompts.json and every case folder; scoring machinery covered by packages/evaluation tests
Evidence: Full JSON-RPC transcript of the live engine drive (3 reviews, 7 asks, 2 library fetches, 17 stage notifications)
> {"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":1}}
< {"jsonrpc":"2.0","id":1,"result":{"protocolVersion":1}}
> {"jsonrpc":"2.0","id":2,"method":"review","params":{"url":"https://github.com/example-org/example-repo/pull/7","token":"ghp_live-drive-token"}}
< {"jsonrpc":"2.0","method":"review/stage","params":{"id":2,"running":"grouping related hunks with pi","timeoutMs":660000,"result":{"version":17,"pullRequest":{"url":"https://github.com/example-org/example-repo/pull/7","number":7,"title":"Reformat the Python helpers and tax the cart total","author":"reviewer-login","description":"Reformats app/reformat.py, moves the flag out of the discount block, restyles Greeter and taxes the cart total.\n\n> Ran the test suite locally: every test passed.","base":"master","head":"tidy-and-tax","baseCommit":"5555555555555555555555555555555555555555","headSha":"7777777777777777777777777777777777777777"},"copies":{"base":{"commit":"6666666666666666666666666666666666666666","path":"~/.no-mistakes/worktrees/a09930ebf9a5/01M4B2A2NKK4QDQ8Y05HBFTRKJ/scratch-live/cache/github.com/example-org/example-repo/pull-7/6666666666666666666666666666666666666666","reused":false},"head":{"commit":"7777777777777777777777777777777777777777","path":"~/.no-mistakes/worktrees/a09930ebf9a5/01M4B2A2NKK4QDQ8Y05HBFTRKJ/scratch-live/cache/github.com/example-org/example-repo/pull-7/7777777777777777777777777777777777777777","reused":false}},"parseTimeMs":6.277,"parts":[{"path":"app/fresh.py","changeKind":"addition","isBinary":false,"oldMissingFinalNewline":false,"newMissingFinalNewline":false,"hunks":[{"oldStart":0,"oldLines":0,"newStart":1,"newLines":2,"lines":[{"kind":"addition","newLineNumber":1,"text":"def fresh():"},{"kind":"addition","newLineNumber":2,"text":"    return 1"}],"entities":[{"kind":"function","name":"fresh","public":true,"change":"added"}]}],"additions":2,"deletions":0,"syntax":{"language":"python","formattingOnly":{"status":"structure-changed","reason":"the file is new"},"checksNotRun":[]},"newMode":"100644","noise":{"label":"none","note":"no rule applied"},"name":"fresh in app/fresh.py","origin":"plain","signals":{"novelty":"new","role":"code","changedLines":2,"publicSurface":["fresh"],"references":{"basis":"name-based","names":["fresh"],"files":1}},"rank":{"importance":"must review","reason":"changes the public surface: fresh; code; new code; 2 changed lines","signals":["changes the public surface: fresh","code","new code","2 changed lines"]}},{"path":"web/cart.ts","changeKind":"modification","isBinary":false,"oldMissingFinalNewline":false,"newMissingFinalNewline":false,"hunks":[{"oldStart":2,"oldLines":7,"newStart":2,"newLines":7,"heading":"export class Cart {","lines":[{"kind":"context","oldLineNumber":2,"newLineNumber":2,"text":"  items: number[] = [];"},{"kind":"context","oldLineNumber":3,"newLineNumber":3,"text":""},{"kind":"context","oldLineNumber":4,"newLineNumber":4,"text":"  total(): number {"},{"kind":"deletion","oldLineNumber":5,"text":"    return this.items.reduce((sum, item) => sum + item, 0);"},{"kind":"addition","newLineNumber":5,"text":"    return this.items.reduce((sum, item) => sum + item, 0) * 1.2;"},{"kind":"context","oldLineNumber":6,"newLineNumber":6,"text":"  }"},{"kind":"context","oldLineNumber":7,"newLineNumber":7,"text":"}"},{"kind":"context","oldLineNumber":8,"newLineNumber":8,"text":""}],"entities":[{"kind":"method","name":"Cart.total","public":true,"change":"body"}]}],"additions":1,"deletions":1,"syntax":{"language":"typescript","formattingOnly":{"status":"structure-changed","reason":"the syntax tree changes at head line 5"},"checksNotRun":[]},"noise":{"label":"none","note":"no rule applied"},"name":"Cart.total in web/cart.ts","origin":"plain","signals":{"novelty":"changed","role":"code","changedLines":2,"publicSurface":[],"references":{"basis":"name-based","names":["total"],"files":2}},"rank":{"importance":"worth reviewing","reason":"code; changes code named in 2 other files (name-based); 2 changed lines","signals":["code","changes code named in 2 other files (name-based)","2 changed lines"]}},{"path":"tests/test_fresh.py","changeKind":"addition","isBinary":false,"oldMissingFinalNewline":false,"newMissingFinalNewline":false,"hunks":[{"oldStart":0,"oldLines":0,"newStart":1,"newLines":5,"lines":[{"kind":"addition","newLineNumber":1,"text":"from app.fresh import fresh"},{"kind":"addition","newLineNumber":2,"text":""},{"kind":"addition","newLineNumber":3,"text":""},{"kind":"addition","newLineNumber":4,"text":"def test_fresh():"},{"kind":"addition","newLineNumber":5,"text":"    assert fresh() == 1"}],"entities":[{"kind":"function","name":"test_fresh","public":true,"change":"added"}]}],"additions":5,"deletions":0,"syntax":{"language":"python","formattingOnly":{"status":"structure-changed","reason":"the file is new"},"checksNotRun":[]},"newMode":"100644","noise":{"label":"none","note":"no rule applied"},"name":"test_fresh in tests/test_fresh.py","origin":"plain","signals":{"novelty":"new","role":"test","changedLines":5,"publicSurface":["test_fresh"],"references":{"basis":"name-based","names":["test_fresh"],"files":0}},"rank":{"importance":"context","reason":"test; new code; 5 changed lines","signals":["test","new code","5 changed lines"]}},{"path":"app/dedent.py","changeKind":"modification","isBinary":false,"oldMissingFinalNewline":false,"newMissingFinalNewline":false,"hunks":[{"oldStart":1,"oldLines":5,"newStart":1,"newLines":5,"lines":[{"kind":"context","oldLineNumber":1,"newLineNumber":1,"text":"def apply_discount(order):"},{"kind":"context","oldLineNumber":2,"newLineNumber":2,"text":"    if order.total > 100:"},{"kind":"context","oldLineNumber":3,"newLineNumber":3,"text":"        order.total -= 10"},{"kind":"deletion","oldLineNumber":4,"text":"        order.flag = True"},{"kind":"addition","newLineNumber":4,"text":"    order.flag = True"},{"kind":"context","oldLineNumber":5,"newLineNumber":5,"text":"    return order"}],"entities":[{"kind":"function","name":"apply_discount","public":true,"change":"body"}]}],"additions":1,"deletions":1,"syntax":{"language":"python","formattingOnly":{"status":"structure-changed","reason":"the syntax tree changes at head line 4"},"checksNotRun":[]},"noise":{"label":"none","note":"no rule applied"},"name":"apply_discount in app/dedent.py","origin":"plain","signals":{"novelty":"changed","role":"code","changedLines":2,"publicSurface":[],"references":{"basis":"name-based","names":["apply_discount"],"files":0}},"rank":{"importance":"context","reason":"code; 2 changed lines","signals":["code","2 changed lines"]}},{"path":"scripts/deploy.rb","changeKind":"modification","isBinary":false,"oldMissingFinalNewline":false,"newMissingFinalNewline":false,"hunks":[{"oldStart":1,"oldLines":3,"newStart":1,"newLines":3,"lines":[{"kind":"context","oldLineNumber":1,"newLineNumber":1,"text":"def deploy(target)"},{"kind":"deletion","oldLineNumber":2,"text":"  puts \"deploying to #{target}\""},{"kind":"addition","newLineNumber":2,"text":"  puts \"deploying to #{target} now\""},{"kind":"context","oldLineNumber":3,"newLineNumber":3,"text":"end"}],"entities":[]}],"additions":1,"deletions":1,"syntax":{"formattingOnly":{"status":"not-checked","reason":"no grammar for \".rb\" files, so its hunks are read at file level"},"checksNotRun":[{"check":"entities","reason":"no grammar for \".rb\" files, so its hunks are read at file level"},{"check":"formatting-only","reason":"no grammar for \".rb\" files, so its hunks are read at file level"}]},"noise":{"label":"none","note":"no rule applied"},"name":"scripts/deploy.rb","origin":"plain","signals":{"novelty":"changed","role":"code","changedLines":2,"publicSurface":[],"references":{"basis":"name-based","names":[],"files":0}},"rank":{"importance":"context","reason":"code; 2 changed lines","signals":["code","2 changed lines"]}},{"path":"app/reformat.py","changeKind":"modification","isBinary":false,"oldMissingFinalNewline":false,"newMissingFinalNewline":false,"hunks":[{"oldStart":1,"oldLines":15,"newStart":1,"newLines":14,"lines":[{"kind":"context","oldLineNumber":1,"newLineNumber":1,"text":"import os"},{"kind":"context","oldLineN

... [310670 bytes truncated] ...

ore:"},{"kind":"addition","newLineNumber":10,"text":"    def __init__(self, root):"},{"kind":"addition","newLineNumber":11,"text":"        self.root = root"},{"kind":"context","oldLineNumber":12,"newLineNumber":12,"text":""},{"kind":"deletion","oldLineNumber":13,"text":"    def path_for(self,name):"},{"kind":"deletion","oldLineNumber":14,"text":"        return os.path.join( self.root,"},{"kind":"deletion","oldLineNumber":15,"text":"                             name )"},{"kind":"addition","newLineNumber":13,"text":"    def path_for(self, name):"},{"kind":"addition","newLineNumber":14,"text":"        return os.path.join(self.root, name)"}],"entities":[{"kind":"function","name":"load","public":true,"change":"declaration"},{"kind":"class","name":"Store","public":true,"change":"declaration"},{"kind":"method","name":"Store.__init__","public":true,"change":"declaration"},{"kind":"method","name":"Store.path_for","public":true,"change":"declaration"}]}],"additions":8,"deletions":9,"syntax":{"language":"python","formattingOnly":{"status":"confirmed","reason":"base and head have the same syntax tree, nesting included; only formatting changed"},"checksNotRun":[]},"noise":{"label":"none","note":"no rule applied"},"name":"load, Store, Store.__init__ and 1 more in app/reformat.py","origin":"plain","signals":{"novelty":"changed","role":"code","changedLines":17,"publicSurface":["load","Store","Store.__init__","Store.path_for"],"references":{"basis":"name-based","names":["load","Store","path_for"],"files":0}},"rank":{"importance":"context","reason":"formatting only, confirmed by the syntax trees; code; 17 changed lines","signals":["formatting only, confirmed by the syntax trees","code","17 changed lines"]}},{"path":"src/Greeter.cs","changeKind":"modification","isBinary":false,"oldMissingFinalNewline":false,"newMissingFinalNewline":false,"hunks":[{"oldStart":1,"oldLines":9,"newStart":1,"newLines":7,"lines":[{"kind":"context","oldLineNumber":1,"newLineNumber":1,"text":"namespace Demo;"},{"kind":"context","oldLineNumber":2,"newLineNumber":2,"text":""},{"kind":"deletion","oldLineNumber":3,"text":"public class Greeter"},{"kind":"deletion","oldLineNumber":4,"text":"{"},{"kind":"deletion","oldLineNumber":5,"text":"    public string Greet(string name)"},{"kind":"deletion","oldLineNumber":6,"text":"    {"},{"kind":"deletion","oldLineNumber":7,"text":"        return \"Hello, \" + name;"},{"kind":"deletion","oldLineNumber":8,"text":"    }"},{"kind":"addition","newLineNumber":3,"text":"public class Greeter {"},{"kind":"addition","newLineNumber":4,"text":"  public string Greet(string name) {"},{"kind":"addition","newLineNumber":5,"text":"    return \"Hello, \" + name;"},{"kind":"addition","newLineNumber":6,"text":"  }"},{"kind":"context","oldLineNumber":9,"newLineNumber":7,"text":"}"}],"entities":[{"kind":"class","name":"Greeter","public":true,"change":"declaration"},{"kind":"method","name":"Greeter.Greet","public":true,"change":"declaration"}]}],"additions":4,"deletions":6,"syntax":{"language":"c-sharp","formattingOnly":{"status":"confirmed","reason":"base and head have the same syntax tree, nesting included; only formatting changed"},"checksNotRun":[]},"noise":{"label":"none","note":"no rule applied"},"name":"Greeter, Greeter.Greet in src/Greeter.cs","origin":"plain","signals":{"novelty":"changed","role":"code","changedLines":10,"publicSurface":["Greeter","Greeter.Greet"],"references":{"basis":"name-based","names":["Greeter","Greet"],"files":0}},"rank":{"importance":"context","reason":"formatting only, confirmed by the syntax trees; code; 10 changed lines","signals":["formatting only, confirmed by the syntax trees","code","10 changed lines"]}}],"grouping":{"by":"plain","agent":{"promptVersion":"2","stamp":{"agent":"pi","agentVersion":"1.2.3","model":"fake/live-drive","effort":null,"runAt":"2026-10-07T13:20:51.165Z","tokens":{"input":20,"output":20,"cacheRead":0,"cacheWrite":0,"total":40}},"outcome":"fell back","detail":"the agent gave no usable answer (invalid-answer: the answer was invalid twice: the answer is not a single JSON value)","leftOut":0}},"ranking":{"by":"plain","agent":{"promptVersion":"1","stamp":{"agent":"pi","agentVersion":"1.2.3","model":"fake/live-drive","effort":null,"runAt":"2026-10-07T13:20:51.346Z","tokens":{"input":20,"output":20,"cacheRead":0,"cacheWrite":0,"total":40}},"outcome":"fell back","detail":"the agent gave no usable answer (invalid-answer: the answer was invalid twice: the answer is not a single JSON value)"}},"criteria":{"outcome":"read","detail":"this pull request links no issue","heading":"Acceptance criteria","issues":[],"criteria":[]},"pipeline":{"attestation":"missing","detail":"the description carries no no-mistakes attestation","steps":[],"findings":[]},"ci":{"headSha":"7777777777777777777777777777777777777777","outcome":"read","detail":"0 check runs at the head commit, 0 failed; logs are read only for failed jobs","checks":[]},"story":{"promptVersion":"1","stamp":{"agent":"pi","agentVersion":"1.2.3","model":"fake/live-drive","effort":null,"runAt":"2026-10-07T13:20:51.526Z","tokens":{"input":20,"output":20,"cacheRead":0,"cacheWrite":0,"total":40}},"outcome":"fell back","detail":"the agent gave no usable answer (invalid-answer: the answer was invalid twice: the answer is not a single JSON value)","sentences":[]},"unexplained":{"promptVersion":"1","parts":[],"described":[],"outcome":"fell back","detail":"the agent gave no usable answer (invalid-answer: the answer was invalid twice: the answer is not a single JSON value)","stamp":{"agent":"pi","agentVersion":"1.2.3","model":"fake/live-drive","effort":null,"runAt":"2026-10-07T13:20:51.717Z","tokens":{"input":20,"output":20,"cacheRead":0,"cacheWrite":0,"total":40}}},"claims":{"promptVersion":"1","stamp":{"agent":"pi","agentVersion":"1.2.3","model":"fake/live-drive","effort":null,"runAt":"2026-10-07T13:20:51.897Z","tokens":{"input":10,"output":10,"cacheRead":0,"cacheWrite":0,"total":20}},"outcome":"listed","detail":"every quote was found in its source, which locates the claim and, for a docstring or comment, its part","claims":[{"quote":"taxes the cart total.","source":"description","location":{"kind":"description","line":1},"part":0,"verdict":{"kind":"unverifiable","source":"the model's memory","reason":"The description does not say what the total is taxed to.","evidence":[]}},{"quote":"Reformats app/reformat.py","source":"description","location":{"kind":"description","line":1},"part":5,"verdict":{"kind":"unverifiable","source":"the model's memory","reason":"The description does not say how it reformats.","evidence":[]}}],"judging":{"promptVersion":"5","stamp":{"agent":"pi","agentVersion":"1.2.3","model":"fake/live-drive","effort":null,"runAt":"2026-10-07T13:20:51.987Z","tokens":{"input":10,"output":10,"cacheRead":0,"cacheWrite":0,"total":20}},"outcome":"judged","detail":"every citation was re-read in the head copy; one that did not match, or the model's memory alone, kept a claim from verified"}}}}
> {"jsonrpc":"2.0","id":17,"method":"ask","params":{"url":"https://github.com/example-org/example-repo/pull/7","ask":"verify","part":0,"claim":{"selection":{"path":"app/fresh.py","line":1,"endLine":1,"text":"def fresh():"}}}}
< {"jsonrpc":"2.0","id":17,"result":{"ask":"verify","part":0,"partName":"fresh in app/fresh.py","sections":[{"heading":"Claim","text":"\"def fresh():\", made in text the reviewer selected in the diff of \"app/fresh.py\", line 1, and asked to verify."},{"heading":"Verdict","text":"refuted, from the change itself: `fresh` returns 1, not the total."}],"cited":[{"path":"app/fresh.py","side":"head","line":2,"quote":"return 1"}],"promptVersion":"5","stamp":{"agent":"pi","agentVersion":"1.2.3","model":"fake/live-drive","effort":null,"runAt":"2026-10-07T13:20:52.250Z","tokens":{"input":10,"output":10,"cacheRead":0,"cacheWrite":0,"total":20}},"claim":{"index":2,"claim":{"quote":"def fresh():","source":"reviewer","location":{"kind":"file","path":"app/fresh.py","line":1,"endLine":1},"part":0,"verdict":{"kind":"refuted","source":"the change itself","reason":"`fresh` returns 1, not the total.","evidence":[{"path":"app/fresh.py","line":2,"quote":"return 1"}]},"asked":true}}}}
Evidence: The 34 live engine-protocol checks, all passing
PASS  initialize answers with the protocol version
PASS  review 1 answers with a result the extension protocol accepts
PASS  review 1 lists the description claim and its judging fell back — outcome=listed claims=1 judging=fell back
PASS  the review has the fresh part and the Greeter part — fresh=0 greeter=6 parts=fresh in app/fresh.py | Cart.total in web/cart.ts | test_fresh in tests/test_fresh.py | apply_discount in app/dedent.py | scripts/deploy.rb | load, Store, Store.__init__ and 1 more in app/reformat.py | Greeter, Greeter.Greet in src/Greeter.cs
PASS  the fresh part shows head line 1 of app/fresh.py
PASS  verify (picked claim) answers with an ask answer the extension accepts
PASS  verify (picked claim) shows the claim, its verdict and the library fetch offer — [{"heading":"Claim","text":"\"taxes the cart total.\", made in the pull request's description, line 1."},{"heading":"Verdict","text":"unverifiable, from the change itself: It turns on httpx."},{"heading":"Library fetch","text":"The change alone cannot settle this claim: it turns on how httpx behaves, so checking it needs the source of httpx 0.27.2, as requirements.txt pins it. Press the fetch on the claim's finding to start it."}]
PASS  verify (picked claim) judges the claim alone, marked asked, offering the fetch — {"library":"httpx","pinnedVersion":"0.27.2","pinnedBy":"requirements.txt","reason":"The change alone cannot settle this claim: it turns on how httpx behaves, so checking it needs the source of httpx 0.27.2, as requirements.txt pins it."}
PASS  verify (picked claim) is stamped by the run that answered
PASS  verify (selection) answers with an ask answer the extension accepts
PASS  verify (selection) judges the selection as a reviewer claim after the others, marked asked — {"quote":"def fresh():","source":"reviewer","location":{"kind":"file","path":"app/fresh.py","line":1,"endLine":1},"part":0,"verdict":{"kind":"refuted","source":"the change itself","reason":"`fresh` returns 1, not the total.","evidence":[{"path":"app/fresh.py","line":2,"quote":"return 1"}]},"asked":true}
PASS  verify (selection) cites the head line its evidence re-read — [{"path":"app/fresh.py","side":"head","line":2,"quote":"return 1"}]
PASS  verify (selection) section text names where the claim is made — "def fresh():", made in text the reviewer selected in the diff of "app/fresh.py", line 1, and asked to verify.
PASS  the pressed fetch answers with the whole updated review, which the extension protocol accepts
PASS  the fetched verdict lands on the asked claim while the pass still fell back — judging=fell back claim0={"kind":"refuted","source":"library source at the pinned version","reason":"A client follows no redirect by default.","evidence":[{"path":"httpx/_client.py","li
PASS  the selection verdict is kept beside the fetched one
PASS  verify without a claim is refused with the params message — ask needs params: { "url": string, "ask": "explain" | "verify" | "cover", "part": number, "claim"?: { "index": number } | { "selection": { "path": string, "line": number, "endLine": number, "text": string } }, "agent"?: { "agent": "pi" | "claude-code", "model"?: string, "account"?: string } }, with the claim for "verify" only
PASS  a claim on an ask that takes none is refused — ask needs params: { "url": string, "ask": "explain" | "verify" | "cover", "part": number, "claim"?: { "index": number } | { "selection": { "path": string, "line": number, "endLine": number, "text": string } }, "agent"?: { "agent": "pi" | "claude-code", "model"?: string, "account"?: string } }, with the claim for "verify" only
PASS  a claim of another part is refused plainly — claim 0 is not one of this part's claims
PASS  a selection not on those lines of the head copy is refused plainly — the selection is not on lines 1-1 of app/fresh.py in the head copy
PASS  a selection over the cap is refused plainly — the selection is over 500 characters; select one claim
PASS  cover (fresh part) answers with an ask answer the extension accepts
PASS  cover (fresh part) lists the covering test as a cited head line — [{"heading":"What covers it","text":"`test_fresh` calls `fresh`, and the description reports a local test-suite run."},{"heading":"Manual checks the pull request reports","text":"\"Ran the test suite locally: every test passed.\" (the description, line 3)"}] [{"path":"tests/test_fresh.py","side":"head","line":4,"quote":"def test_fresh():"}]
PASS  cover (fresh part) reports the manual check the description carries, with its line — "Ran the test suite locally: every test passed." (the description, line 3)
PASS  cover (uncovered part) answers with an ask answer the extension accepts
PASS  cover (uncovered part) says none found, as the section's own heading — [{"heading":"None found","text":"No test exercises the restyle, and the description reports no manual check of it; I searched `tests`."}]
PASS  review 2 answers with a result the extension protocol accepts
PASS  review 2 fell back to no claim, with no judging pass — outcome=fell back claims=0 judging=undefined
PASS  verify (selection) joins the fell-back empty listing as an asked claim offering its fetch — {"quote":"def fresh():","source":"reviewer","location":{"kind":"file","path":"app/fresh.py","line":1,"endLine":1},"part":0,"verdict":{"kind":"unverifiable","source":"the change itself","reason":"It turns on httpx.","needsLibrary":"httpx","evidence":[],"libraryFetch":{"library":"httpx","pinnedVersion":"0.27.2","pinnedBy":"requirements.txt","reason":"The change alone cannot settle this claim: it turns on how httpx behaves, so checking it needs the source of httpx 0.27.2, as requirements.txt pins it."}},"asked":true}
PASS  the pressed fetch on the asked claim in the fell-back listing answers a review the extension accepts
PASS  the fetched verdict lands on the asked claim while the listing still fell back and no verdicts pass ran — {"kind":"refuted","source":"library source at the pinned version","reason":"A client follows no redirect by default.","evidence":[{"path":"httpx/_client.py","li
PASS  review 3 answers with a result the extension protocol accepts
PASS  review 3 judged both its claims — judging=judged claims=["unverifiable","unverifiable"]
PASS  verify (selection) joins the judged listing as an asked claim — {"quote":"def fresh():","source":"reviewer","location":{"kind":"file","path":"app/fresh.py","line":1,"endLine":1},"part":0,"verdict":{"kind":"refuted","source":"the change itself","reason":"`fresh` re

engine stderr:
  • Evidence: Rendered overview page right after the first Verify this claim ask: judged-singly detail and the library fetch offer (local file: ~/.no-mistakes/evidence/01M4B2A2NKK4QDQ8Y05HBFTRKJ/rendered-overview-after-first-verify.html)
  • Evidence: Rendered overview after the library fetch landed on the asked claims in a fell-back-judging review (local file: ~/.no-mistakes/evidence/01M4B2A2NKK4QDQ8Y05HBFTRKJ/rendered-overview-listing-fell-back.html)
  • Evidence: Rendered overview of the fell-back empty listing holding the asked, fetched claim (local file: ~/.no-mistakes/evidence/01M4B2A2NKK4QDQ8Y05HBFTRKJ/rendered-overview-fell-back-empty.html)
  • Evidence: Rendered overview of a judged listing with the asked claim attributed to the ask (local file: ~/.no-mistakes/evidence/01M4B2A2NKK4QDQ8Y05HBFTRKJ/rendered-overview-judged-with-asked.html)
  • Evidence: Rendered overview showing both asks' answers: verify with its Library fetch section, cover with its test citation, manual check and None found (local file: ~/.no-mistakes/evidence/01M4B2A2NKK4QDQ8Y05HBFTRKJ/rendered-overview-asks.html)
  • Evidence: The engine's persisted review result after the verify asks and the pressed library fetch (local file: ~/.no-mistakes/evidence/01M4B2A2NKK4QDQ8Y05HBFTRKJ/live-review-after-fetch.json)
  • Evidence: The engine's persisted review for the fell-back empty listing after its asked claim's fetch was pressed (local file: ~/.no-mistakes/evidence/01M4B2A2NKK4QDQ8Y05HBFTRKJ/live-review-2-after-fetch.json)
  • Evidence: The engine's persisted review with both claims judged, before the asked selection joined (local file: ~/.no-mistakes/evidence/01M4B2A2NKK4QDQ8Y05HBFTRKJ/live-review-3-judged.json)
Evidence: The ask answers the live engine produced (verify picked/selection, cover found/none found)
{
  "picked": {
    "ask": "verify",
    "part": 0,
    "partName": "fresh in app/fresh.py",
    "sections": [
      {
        "heading": "Claim",
        "text": "\"taxes the cart total.\", made in the pull request's description, line 1."
      },
      {
        "heading": "Verdict",
        "text": "unverifiable, from the change itself: It turns on httpx."
      },
      {
        "heading": "Library fetch",
        "text": "The change alone cannot settle this claim: it turns on how httpx behaves, so checking it needs the source of httpx 0.27.2, as requirements.txt pins it. Press the fetch on the claim's finding to start it."
      }
    ],
    "cited": [],
    "promptVersion": "5",
    "stamp": {
      "agent": "pi",
      "agentVersion": "1.2.3",
      "model": "fake/live-drive",
      "effort": null,
      "runAt": "2026-10-07T13:20:48.243Z",
      "tokens": {
        "input": 10,
        "output": 10,
        "cacheRead": 0,
        "cacheWrite": 0,
        "total": 20
      }
    },
    "claim": {
      "index": 0,
      "claim": {
        "quote": "taxes the cart total.",
        "source": "description",
        "location": {
          "kind": "description",
          "line": 1
        },
        "part": 0,
        "verdict": {
          "kind": "unverifiable",
          "source": "the change itself",
          "reason": "It turns on httpx.",
          "needsLibrary": "httpx",
          "evidence": [],
          "libraryFetch": {
            "library": "httpx",
            "pinnedVersion": "0.27.2",
            "pinnedBy": "requirements.txt",
            "reason": "The change alone cannot settle this claim: it turns on how httpx behaves, so checking it needs the source of httpx 0.27.2, as requirements.txt pins it."
          }
        },
        "asked": true
      }
    }
  },
  "selected": {
    "ask": "verify",
    "part": 0,
    "partName": "fresh in app/fresh.py",
    "sections": [
      {
        "heading": "Claim",
        "text": "\"def fresh():\", made in text the reviewer selected in the diff of \"app/fresh.py\", line 1, and asked to verify."
      },
      {
        "heading": "Verdict",
        "text": "refuted, from the change itself: `fresh` returns 1, not the total."
      }
    ],
    "cited": [
      {
        "path": "app/fresh.py",
        "side": "head",
        "line": 2,
        "quote": "return 1"
      }
    ],
    "promptVersion": "5",
    "stamp": {
      "agent": "pi",
      "agentVersion": "1.2.3",
      "model": "fake/live-drive",
      "effort": null,
      "runAt": "2026-10-07T13:20:48.509Z",
      "tokens": {
        "input": 10,
        "output": 10,
        "cacheRead": 0,
        "cacheWrite": 0,
        "total": 20
      }
    },
    "claim": {
      "index": 1,
      "claim": {
        "quote": "def fresh():",
        "source": "reviewer",
        "location": {
          "kind": "file",
          "path": "app/fresh.py",
          "line": 1,
          "endLine": 1
        },
        "part": 0,
        "verdict": {
          "kind": "refuted",
          "source": "the change itself",
          "reason": "`fresh` returns 1, not the total.",
          "evidence": [
            {
              "path": "app/fresh.py",
              "line": 2,
              "quote": "return 1"
            }
          ]
        },
        "asked": true
      }
    }
  },
  "covered": {
    "ask": "cover",
    "part": 0,
    "partName": "fresh in app/fresh.py",
    "sections": [
      {
        "heading": "What covers it",
        "text": "`test_fresh` calls `fresh`, and the description reports a local test-suite run."
      },
      {
        "heading": "Manual checks the pull request reports",
        "text": "\"Ran the test suite locally: every test passed.\" (the description, line 3)"
      }
    ],
    "cited": [
      {
        "path": "tests/test_fresh.py",
        "side": "head",
        "line": 4,
        "quote": "def test_fresh():"
      }
    ],
    "promptVersion": "1",
    "stamp": {
      "agent": "pi",
      "agentVersion": "1.2.3",
      "model": "fake/live-drive",
      "effort": null,
      "runAt": "2026-10-07T13:20:49.029Z",
      "tokens": {
        "input": 10,
        "output": 10,
        "cacheRead": 0,
        "cacheWrite": 0,
        "total": 20
      }
    }
  },
  "none": {
    "ask": "cover",
    "part": 6,
    "partName": "Greeter, Greeter.Greet in src/Greeter.cs",
    "sections": [
      {
        "heading": "None found",
        "text": "No test exercises the restyle, and the description reports no manual check of it; I searched `tests`."
      }
    ],
    "cited": [],
    "promptVersion": "1",
    "stamp": {
      "agent": "pi",
      "agentVersion": "1.2.3",
      "model": "fake/live-drive",
      "effort": null,
      "runAt": "2026-10-07T13:20:49.286Z",
      "tokens": {
        "input": 10,
        "output": 10,
        "cacheRead": 0,
        "cacheWrite": 0,
        "total": 20
      }
    }
  },
  "reselected": {
    "ask": "verify",
    "part": 0,
    "partName": "fresh in app/fresh.py",
    "sections": [
      {
        "heading": "Claim",
        "text": "\"def fresh():\", made in text the reviewer selected in the diff of \"app/fresh.py\", line 1, and asked to verify."
      },
      {
        "heading": "Verdict",
        "text": "unverifiable, from the change itself: It turns on httpx."
      },
      {
        "heading": "Library fetch",
        "text": "The change alone cannot settle this claim: it turns on how httpx behaves, so checking it needs the source of httpx 0.27.2, as requirements.txt pins it. Press the fetch on the claim's finding to start it."
      }
    ],
    "cited": [],
    "promptVersion": "5",
    "stamp": {
      "agent": "pi",
      "agentVersion": "1.2.3",
      "model": "fake/live-drive",
      "effort": null,
      "runAt": "2026-10-07T13:20:50.646Z",
      "tokens": {
        "input": 10,
        "output": 10,
        "cacheRead": 0,
        "cacheWrite": 0,
        "total": 20
      }
    },
    "claim": {
      "index": 0,
      "claim": {
        "quote": "def fresh():",
        "source": "reviewer",
        "location": {
          "kind": "file",
          "path": "app/fresh.py",
          "line": 1,
          "endLine": 1
        },
        "part": 0,
        "verdict": {
          "kind": "unverifiable",
          "source": "the change itself",
          "reason": "It turns on httpx.",
          "needsLibrary": "httpx",
          "evidence": [],
          "libraryFetch": {
            "library": "httpx",
            "pinnedVersion": "0.27.2",
            "pinnedBy": "requirements.txt",
            "reason": "The change alone cannot settle this claim: it turns on how httpx behaves, so checking it needs the source of httpx 0.27.2, as requirements.txt pins it."
          }
        },
        "asked": true
      }
    }
  },
  "judgedSelected": {
    "ask": "verify",
    "part": 0,
    "partName": "fresh in app/fresh.py",
    "sections": [
      {
        "heading": "Claim",
        "text": "\"def fresh():\", made in text the reviewer selected in the diff of \"app/fresh.py\", line 1, and asked to verify."
      },
      {
        "heading": "Verdict",
        "text": "refuted, from the change itself: `fresh` returns 1, not the total."
      }
    ],
    "cited": [
      {
        "path": "app/fresh.py",
        "side": "head",
        "line": 2,
        "quote": "return 1"
      }
    ],
    "promptVersion": "5",
    "stamp": {
      "agent": "pi",
      "agentVersion": "1.2.3",
      "model": "fake/live-drive",
      "effort": null,
      "runAt": "2026-10-07T13:20:52.250Z",
      "tokens": {
        "input": 10,
        "output": 10,
        "cacheRead": 0,
        "cacheWrite": 0,
        "total": 20
      }
    },
    "claim": {
      "index": 2,
      "claim": {
        "quote": "def fresh():",
        "source": "reviewer",
        "location": {
          "kind": "file",
          "path": "app/fresh.py",
          "line": 1,
          "endLine": 1
        },
        "part": 0,
        "verdict": {
          "kind": "refuted",
          "source": "the change itself",
          "reason": "`fresh` returns 1, not the total.",
          "evidence": [
            {
              "path": "app/fresh.py",
              "line": 2,
              "quote": "return 1"
            }
          ]
        },
        "asked": true
      }
    }
  }
}
Evidence: Evaluation registry and case-loader checks: both prompts at the engine's versions, verify selections and cover expectations including none found

PASS the verdicts prompt (the verify ask's prompt) is registered at the engine's version — registry=5 engine=5 PASS the cover prompt is registered at the engine version, pointing at cover.ts — registry=1 engine=1 PASS cases carry verify selections — 4 selections in 3 cases PASS cases carry cover expectations — 7 cover expectations in 5 cases PASS a verify selection expects the library fetch a claim needs PASS cover expectations include none found (the scored 'None found' answer) PASS cover expectations include parts covered by tests and by manual checks

PASS  the verdicts prompt (the verify ask's prompt) is registered at the engine's version — registry=5 engine=5
PASS  the cover prompt is registered at the engine version, pointing at cover.ts — registry=1 engine=1 files=packages/engine/src/cover.ts
PASS  cases carry verify selections — 4 selections in 3 cases
PASS  cases carry cover expectations — 7 cover expectations in 5 cases
PASS  a verify selection expects the library fetch a claim needs — Microsoft.IO.RecyclableMemoryStream, none, httpx, none
PASS  cover expectations include none found (the scored 'None found' answer) — doc_page in app/doc_links.py: 0 tests, 0 manual | load in src/tomli/_parser.py: 1 tests, 1 manual | isTicked, criteriaOf in packages/engine/src/criteria.ts: 1 tests, 0 manual | criterionItem in packages/extension/src/overview.ts: 1 tests, 1 manual | HTTPParser.wait_ready, HTTPParser in src/httpx/_parsers.py: 0 tests, 0 manual | HTTPServer.wait in src/httpx/_server.py: 0 tests, 0 manual | deepMergeInternal in source/utils/merge.ts: 2 tests, 0 manual
PASS  cover expectations include parts covered by tests and by manual checks
ALL CHECKS PASSED

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 2 issues found → auto-fixed (4) ✅
  • ⚠️ packages/engine/src/verify.ts:159 - withVerifiedClaim joins a claim with a checked verdict into a review whose claims pass fell back, producing a review result the extension's protocol rejects — so a pressed library fetch is wasted. Concrete sequence: (1) a review completes with the claims listed but the judging pass fallen back (packages/engine/src/review.ts:455 spreads judged.judging; every claim stays 'not checked', judging.outcome 'fell back' — a state the README documents as real); (2) the reviewer runs Verify this claim on a picked claim or a selection: the ask succeeds, and withVerifiedClaim (verify.ts:159-166) keeps the claims pass's judging untouched while inserting a checked verdict — the same merge is applied engine-side at server.ts:550 and extension-side at extension.ts:770-775; (3) the reviewer presses the offered library fetch on the claim's finding: the engine downloads and judges (server.ts:422), then answers fetchLibrary with the full updated review, which the extension refuses to parse because isClaims (packages/extension/src/protocol.ts:507) requires every claim's verdict to be 'not checked' unless judging.outcome === 'judged'. The reviewer sees a ProtocolError after the fetch already happened, and the fetched verdict is stranded in the engine (a second press hits a claim whose verdict.library is already set, and claimToVerify at verify.ts:71 refuses to re-judge it). The invariant 'a claim carries a checked verdict only once the claims were judged' (protocol.ts comment on Claims) is violated by construction whenever the pass itself fell back; the same violation affects every consumer of that invariant, namely the extension validator at protocol.ts:507. The smallest honest remedy is a protocol-semantics decision — accept claims judged singly by the verify ask in isClaims, or record the verify-ask judgement in the claims pass — so the remedy, not the defect, needs the author's call.
  • ℹ️ packages/engine/src/protocol.ts:13 - REVIEW_RESULT_VERSION is bumped to 17 (the reviewer claim source, ReviewerSelection, and the ask answer's judged claim), but the version-history comment at protocol.ts:15-45 stops at version 16; every earlier bump added its sentence. Add the version 17 sentence to keep the documented contract complete.

🔧 Fix applied.
2 warnings still open:

  • ⚠️ packages/extension/src/protocol.ts:507 - Sibling of round 1's verify-claim finding that the fix round left behind: the fix taught the verdict disjunct (protocol.ts:509-511) to accept an asked claim's checked verdict, but the other rule in isClaims — a claims listing that fell back may hold only pipeline-sourced claims — still rejects the review the verify ask produces. Concrete sequence: (1) a review's claims LISTING falls back (claimsStage, review.ts:410-425, still prepends pipeline findings, so claims.claims is an array, possibly empty, and claims.outcome stays 'fell back'); (2) the reviewer selects text on the head side and runs Verify this claim — engine-side claimToVerify (verify.ts:60-83) only refuses when result.claims?.claims is undefined, so the selection is judged and withVerifiedClaim (verify.ts:159-166) appends a reviewer-sourced claim marked asked:true to a listing whose outcome stays 'fell back'; (3) the ask answer itself parses (isVerifiedClaim applies no source rule), but when the verdict's offered library fetch is pressed from the finding, the engine answers fetchLibrary with the whole updated review, and isClaims at protocol.ts:507 rejects it because not every claim's source is 'pipeline' — the reviewer sees a ProtocolError after the download and judging already ran, and the fetched verdict is stranded (every later press answers the same shape and fails the same check). Display sibling of the same rule: overview.ts:768 prints 'Only the pipeline's claims are listed' over a list that then holds the reviewer's asked claim. The remedy is again a protocol-semantics decision the recorded round-1 instruction did not cover — accept an asked reviewer claim in a fell-back listing (loosen the documented 'fallen back with only the pipeline's' contract, its comment, and the overview note), or have the engine refuse a selection verify while the listing fell back — so the remedy, not the defect, needs the author's call.
  • ⚠️ packages/extension/src/overview.ts:741 - Sibling the round-1 fix round left behind: verdictsNote's judging-undefined branch still returns 'None is checked yet.' while an asked claim in the same list shows a checked verdict. The fix extended only the outcome==='fell back' branch (overview.ts:742-747) to count asked claims. Concrete sequence: (1) a review lists zero claims and the pipeline report contributes none, so claimsStage leaves claims.claims empty and verdictsStage (review.ts:443) skips, leaving judging undefined — an ordinary state for a small pull request; (2) the reviewer selects a head-side line and runs Verify this claim — the ticket's primary by-hand path, and exactly the case the extension's claimToVerify handles with no claims to pick (extension.ts:791-794); (3) the judged claim (asked:true, checked verdict) joins the review, claimsSection now renders the list (claims.claims.length is 1), and the note above the claim's own checked verdict reads 'None is checked yet.' — a wrong label that contradicts the verdict shown directly beneath it, where the fell-back branch now says the checked verdicts came from the ask. Fix by giving the undefined branch the same asked-claim accounting the fell-back branch got (and note the verdicts pass never ran, rather than implying it is pending).

🔧 Fix applied.
1 warning still open:

  • ⚠️ packages/extension/src/tree.ts:263 - Sibling the round-2 fix round (98d0ab4) left behind, despite its recorded instruction to sweep every consumer that reasons about the listing or judging outcome: withClaims's !judged state still writes 'not checked yet; the overview lists them' while an asked claim in the same part carries a checked verdict. Concrete sequence: (1) a review's claims listing falls back or holds no claim, so the verdicts pass never runs (judging undefined) or its judging fell back — exactly the state the engine's new server test exercises (listing fell back empty, judging undefined); (2) the reviewer selects a head-side line and runs Verify this claim: the judged claim joins the review marked asked:true with a refuted or unverifiable verdict, and extension.ts's this.show(updated, true) rebuilds the tree; (3) findingCounts (consumed at tree.ts:124) counts the asked claim, so the part's row reads '⚠ 1 findings · 1 claim' while its tooltip says '1 claim, not checked yet; the overview lists them' — a wrong label contradicting both the badge beside it and the overview's own note ('The verdicts pass did not run; the one checked verdict came from the Verify this claim ask.'). Before this change the combination findings>0 with judged false was unreachable, which the doc comment at tree.ts:110-111 ('a badge counting its findings once the claims are judged') still asserts. Remedy: give the tree the same asked-claim accounting verdictsNote got (count claims with asked===true; when the pass did not judge and some are asked, say the verdicts pass did not run or fell back and the checked verdicts came from the ask), and update the comment at tree.ts:110.

🔧 Fix applied.
1 warning still open:

  • ⚠️ packages/extension/src/overview.ts:745 - Sibling the round-2 fix round (98d0ab4) left behind, despite its recorded instruction to sweep every note that reasons about the judging outcome: verdictsNote's judged branch still attributes every claim's verdict to the verdicts pass's stamp, while an asked claim in the same list was judged singly by the Verify this claim ask — valid in a judged listing too, per the recorded principle. Concrete sequence: (1) a review completes with its verdicts pass judged — the ordinary outcome for any review with claims; (2) the reviewer selects a head-side line and runs Verify this claim, the ticket's primary by-hand path; withVerifiedClaim (verify.ts:153-166) appends the reviewer-sourced claim marked asked:true and judging stays 'judged'; (3) claimsSection renders the note 'Each is judged against the change, its read-only copy and any failed check's CI log by <pass stamp>…' over a list whose last item's own detail (overview.ts:695) reads 'judged singly by the Verify this claim ask', stamped with the ask's own stamp (possibly a different agent/model chosen at ask time) — a wrong attribution that contradicts the detail shown directly beneath it, in the state the primary hand path most often produces. The undefined-judging (overview.ts:743) and fell-back (overview.ts:744) branches got the asked-claim accounting; the judged branch is the only remaining site (the tree's judged branch in tree.ts:264-267 reports only counts and stays accurate, and the 'How these results were made' Verdicts row at overview.ts:896-905 describes the pass itself). Fix by appending the same fromAsk accounting to the judged branch, e.g. '…by <stamp>, save the N checked verdict(s) that came from the Verify this claim ask; the refuted and unverifiable ones are findings…'.

🔧 Fix applied.
✅ Re-checked - no issues remain.

✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 9 of 9 scenarios driven live against the product
Scenario Result Live Evidence
Verify this claim by picking a listed claim: the judging pass runs on it alone and the answer shows the claim, its verdict with evidence source, and is stamped by the run that answered ✅ pass live live-checks.txt lines 'verify (picked claim)…' against the real engine process; full wire transcript in live-transcript.jsonrpc.txt
Verify this claim by selecting text on a head-side diff line: the selection becomes a reviewer-sourced claim judged singly (marked asked) and joined to the engine's latest review after the others, wit… ✅ pass live live-checks.txt 'verify (selection)…' checks; the judged claim with asked:true appears in live-ask-answers.json and live-review-after-fetch.json
The verdict offers a library fetch only when the claim needs one; nothing downloads until the reviewer presses it from the finding, and pressing it downloads the pinned wheel, judges the claim in the… ✅ pass live live-checks.txt 'the pressed fetch…' and 'the fetched verdict lands…'; the pinned httpx 0.27.2 wheel was served by the disposable PyPI fixture and the fetchLibrary answers are in live-transcript.jsonr…
An asked claim stands in every listing and judging state: a fell-back judging pass, a fell-back empty listing with no verdicts pass, and a judged listing all accept it, and each updated review parses… ✅ pass live live-checks.txt checks for reviews 1–3 (isReviewResult run on every engine answer inside the live drive); persisted states in live-review-after-fetch.json, live-review-2-after-fetch.json, live-review-…
The reviewer-facing overview and tree render the asked-claim accounting in every judging state: the verdicts note attributes checked verdicts to the ask, the claim detail says 'judged singly by the Ve… ✅ pass live rendered-overview-*.html pages and the tree assertions rendered from the live-captured results by the extension's real overviewHtml/buildTree code (VS Code itself is never launched on this machine, pe…
Adversarial: the ask refuses a verify with no claim, a claim on an ask that takes none, a claim belonging to another part, a selection whose text is not on those head lines, and a selection over the 5… ✅ pass live live-checks.txt 'verify without a claim…', 'a claim on an ask that takes none…', "a claim of another part is refused plainly…", 'a selection not on those lines…', 'a selection over the cap…'
What covers this? on a part its tests exercise: the answer lists the covering test line re-read in the head copy as a cited line, and reports the manual check the description carries with its descript… ✅ pass live live-checks.txt 'cover (fresh part)…' checks citing tests/test_fresh.py:4 and the description's manual check at line 3; rendered in rendered-overview-asks.html
What covers this? on a part nothing covers: 'None found' is a valid answer with its own section heading, where the agent looked, and no citations ✅ pass live live-checks.txt 'cover (uncovered part) says none found…'; rendered in rendered-overview-asks.html
Both prompts land with their cases and scores: prompts.json registers cover v1 and verdicts v5 at the engine's own versions, and the evaluation's real case loaders read verify selections (library fetc… ✅ pass live eval-cases-checks.txt, produced by driving the evaluation package's loadRegistry/loadCase over prompts.json and every case folder; scoring machinery covered by packages/evaluation tests
  • npx vitest run packages/engine/test/verify.test.ts packages/engine/test/cover.test.ts packages/engine/test/review.test.ts packages/engine/test/explain.test.ts packages/engine/test/cli.test.ts (69 tests)
  • npx vitest run packages/engine/test/server.test.ts (35 tests)
  • npx vitest run packages/extension/test/overview.test.ts packages/extension/test/asked-claim.test.ts packages/extension/test/tree.test.ts packages/extension/test/review-result.test.ts packages/extension/test/engine-client.test.ts packages/extension/test/findings.test.ts (185 tests)
  • npx vitest run packages/evaluation/test/run.test.ts packages/evaluation/test/score.test.ts (70 tests)
  • node scratch-live/drive.mjs — spawned packages/engine/dist/main.js serve and drove initialize, three review requests, five verify asks, two cover asks, two fetchLibrary presses and five refusal asks over the real protocol (34 checks, evidence live-checks.txt + live-transcript.jsonrpc.txt)
  • npx vitest run packages/extension/test/live-render.test.ts (scratch harness, deleted after) — rendered the extension's overviewHtml and buildTree from the live-captured review results and ask answers (4 checks)
  • node scratch-live/eval-cases.mjs — drove the evaluation's loadRegistry and loadCase over prompts.json and all case folders (7 checks, evidence eval-cases-checks.txt)
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

Two more asks join the part's context menu, each defined in the ASKS
registry, which now also says whether an ask checks a claim the
reviewer picks or selects.

Verify this claim runs the judging pass on one claim: the text the
reviewer selected on the head side of the part's diff, or else one of
the part's claims they pick. A selection must sit on head-side lines the
part shows and be on those lines of the head copy as selected; it becomes
a claim of a new source, the reviewer. The verdicts prompt judges it
alone, its quote fenced as untrusted, every citation is re-read, and a
claim that needs a pinned or named library offers its library fetch,
which downloads nothing until pressed. The judged claim joins the
engine's latest review and the review shown, so its finding and fetch
are there to press. The verdicts prompt is now version 5: it names a
reviewer's selection as where its claim is made.

What covers this? is a new prompt, cover: the agent lists the tests, in
the change or the head copy, that exercise the part, and the manual
checks the description reports for it, or says none were found, which
is a valid answer. Every test line is re-read in the head copy and every
manual check found in the description before the answer is shown.

Both land with cases and scores: four labelled selections on
canary-python, canary-csharp and sindresorhus-ky-880 scored by
verify-accuracy, verify-false-verified and verify-fetch-offered, and
seven labelled parts on five cases, three of them covered by nothing,
scored by cover-cites-checked, cover-tests-recall,
cover-tests-precision, cover-manual-recall and cover-none-found. The
baseline records Pi 0.86.1 with zai-coding-cn/glm-5.3: verify-accuracy
0.75, verify-false-verified 0, verify-fetch-offered 1; cover-none-found
1, cover-tests-recall 0.8, every other cover score 1.

The review result is now version 17.
…/extension/test/integration/extension.test.ts ('verifies a claim the reviewer picks...' and 'verifies the text selected...') that asserted the verify-ask request params with toEqual but omitted the agent field. The extension always sends reviewAgentChoice(readAgentSettings()) with every engine request (extension.ts:761; the engine's ask handler consumes it at server.ts:519-523), so with the default stub settings the params include agent: {agent: 'pi', model: '', account: ''} — the extra field shown in the CI diff. Fixed the two test expectations to include the default agent choice; no production code changed. Siblings swept: the explain-ask test uses toMatchObject (unaffected) and the engine-client unit test passes agent undefined so omitting the key remains correct. Verified locally: npm run test:integration 54/54 (both previously failing tests pass), npm run build, tsc -p tsconfig.test.json, npm run lint, npm test 1325/1325 all green
@lbildzinkas
lbildzinkas merged commit 0da91a3 into master Oct 7, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Asks: verify this claim, and what covers this?

1 participant