Repository navigation
feat(engine): map acceptance criteria to code, tests and manual checks - #86
Merged
Merged
Conversation
…at show it An agent pass, last after the claims are judged, gives each acceptance criterion read from the linked issues a verdict: met, partly met, not met, can't tell or needs manual check, with its reason, the lines of code that implement it and of the tests that cover it, and the manual checks the description reports. The engine re-reads every citation in the head copy and finds every manual check in the description; one that does not match makes the criterion can't tell, as does a met or partly met verdict that shows nothing. Issue and description text reach the agent only inside untrusted blocks. The review result (version 15) carries each criterion's verdict and the criteria's mapping, and the overview lists them with their counts, each cited line a button that opens it in the head copy. The criteria-mapping prompt lands with its cases, two planted cases whose linked issues ask for every kind of verdict, one with a manual check its description reports, and planted-typescript's two criteria, and its scores: accuracy and false-met against the hand verdicts, and the recall of the labelled code, tests and manual checks.
Pi 0.86.1 with zai-coding-cn/glm-5.3 at its default effort, over the prompt's three cases: criteria-accuracy 1 over fourteen criteria, criteria-false-met 0, criteria-code-recall 0.89, and criteria-tests-recall and criteria-manual-recall 1. The plain rows are rewritten from a full model-free run; every other agent row stays.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Closes #2.
For each acceptance criterion read from the issues a pull request links, the companion shows whether the change meets it and where. Criteria are proven by code and automated tests, and also by manual checks that a person performed and the pull request reports, so those count as evidence too. Reading and listing the criteria was added earlier; this change adds the mapping and the verdicts.
The acceptance criteria are:
Validation installs nothing on the machine and never launches the VS Code application. Extension tests that need a real VS Code run only in CI; local validation goes through the engine protocol and unit tests.
What Changed
criteria-python,criteria-typescript), scored the mapped verdicts and evidence against hand labels, and refreshed the evaluation baselineRisk Assessment
Testing
Drove the change's full surface live: the engine's review command and its JSON-RPC serve process ran as real processes against a disposable recorded-GitHub pull request and a scripted agent, exercising the happy path (all five verdict kinds, re-checked code/test/manual-check evidence kept for a met criterion), the guards (citation off by one, manual check absent from the description, met with no evidence — each dropped to can't tell with the reason named), and the retry-then-fall-back path; the extension's overview page was rendered from that live result and its delivered script executed, proving the criteria sit at the top of the panel with linked code, tests and manual checks, including the human-decided manual-check link that jumps to the description. The evaluation's two new criteria cases and five criteria scores run under the offline CI evaluation with the baseline holding. No screenshot was taken: no browser is on PATH and the intent forbids launching VS Code, so the rendered webview HTML artifact plus the executed page script serve as the visual evidence.
Evidence: Rendered overview page (extension webview HTML) generated from the live engine result
Evidence: Engine review command's printed JSON result (live run, criteria mapped)
~/.no-mistakes/evidence/01M4876EJJR6C1B6R8HTFG3E21/engine-serve-protocol-transcript.jsonl)Evidence: Stage announcements of the live CLI review run
second-look-engine: plain parts ready; grouping related hunks with pi second-look-engine: plain parts ready; ranking the parts with pi second-look-engine: plain parts ready; writing the story with pi second-look-engine: plain parts ready; comparing the change with its description and issues with pi second-look-engine: plain parts ready; listing the claims with pi second-look-engine: plain parts ready; mapping the acceptance criteria with piEvidence: Live adversarial run: invalid mapping answer retried once, then fell back with every criterion not checked
Evidence: Offline CI evaluation against the stored baseline, including the new criteria cases
baseline: 0 dropped, 0 missing, 0 gained, 95 unchanged, 0 without a baseline, 0 unstamped and not comparedPipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
packages/extension/src/overview.ts:386- The ticket's acceptance criterion 2 reads "Each criterion links to the code that implements it, the automated tests that cover it, and any manual check the pull request reports." The implementation links code and tests (buttons that open the cited head-copy line via openHeadLine, extension.ts:449-458), but a manual check is rendered only as an inline quote plus plain text "description, line N" (overview.ts:386) with no link to jump to the description section or open the pull request description at that line. The choice is consistent with the overview's existing precedent for description-line references (describedWhere, overview.ts:484, is also plain text) and with the README's own wording ("each quoted from the description"), so this may be a deliberate reading that quoting the check satisfies the criterion; but since the quoted criterion literally says "links to ... any manual check", a human should confirm. If a link is required, the smallest remedy is making the manual-check quote (or its line reference) a link to the description section/line already rendered on the same page.packages/engine/src/criteria-mapping.ts:195- describePart (criteria-mapping.ts:195-199) is a near-verbatim copy of the part-rendering block in verdictsPrompt (packages/engine/src/verdicts.ts:276-281) — identical except the noise note's noun ("a criterion needs them" vs "a claim needs them") — and oneLine (criteria-mapping.ts:261) is a third private copy of the same helper (claims.ts:240, verdicts.ts:314). A shared part-description helper parameterised by the noun (and a shared oneLine) would keep the agent-facing part-presentation rule in one place and avoid drift between the two prompts. Mechanical, non-functional refactor.🔧 Fix applied.
1 info still open:
packages/engine/src/criteria-mapping.ts:195- describePart (criteria-mapping.ts:195-199) is a near-verbatim copy of the part-rendering block in verdictsPrompt (packages/engine/src/verdicts.ts:276-281) — identical except the noise note's noun ("a criterion needs them" vs "a claim needs them") — and oneLine (criteria-mapping.ts:261) is a third private copy of the same helper (claims.ts:240, verdicts.ts:314). A shared part-description helper parameterised by the noun (and a shared oneLine) would keep the agent-facing part-presentation rule in one place and avoid drift between the two prompts. Mechanical, non-functional refactor.✅ **Test** - passed
✅ No issues found.
npm run buildLive CLI run: node --import <hook> packages/engine/dist/main.js review https://github.com/example-org/example-repo/pull/9 --token test-token --agent pi against a disposable recorded-GitHub fixture (.live-validation, removed after) with a scripted fake pi on PATH — exit 0, result v15, criteria mapped (evidence: engine-review-cli-result.json, engine-review-cli-stages.txt)Live adversarial CLI run with FAKE_PI_MODE=fallback: mapping answer invalid twice — retried once with problems named, then fell back, every criterion not checked, review still succeeds (evidence: engine-review-fallback.txt)Live protocol run: spawned packages/engine/dist/main.js serve over real stdio, handshake + review request carrying the agent choice — review/stage notifications ending with 'mapping the acceptance criteria with pi', final response carrying the mapped criteria and the request's account label (evidence: engine-serve-protocol-transcript.jsonl)Rendered the extension's real overviewHtml from the live engine result through the extension protocol guard, wrote it as evidence, and executed the delivered page script: manual-check button scrolls the description section into view, cite button posts openEvidence (temporary vitest driver, removed after; evidence: overview-live.html)npx vitest run packages/engine/test/criteria-mapping.test.ts packages/engine/test/review.test.ts packages/engine/test/server.test.ts packages/engine/test/cli.test.ts packages/engine/test/criteria.test.ts — 85 passednpx vitest run packages/extension/test/overview.test.ts packages/extension/test/review-result.test.ts packages/extension/test/engine-client.test.ts packages/evaluation/test/score.test.ts packages/evaluation/test/run.test.ts packages/evaluation/test/cli.test.ts — 175 passednpm run eval (offline, model-free, as CI runs it) — baseline: 0 dropped, 0 missing, 0 gained, 95 unchanged (evidence: evaluation-baseline-check.txt)✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.