Skip to content

feat(engine): map acceptance criteria to code, tests and manual checks - #86

Merged
lbildzinkas merged 5 commits into
masterfrom
fm/second-look-criteria-mapping
Oct 6, 2026
Merged

lbildzinkas merged 5 commits into
masterfrom
fm/second-look-criteria-mapping

Conversation

@lbildzinkas

@lbildzinkas lbildzinkas commented Oct 6, 2026 •

Copy link
Copy Markdown
Owner

Intent

Closes #2.

For each acceptance criterion read from the issues a pull request links, the companion shows whether the change meets it and where. Criteria are proven by code and automated tests, and also by manual checks that a person performed and the pull request reports, so those count as evidence too. Reading and listing the criteria was added earlier; this change adds the mapping and the verdicts.

The acceptance criteria are:

  • Each criterion shows its quoted text and a verdict: met, partly met, not met, can't tell, or needs manual check.
  • Each criterion links to the code that implements it, the automated tests that cover it, and any manual check the pull request reports.
  • The criteria and their verdicts appear in the overview panel at the top of the review, not only as comments.
  • The mapping prompt lands with its own evaluation cases and score, as ADR 0006 requires.

Validation installs nothing on the machine and never launches the VS Code application. Extension tests that need a real VS Code run only in CI; local validation goes through the engine protocol and unit tests.

What Changed

  • Added a criteria-mapping pass to the engine, run as the last review stage (review-result protocol v15): each acceptance criterion read from the linked issues gets a verdict — met, partly met, not met, can't tell or needs manual check — with a reason, code and test citations the engine re-reads in the head copy, and the manual checks the PR description reports, each quote verified there; a citation or quote that does not match drops the verdict to can't tell
  • The extension's overview panel now shows each criterion with its verdict at the top of the review: cited code and test lines open in the head copy, manual checks jump to the description, unmet verdicts are highlighted as findings, and a criteria stage chip reports the mapping outcome
  • Registered the versioned criteria-mapping prompt (ADR 0006) with its own evaluation cases (criteria-python, criteria-typescript), scored the mapped verdicts and evidence against hand labels, and refreshed the evaluation baseline

Risk Assessment

⚠️ Medium: The new stage is well-built and follows the repo's established patterns (untrusted blocks, engine-side re-checks, retry-then-fallback, versioned prompt with evaluation and baseline), but one ticket criterion ("links to ... any manual check") is only arguably satisfied since manual checks are quoted rather than linked, which needs a human decision before merge.

Testing

Drove the change's full surface live: the engine's review command and its JSON-RPC serve process ran as real processes against a disposable recorded-GitHub pull request and a scripted agent, exercising the happy path (all five verdict kinds, re-checked code/test/manual-check evidence kept for a met criterion), the guards (citation off by one, manual check absent from the description, met with no evidence — each dropped to can't tell with the reason named), and the retry-then-fall-back path; the extension's overview page was rendered from that live result and its delivered script executed, proving the criteria sit at the top of the panel with linked code, tests and manual checks, including the human-decided manual-check link that jumps to the description. The evaluation's two new criteria cases and five criteria scores run under the offline CI evaluation with the baseline holding. No screenshot was taken: no browser is on PATH and the intent forbids launching VS Code, so the rendered webview HTML artifact plus the executed page script serve as the visual evidence.

  • Live validation: ✅ go - 7 of 7 scenarios driven live against the product
Scenario Result Live Evidence
Reviewer runs a review of a PR whose linked issue holds an acceptance-criteria checklist: each criterion is shown with its quoted text and one of the five verdicts, and a met verdict keeps the code, t… ✅ pass live engine-review-cli-result.json (live CLI run: six criteria → met, can't tell ×3, not met, needs manual check; a1 keeps app/send.py:5, tests/test_send.py:15 and the description's manual check at line 5)
Agent cites a line one off, quotes a manual check the description does not hold, and calls a criterion met with no evidence: the engine drops each verdict to can't tell and says why, keeping the revie… ✅ pass live engine-review-cli-result.json — a2 recheck: 'the quote of the citation app/send.py:6 is not on that line; the manual check "Tested on production" is not in the description'; a3 recheck: 'the verdict c…
Agent's mapping answer is invalid: the engine retries once with the problems named, then falls back with every criterion not checked and the mapping outcome saying why, without failing the review ✅ pass live engine-review-fallback.txt (exit 0, outcome 'fell back', all six criteria 'not checked')
Extension asks the engine to review over the JSON-RPC protocol: the criteria-mapping stage is announced while it runs and the final response carries the version-15 result with the criteria mapped and… ✅ pass live engine-serve-protocol-transcript.jsonl (real serve process on stdio; last stage 'mapping the acceptance criteria with pi' shows the criteria still not checked; final id=2 response carries mapped crite…
Reviewer opens the overview panel: the criteria and their verdicts appear at the top of the review, right after the story, with the verdict counts beside the heading and each criterion's reason and ev… ✅ pass live overview-live.html — the extension's real overviewHtml rendered from the live engine result, accepted by the extension's protocol guard; section order story → criteria → unexplained → claims; heading…
Each criterion links to the code that implements it and the tests that cover it (buttons that open the cited head-copy line) and to any manual check the PR reports (a link the reviewer follows to the… ✅ pass live overview-live.html (buttons data-criterion/data-evidence/data-index and button.manual 'description, line 5'); the delivered page script was executed and a manual-check click scrolled the description s…
The criteria-mapping prompt lands with its own evaluation cases and score (ADR 0006): versioned prompt registry entry, hand-labelled cases, its own score set, and the CI evaluation comparing them to t… ✅ pass live evaluation-baseline-check.txt — npm run eval (the real evaluation CLI, offline as CI runs it) over the new criteria-python and criteria-typescript cases: 0 dropped, 0 missing, 0 gained, 95 unchanged;…
Evidence: Rendered overview page (extension webview HTML) generated from the live engine result
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8">
<meta http-equiv="Content-Security-Policy" content="default-src 'none'; style-src 'nonce-liveValidationNonce'; script-src 'nonce-liveValidationNonce';">
<style nonce="liveValidationNonce">
  body {
    color: var(--vscode-foreground);
    background-color: var(--vscode-editor-background);
    font-family: var(--vscode-font-family);
    font-size: var(--vscode-font-size, 13px);
    margin: 0;
    padding: 0 26px;
  }
  main { max-width: 900px; padding: 18px 0 48px; }
  h1 { font-size: 20px; font-weight: 600; margin: 0 0 4px; }
  h2 { font-size: 15px; font-weight: 600; margin: 16px 0 8px; display: flex; align-items: center; gap: 8px; flex-wrap: wrap; }
  .meta, .note { color: var(--vscode-descriptionForeground); font-size: 12px; }
  .note { margin: 0 0 6px; }
  .stages { display: flex; gap: 6px; flex-wrap: wrap; margin: 10px 0 4px; }
  .stg { font-size: 12px; border: 1px solid var(--vscode-panel-border); border-radius: 12px; padding: 1px 9px; }
  .stg.done::before { content: "✓ "; color: var(--vscode-testing-iconPassed, #89d185); }
  .stg.run { color: var(--vscode-descriptionForeground); }
  .stamp { font-size: 11px; font-weight: 400; color: var(--vscode-descriptionForeground); }
  .story {
    border: 1px solid var(--vscode-panel-border);
    border-left: 3px solid var(--vscode-textLink-foreground);
    border-radius: 3px;
    padding: 10px 12px;
    font-size: 13.5px;
    line-height: 1.55;
  }
  .sentence.focus { background-color: var(--vscode-editor-findMatchHighlightBackground); }
  .pt {
    font: inherit;
    color: var(--vscode-textLink-foreground);
    background: none;
    border: none;
    border-bottom: 1px dotted var(--vscode-textLink-foreground);
    padding: 0;
    cursor: pointer;
  }
  .pt.focus { font-weight: 600; }
  .issue { font-style: italic; }
  code, .shown { font-family: var(--vscode-editor-font-family, monospace); font-size: 12px; }
  .description {
    white-space: pre-wrap;
    overflow-wrap: anywhere;
    border: 1px solid var(--vscode-panel-border);
    border-radius: 3px;
    padding: 10px 12px;
  }
  .alert {
    border-left: 3px solid var(--vscode-editorWarning-foreground);
    padding: 4px 10px;
    margin: 0 0 8px;
  }
  .hidden {
    border: 1px dashed var(--vscode-editorWarning-foreground);
    border-radius: 3px;
    padding: 0 4px;
  }
  .flag {
    color: var(--vscode-editorWarning-foreground);
    font-size: 11px;
    font-weight: 600;
    margin-right: 6px;
  }
  .claims { padding-left: 22px; margin: 0; }
  .claims li { margin-bottom: 8px; }
  .quote { overflow-wrap: anywhere; }
  .where { color: var(--vscode-descriptionForeground); font-size: 12px; margin-top: 2px; }
  .verdict { font-style: italic; }
  .verdict.finding { color: var(--vscode-editorWarning-foreground); font-weight: 600; }
  .why { color: var(--vscode-descriptionForeground); font-size: 12px; overflow-wrap: anywhere; }
  .evidence { display: grid; grid-template-columns: max-content minmax(0, 1fr); gap: 2px 10px; font-size: 12px; margin-top: 2px; }
  .evidence .label, .cited { color: var(--vscode-descriptionForeground); }
  .evidence span { min-width: 0; overflow-wrap: anywhere; }
  .cite { font-family: var(--vscode-editor-font-family, monospace); font-size: 12px; }
  .none { color: var(--vscode-editorWarning-foreground); }
  .findings, .checks { padding-left: 18px; margin: 0; }
  .findings li, .checks li { margin-bottom: 6px; overflow-wrap: anywhere; }
  .att, .sev, .check { font-size: 11px; font-weight: 600; border: 1px solid var(--vscode-panel-border); border-radius: 10px; padding: 0 7px; }
  .att.stale, .att.malformed, .sev.error, .sev.warning, .check.failed { color: var(--vscode-editorWarning-foreground); }
  .att.fresh, .check.passed { color: var(--vscode-testing-iconPassed, #89d185); }
  .log {
    font-family: var(--vscode-editor-font-family, monospace);
    font-size: 12px;
    white-space: pre-wrap;
    overflow-wrap: anywhere;
    border: 1px solid var(--vscode-panel-border);
    border-radius: 3px;
    padding: 6px 10px;
    margin: 4px 0;
    max-height: 320px;
    overflow-y: auto;
  }
  .stamps { padding-left: 18px; margin: 0; }
  .stamps li { margin-bottom: 4px; }
</style>
</head>
<body>
<main>
  <h1>Retry failed webhook sends</h1>
  <div class="meta">example-org/example-repo #9 · contributor-login · retry-send → master · head abcdabc</div>
  <div class="stages"><span class="stg done">parts</span><span class="stg done">noise checks</span><span class="stg done">plain grouping kept</span><span class="stg done">plain ranking kept</span><span class="stg done">no story</span><span class="stg done">no comparison</span><span class="stg done">no claims</span><span class="stg done">criteria mapped</span></div>
  <section id="story"><h2>Story <span class="stamp">pi · fake-provider/glm-5.3 · story prompt v1</span></h2><p class="note">No story: the agent gave no usable answer (invalid-answer: the answer was invalid twice: the answer is missing &quot;sentences&quot;).</p></section>
  <section id="criteria"><h2>Acceptance criteria <span class="stamp">1 not met · 1 needs manual check · 3 can&#39;t tell · 1 met</span> <span class="stamp">pi · fake-provider/glm-5.3 · criteria-mapping prompt v1</span></h2><p class="note">1 issue this pull request closes.</p><p class="note">Each condition listed in the issues this pull request links, quoted from the checklist under &quot;Acceptance criteria&quot;. Issue text is untrusted: its hidden content is shown and flagged. Each is judged against the change, its read-only copy and the manual checks the description reports, by pi · fake-provider/glm-5.3 · criteria-mapping prompt v1; the not met and partly met ones are findings.</p><ol class="claims criteria"><li><q class="quote">A send that fails is retried three times.</q><div class="where"><button type="button" class="pt issue" data-issue="0">#30 in example-org/example-repo</button> · closes · <span class="verdict">met</span></div><div class="why">The loop retries a failed send three times.</div><div class="evidence"><span class="label">Code</span><span><button type="button" class="pt cite" data-criterion="0" data-evidence="code" data-index="0">app/send.py:5</button> <span class="cited">for attempt in range(3):</span></span><span class="label">Tests</span><span><button type="button" class="pt cite" data-criterion="0" data-evidence="tests" data-index="0">tests/test_send.py:15</button> <span class="cited">assert Failing.calls == 3</span></span><span class="label">Manual check</span><span><q class="quote">Tried by hand against the staging endpoint: the third retry gave up and the send was dropped after three tries.</q> <button type="button" class="pt manual">description, line 5</button></span></div></li><li><q class="quote">Each retry waits one second before it starts.</q><div class="where"><button type="button" class="pt issue" data-issue="0">#30 in example-org/example-repo</button> · closes · <span class="verdict">can&#39;t tell</span></div><div class="why">Each failed try sleeps one second.</div><div class="evidence"><span class="label">Code</span><span><span class="none">none</span></span><span class="label">Tests</span><span><span class="none">none</span></span><span class="label">Manual check</span><span><span class="cited">none reported in the pull request</span></span></div><div class="why">dropped to can&#39;t tell: the quote of the citation app/send.py:6 is not on that line; the manual check &quot;Tested on production&quot; is not in the description</div></li><li><q class="quote">The retries are logged with the reason they were needed.</q><div class="where"><button type="button" class="pt issue" data-issue="0">#30 in example-org/example-repo</button> · closes · <span class="verdict">can&#39;t tell</span></div><div class="why">The retries are logged.</div><div class="evidence"><span class="label">Code</span><span><span class="none">none</span></span><span class="label">Tests</span><span><span class="none">none</span></span><span class="label">Manual check</span><span><span class="cited">none reported in the pull request</span></span></div><div class="why">dropped to can&#39;t tell: the verdict cites no code, test or manual check</div></li><li><q class="quote">The retry backoff is capped at one minute.</q><div class="where"><button type="button" class="pt issue" data-issue="0">#30 in example-org/example-repo</button> · closes · <span class="verdict finding">not met</span></div><div class="why">Nothing caps the wait between retries.</div><div class="evidence"><span class="label">Code</span><span><span class="none">none</span></span><span class="label">Tests</span><span><span class="none">none</span></span><span class="label">Manual check</span><span><span class="cited">none reported in the pull request</span></span></div></li><li><q class="quote">The send timeout follows the OS TCP settings.</q><div class="where"><button type="button" class="pt issue" data-issue="0">#30 in example-org/example-repo</button> · closes · <span class="verdict">can&#39;t tell</span></div><div class="why">The OS TCP settings are not part of the change.</div><div class="evidence"><span class="label">Code</span><span><span class="none">none</span></span><span class="label">Tests</span><span><span class="none">none</span></span><span class="label">Manual check</span><span><span class="cited">none reported in the pull request</span></span></div></li><li><q class="quote">The retry warning reads clearly on screen.</q><div class="where"><button type="button" class="pt issue" data-issue="0">#30 in example-org/example-repo</button> · closes · <span class="verdict">needs manual check</span></div><div class="why">How the warning reads on screen only a person trying the change can settle.</div><div class="evidence"><span class="label">Code</span><span><span class="none">none</span></span><span class="label">Tests</span><span><span class="none">none</span></span><span class="label">Manual check</span><span><span class="cited">none reported in the pull request</span></span></div></li></ol></section>
  <section id="unexplained"><h2>Unexplained changes <span class="stamp">pi · fake-provider/glm-5.3 · unexplained prompt v1</span></h2><p class="note">No comparison: the agent gave no usable answer (invalid-answer: the answer was invalid twice: the answer is missing &quot;unexplained&quot;; the answer is missing &quot;described&quot;).</p></section>
  <section id="claims"><h2>Claims <span class="stamp">pi · fake-provider/glm-5.3 · claims prompt v1</span></h2><p class="note">No claims: the agent gave no usable answer (invalid-answer: the answer was invalid twice: the answer is missing &quot;claims&quot;).</p></section>
  <section id="pipeline"><h2>Pipeline and CI</h2><p><span class="att missing">no-mistakes report: none</span> <span class="note">the description carries no no-mistakes attestation.</span></p><p class="note">Checks listed at head abcdabc, ran on merge commit feedfac: 0 check runs at the head commit, 0 failed; logs are read only for failed jobs.</p></section>
  <section id="description"><h2>Pull request description</h2><div class="description">Closes #30.

Retries a failed webhook send three times, waiting a second between tries.

Tried by hand against the staging endpoint: the third retry gave up and the send was dropped after three tries.</div></section>
  <section id="stamps"><h2>How these results were made</h2><ul class="stamps"><li><b>Parts</b> plain grouping kept: the agent gave no usable answer (invalid-answer: the answer was invalid twice: the answer is missing &quot;parts&quot;)</li><li><b>Ranking</b> plain ranking kept: the agent gave no usable answer (invalid-answer: the answer was invalid twice: the answer is missing &quot;parts&quot;)</li><li><b>Story</b> none: the agent gave no usable answer (invalid-answer: the answer was invalid twice: the answer is missing &quot;sentences&quot;)</li><li><b>Unexplained changes</b> none: the agent gave no usable answer (invalid-answer: the answer was invalid twice: the answer is missing &quot;unexplained&quot;; the answer is missing &quot;described&quot;)</li><li><b>Claims</b> none: the agent gave no usable answer (invalid-answer: the answer was invalid twice: the answer is missing &quot;claims&quot;)</li><li><b>Acceptance criteria</b> mapped by pi · fake-provider/glm-5.3 · criteria-mapping prompt v1: every citation was re-read in the head copy and every manual check found in the description; one that did not match made a criterion can&#39;t tell</li></ul></section>
</main>
<script nonce="liveValidationNonce">
(function () {
  'use strict';
  var vscode = acquireVsCodeApi();
  Array.prototype.forEach.call(document.querySelectorAll('button.pt[data-part]'), function (button) {
    button.addEventListener('click', function () {
      vscode.postMessage({ type: 'openPart', part: Number(button.getAttribute('data-part')) });
    });
  });
  Array.prototype.forEach.call(document.querySelectorAll('button.issue'), function (button) {
    button.addEventListener('click', function () {
      vscode.postMessage({ type: 'openIssue', issue: Number(button.getAttribute('data-issue')) });
    });
  });
  Array.prototype.forEach.call(document.querySelectorAll('button.cite'), function (button) {
    button.addEventListener('click', function () {
      vscode.postMessage({
        type: 'openEvidence',
        criterion: Number(button.getAttribute('data-criterion')),
        evidence: button.getAttribute('data-evidence'),
        index: Number(button.getAttribute('data-index'))
      });
    });
  });
  Array.prototype.forEach.call(document.querySelectorAll('button.manual'), function (button) {
    button.addEventListener('click', function () {
      var description = document.getElementById('description');
      if (description !== null) {
        description.scrollIntoView({ block: 'start' });
      }
    });
  });
  var focused = document.querySelector('.sentence.focus');
  if (focused !== null) {
    focused.scrollIntoView({ block: 'center' });
  }
}());
</script>
</body>
</html>
Evidence: Engine review command's printed JSON result (live run, criteria mapped)
{
  "version": 15,
  "pullRequest": {
    "url": "https://github.com/example-org/example-repo/pull/9",
    "number": 9,
    "title": "Retry failed webhook sends",
    "author": "contributor-login",
    "description": "Closes #30.\n\nRetries a failed webhook send three times, waiting a second between tries.\n\nTried by hand against the staging endpoint: the third retry gave up and the send was dropped after three tries.",
    "base": "master",
    "head": "retry-send",
    "baseCommit": "0fed0fed0fed0fed0fed0fed0fed0fed0fed0fed",
    "headSha": "abcdabcdabcdabcdabcdabcdabcdabcdabcdabcd"
  },
  "copies": {
    "base": {
      "commit": "1234123412341234123412341234123412341234",
      "path": "~/.no-mistakes/worktrees/a09930ebf9a5/01M4876EJJR6C1B6R8HTFG3E21/.live-validation/cache/github.com/example-org/example-repo/pull-9/1234123412341234123412341234123412341234",
      "reused": false
    },
    "head": {
      "commit": "abcdabcdabcdabcdabcdabcdabcdabcdabcdabcd",
      "path": "~/.no-mistakes/worktrees/a09930ebf9a5/01M4876EJJR6C1B6R8HTFG3E21/.live-validation/cache/github.com/example-org/example-repo/pull-9/abcdabcdabcdabcdabcdabcdabcdabcdabcdabcd",
      "reused": false
    }
  },
  "parseTimeMs": 2.731,
  "parts": [
    {
      "path": "tests/test_send.py",
      "changeKind": "addition",
      "isBinary": false,
      "oldMissingFinalNewline": false,
      "newMissingFinalNewline": false,
      "hunks": [
        {
          "oldStart": 0,
          "oldLines": 0,
          "newStart": 1,
          "newLines": 15,
          "lines": [
            {
              "kind": "addition",
              "newLineNumber": 1,
              "text": "from app.send import send"
            },
            {
              "kind": "addition",
              "newLineNumber": 2,
              "text": ""
            },
            {
              "kind": "addition",
              "newLineNumber": 3,
              "text": ""
            },
            {
              "kind": "addition",
              "newLineNumber": 4,
              "text": "class Failing:"
            },
            {
              "kind": "addition",
              "newLineNumber": 5,
              "text": "    calls = 0"
            },
            {
              "kind": "addition",
              "newLineNumber": 6,
              "text": ""
            },
            {
              "kind": "addition",
              "newLineNumber": 7,
              "text": "    def ok(self):"
            },
            {
              "kind": "addition",
              "newLineNumber": 8,
              "text": "        Failing.calls += 1"
            },
            {
              "kind": "addition",
              "newLineNumber": 9,
              "text": "        return False"
            },
            {
              "kind": "addition",
              "newLineNumber": 10,
              "text": ""
            },
            {
              "kind": "addition",
              "newLineNumber": 11,
              "text": ""
            },
            {
              "kind": "addition",
              "newLineNumber": 12,
              "text": "def test_send_retries_three_times():"
            },
            {
              "kind": "addition",
              "newLineNumber": 13,
              "text": "    Failing.calls = 0"
            },
            {
              "kind": "addition",
              "newLineNumber": 14,
              "text": "    assert not send(Failing())"
            },
            {
              "kind": "addition",
              "newLineNumber": 15,
              "text": "    assert Failing.calls == 3"
            }
          ],
          "entities": [
            {
              "kind": "class",
              "name": "Failing",
              "public": true,
              "change": "added"
            },
            {
              "kind": "method",
              "name": "Failing.ok",
              "public": true,
              "change": "added"
            },
            {
              "kind": "function",
              "name": "test_send_retries_three_times",
              "public": true,
              "change": "added"
            }
          ]
        }
      ],
      "additions": 15,
      "deletions": 0,
      "syntax": {
        "language": "python",
        "formattingOnly": {
          "status": "structure-changed",
          "reason": "the file is new"
        },
        "checksNotRun": []
      },
      "newMode": "100644",
      "noise": {
        "label": "none",
        "note": "no rule applied"
      },
      "name": "Failing, Failing.ok, test_send_retries_three_times in tests/test_send.py",
      "origin": "plain",
      "signals": {
        "novelty": "new",
        "role": "test",
        "changedLines": 15,
        "publicSurface": [
          "Failing",
          "Failing.ok",
          "test_send_retries_three_times"
        ],
        "references": {
          "basis": "name-based",
          "names": [
            "Failing",
            "ok",
            "test_send_retries_three_times"
          ],
          "files": 1
        }
      },
      "rank": {
        "importance": "worth reviewing",
        "reason": "test; new code; 15 changed lines",
        "signals": [
          "test",
          "new code",
          "15 changed lines"
        ]
      }
    },
    {
      "path": "app/send.py",
      "changeKind": "modification",
      "isBinary": false,
      "oldMissingFinalNewline": false,
      "newMissingFinalNewline": false,
      "hunks": [
        {
          "oldStart": 1,
          "oldLines": 4,
          "newStart": 1,
          "newLines": 9,
          "lines": [
            {
              "kind": "addition",
              "newLineNumber": 1,
              "text": "import time"
            },
            {
              "kind": "addition",
              "newLineNumber": 2,
              "text": ""
            },
            {
              "kind": "addition",
              "newLineNumber": 3,
              "text": ""
            },
            {
              "kind": "context",
              "oldLineNumber": 1,
              "newLineNumber": 4,
              "text": "def send(request):"
            },
            {
              "kind": "deletion",
              "oldLineNumber": 2,
              "text": "    if request.ok():"
            },
            {
              "kind": "deletion",
              "oldLineNumber": 3,
              "text": "        return True"
            },
            {
              "kind": "addition",
              "newLineNumber": 5,
              "text": "    for attempt in range(3):"
            },
            {
              "kind": "addition",
              "newLineNumber": 6,
              "text": "        if request.ok():"
            },
            {
              "kind": "addition",
              "newLineNumber": 7,
              "text": "            return True"
            },
            {
              "kind": "addition",
              "newLineNumber": 8,
              "text": "        time.sleep(1)"
            },
            {
              "kind": "context",
              "oldLineNumber": 4,
              "newLineNumber": 9,
              "text": "    return False"
            }
          ],
          "entities": [
            {
              "kind": "function",
              "name": "send",
              "public": true,
              "change": "body"
            }
          ]
        }
      ],
      "additions": 7,
      "deletions": 2,
      "syntax": {
        "language": "python",
        "formattingOnly": {
          "status": "structure-changed",
          "reason": "the syntax tree changes at head line 1"
        },
        "checksNotRun": []
      },
      "noise": {
        "label": "none",
        "note": "no rule applied"
      },
      "name": "send in app/send.py",
      "origin": "plain",
      "signals": {
        "novelty": "changed",
        "role": "code",
        "changedLines": 9,
        "publicSurface": [],
        "references": {
          "basis": "name-based",
          "names": [
            "send"
          ],
          "files": 1
        }
      },
      "rank": {
        "importance": "worth reviewing",
        "reason": "code; changes code named in 1 other file (name-based); 9 changed lines",
        "signals": [
          "code",
          "changes code named in 1 other file (name-based)",
          "9 changed lines"
        ]
      }
    }
  ],
  "grouping": {
    "by": "plain",
    "agent": {
      "promptVersion": "2",
      "stamp": {
        "agent": "pi",
        "agentVersion": "9.9.9",
        "model": "fake-provider/glm-5.3",
        "effort": null,
        "runAt": "2026-10-06T09:40:56.778Z",
        "tokens": {
          "input": 4200,
          "output": 360,
          "cacheRead": 0,
          "cacheWrite": 0,
          "total": 4560
        },
        "costUsd": 0.04
      },
      "outcome": "fell back",
      "detail": "the agent gave no usable answer (invalid-answer: the answer was invalid twice: the answer is missing \"parts\")",
      "leftOut": 0
    }
  },
  "ranking": {
    "by": "plain",
    "agent": {
      "promptVersion": "1",
      "stamp": {
        "agent": "pi",
        "agentVersion": "9.9.9",
        "model": "fake-provider/glm-5.3",
        "effort": null,
        "runAt": "2026-10-06T09:40:56.944Z",
        "tokens": {
          "input": 4200,
          "output": 360,
          "cacheRead": 0,
          "cacheWrite": 0,
          "total": 4560
        },
        "costUsd": 0.04
      },
      "outcome": "fell back",
      "detail": "the agent gave no usable answer (invalid-answer: the answer was invalid twice: the answer is missing \"parts\")"
    }
  },
  "criteria": {
    "outcome": "read",
    "detail": "1 issue this pull request closes",
    "heading": "Acceptance criteria",
    "issues": [
      {
        "number": 30,
        "title": "Retry failed webhook sends",
        "url": "https://github.com/example-org/example-repo/issues/30",
        "repository": "example-org/example-repo",
        "body": "The webhook sender gives up on the first failure.\n\n## Acceptance criteria\n\n- [ ] A send that fails is retried three times.\n- [x] Each retry waits one second before it starts.\n- [ ] The retries are logged with the reason they were needed.\n- [ ] The retry backoff is capped at one minute.\n- [ ] The send timeout follows the OS TCP settings.\n- [ ] The retry warning reads clearly on screen.\n\nSome notes after the checklist.\n",
        "link": "closes"
      }
    ],
    "criteria": [
      {
        "quote": "A send that fails is retried three times.",
        "issue": 0,
        "line": 5,
        "verdict": {
          "kind": "met",
          "reason": "The loop retries a failed send three times.",
          "code": [
            {
              "path": "app/send.py",
              "line": 5,
              "quote": "for attempt in range(3):"
            }
          ],
          "tests": [
            {
              "path": "tests/test_send.py",
              "line": 15,
              "quote": "assert Failing.calls == 3"
            }
          ],
          "manualChecks": [
            {
              "quote": "Tried by hand against the staging endpoint: the third retry gave up and the send was dropped after three tries.",
              "line": 5
            }
          ]
        }
      },
      {
        "quote": "Each retry waits one second before it starts.",
        "issue": 0,
        "line": 6,
        "verdict": {
          "kind": "can't tell",
          "reason": "Each failed try sleeps one second.",
          "code": [],
          "tests": [],
          "manualChecks": [],
          "recheck": "the quote of the citation app/send.py:6 is not on that line; the manual check \"Tested on production\" is not in the description"
        }
      },
      {
        "quote": "The retries are logged with the reason they were needed.",
        "issue": 0,
        "line": 7,
        "verdict": {
          "kind": "can't tell",
          "reason": "The retries are logged.",
          "code": [],
          "tests": [],
          "manualChecks": [],
          "recheck": "the verdict cites no code, test or manual check"
        }
      },
      {
        "quote": "The retry backoff is capped at one minute.",
        "issue": 0,
        "line": 8,
        "verdict": {
          "kind": "not met",
          "reason": "Nothing caps the wait between retries.",
          "code": [],
          "tests": [],
          "manualChecks": []
        }
      },
      {
        "quote": "The send timeout follows the OS TCP settings.",
        "issue": 0,
        "line": 9,
        "verdict": {
          "kind": "can't tell",
          "reason": "The OS TCP settings are not part of the change.",
          "code": [],
          "tests": [],
          "manualChecks": []
        }
      },
      {
        "quote": "The retry warning reads clearly on screen.",
        "issue": 0,
        "line": 10,
        "verdict": {
          "kind": "needs manual check",
          "reason": "How the warning reads on screen only a person trying the change can settle.",
          "code": [],
          "tests": [],
          "manualChecks": []
        }
      }
    ],
    "mapping": {
      "promptVersion": "1",
      "stamp": {
        "agent": "pi",
        "agentVersion": "9.9.9",
        "model": "fake-provider/glm-5.3",
        "effort": null,
        "runAt": "2026-10-06T09:40:57.606Z",
        "tokens": {
          "input": 2100,
          "output": 180,
          "cacheRead": 0,
          "cacheWrite": 0,
          "total": 2280
        },
        "costUsd": 0.02
      },
      "outcome": "mapped",
      "detail": "every citation was re-read in the head copy and every manual check found in the description; one that did not match made a criterion can't tell"
    }
  },
  "pipeline": {
    "attestation": "missing",
    "detail": "the description carries no no-mistakes attestation",
    "steps": [],
    "findings": []
  },
  "ci": {
    "headSha": "abcdabcdabcdabcdabcdabcdabcdabcdabcdabcd",
    "mergeCommit": "feedfacefeedfacefeedfacefeedfacefeedface",
    "outcome": "read",
    "detail": "0 check runs at the head commit, 0 failed; logs are read only for failed jobs",
    "checks": []
  },
  "story": {
    "promptVersion": "1",
    "stamp": {
      "agent": "pi",
      "agentVersion": "9.9.9",
      "model": "fake-provider/glm-5.3",
      "effort": null,
      "runAt": "2026-10-06T09:40:57.108Z",
      "tokens": {
        "input": 4200,
        "output": 360,
        "cacheRead": 0,
        "cacheWrite": 0,
        "total": 4560
      },
      "costUsd": 0.04
    },
    "outcome": "fell back",
    "detail": "the agent gave no usable answer (invalid-answer: the answer was invalid twice: the answer is missing \"sentences\")",
    "sentences": []
  },
  "unexplained": {
    "promptVersion": "1",
    "parts": [],
    "described": [],
    "outcome": "fell back",
    "detail": "the agent gave no usable answer (invalid-answer: the answer was invalid twice: the answer is missing \"unexplained\"; the answer is missing \"described\")",
    "stamp": {
      "agent": "pi",
      "agentVersion": "9.9.9",
      "model": "fake-provider/glm-5.3",
      "effort": null,
      "runAt": "2026-10-06T09:40:57.271Z",
      "tokens": {
        "input": 4200,
        "output": 360,
        "cacheRead": 0,
        "cacheWrite": 0,
        "total": 4560
      },
      "costUsd": 0.04
    }
  },
  "claims": {
    "promptVersion": "1",
    "stamp": {
      "agent": "pi",
      "agentVersion": "9.9.9",
      "model": "fake-provider/glm-5.3",
      "effort": null,
      "runAt": "2026-10-06T09:40:57.436Z",
      "tokens": {
        "input": 4200,
        "output": 360,
        "cacheRead": 0,
        "cacheWrite": 0,
        "total": 4560
      },
      "costUsd": 0.04
    },
    "outcome": "fell back",
    "detail": "the agent gave no usable answer (invalid-answer: the answer was invalid twice: the answer is missing \"claims\")",
    "claims": []
  }
}
  • Evidence: Engine serve process JSON-RPC transcript: staged review/stage notifications ending with the criteria-mapping stage, then the mapped final result (local file: ~/.no-mistakes/evidence/01M4876EJJR6C1B6R8HTFG3E21/engine-serve-protocol-transcript.jsonl)
Evidence: Stage announcements of the live CLI review run

second-look-engine: plain parts ready; grouping related hunks with pi second-look-engine: plain parts ready; ranking the parts with pi second-look-engine: plain parts ready; writing the story with pi second-look-engine: plain parts ready; comparing the change with its description and issues with pi second-look-engine: plain parts ready; listing the claims with pi second-look-engine: plain parts ready; mapping the acceptance criteria with pi

second-look-engine: plain parts ready; grouping related hunks with pi
second-look-engine: plain parts ready; ranking the parts with pi
second-look-engine: plain parts ready; writing the story with pi
second-look-engine: plain parts ready; comparing the change with its description and issues with pi
second-look-engine: plain parts ready; listing the claims with pi
second-look-engine: plain parts ready; mapping the acceptance criteria with pi
Evidence: Live adversarial run: invalid mapping answer retried once, then fell back with every criterion not checked
Live run: the same review command with a scripted agent whose criteria-mapping
answer leaves criteria out (first answer) and names an unknown criterion id
(retried answer). The engine retried once with the problems named, then fell
back: every criterion stays "not checked", the review itself still succeeds.

$ FAKE_PI_MODE=fallback node --import hook.mjs packages/engine/dist/main.js \
    review https://github.com/example-org/example-repo/pull/9 --token test-token --agent pi
exit=0
criteria.mapping.outcome = "fell back"
criteria.mapping.detail  = "the agent gave no usable answer (invalid-answer: the answer was invalid twice: \"a99\" is not a criterion id; no verdict for a1, a2, a3, a4, a5, a6)"
criteria[].verdict.kind  = ["not checked", "not checked", "not checked", "not checked", "not checked", "not checked"]
Evidence: Offline CI evaluation against the stored baseline, including the new criteria cases

baseline: 0 dropped, 0 missing, 0 gained, 95 unchanged, 0 without a baseline, 0 unstamped and not compared

(all)  claims-verdict:refuted  0  (expected failure: the review reports no claims)
(all)  claims-verdict:verified  0  (expected failure: the review reports no claims)
(all)  claims-evidence  0  (expected failure: the review reports no claims)
(all)  claims-fetch-offered  0  (expected failure: the review reports no claims)
results and trace: ~/Library/Caches/second-look/evaluation/2026-10-06T09-42-58-245Z
baseline: 0 dropped, 0 missing, 0 gained, 95 unchanged, 0 without a baseline, 0 unstamped and not compared

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - 1 info
  • ⚠️ packages/extension/src/overview.ts:386 - The ticket's acceptance criterion 2 reads "Each criterion links to the code that implements it, the automated tests that cover it, and any manual check the pull request reports." The implementation links code and tests (buttons that open the cited head-copy line via openHeadLine, extension.ts:449-458), but a manual check is rendered only as an inline quote plus plain text "description, line N" (overview.ts:386) with no link to jump to the description section or open the pull request description at that line. The choice is consistent with the overview's existing precedent for description-line references (describedWhere, overview.ts:484, is also plain text) and with the README's own wording ("each quoted from the description"), so this may be a deliberate reading that quoting the check satisfies the criterion; but since the quoted criterion literally says "links to ... any manual check", a human should confirm. If a link is required, the smallest remedy is making the manual-check quote (or its line reference) a link to the description section/line already rendered on the same page.
  • ℹ️ packages/engine/src/criteria-mapping.ts:195 - describePart (criteria-mapping.ts:195-199) is a near-verbatim copy of the part-rendering block in verdictsPrompt (packages/engine/src/verdicts.ts:276-281) — identical except the noise note's noun ("a criterion needs them" vs "a claim needs them") — and oneLine (criteria-mapping.ts:261) is a third private copy of the same helper (claims.ts:240, verdicts.ts:314). A shared part-description helper parameterised by the noun (and a shared oneLine) would keep the agent-facing part-presentation rule in one place and avoid drift between the two prompts. Mechanical, non-functional refactor.

🔧 Fix applied.
1 info still open:

  • ℹ️ packages/engine/src/criteria-mapping.ts:195 - describePart (criteria-mapping.ts:195-199) is a near-verbatim copy of the part-rendering block in verdictsPrompt (packages/engine/src/verdicts.ts:276-281) — identical except the noise note's noun ("a criterion needs them" vs "a claim needs them") — and oneLine (criteria-mapping.ts:261) is a third private copy of the same helper (claims.ts:240, verdicts.ts:314). A shared part-description helper parameterised by the noun (and a shared oneLine) would keep the agent-facing part-presentation rule in one place and avoid drift between the two prompts. Mechanical, non-functional refactor.
✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 7 of 7 scenarios driven live against the product
Scenario Result Live Evidence
Reviewer runs a review of a PR whose linked issue holds an acceptance-criteria checklist: each criterion is shown with its quoted text and one of the five verdicts, and a met verdict keeps the code, t… ✅ pass live engine-review-cli-result.json (live CLI run: six criteria → met, can't tell ×3, not met, needs manual check; a1 keeps app/send.py:5, tests/test_send.py:15 and the description's manual check at line 5)
Agent cites a line one off, quotes a manual check the description does not hold, and calls a criterion met with no evidence: the engine drops each verdict to can't tell and says why, keeping the revie… ✅ pass live engine-review-cli-result.json — a2 recheck: 'the quote of the citation app/send.py:6 is not on that line; the manual check "Tested on production" is not in the description'; a3 recheck: 'the verdict c…
Agent's mapping answer is invalid: the engine retries once with the problems named, then falls back with every criterion not checked and the mapping outcome saying why, without failing the review ✅ pass live engine-review-fallback.txt (exit 0, outcome 'fell back', all six criteria 'not checked')
Extension asks the engine to review over the JSON-RPC protocol: the criteria-mapping stage is announced while it runs and the final response carries the version-15 result with the criteria mapped and… ✅ pass live engine-serve-protocol-transcript.jsonl (real serve process on stdio; last stage 'mapping the acceptance criteria with pi' shows the criteria still not checked; final id=2 response carries mapped crite…
Reviewer opens the overview panel: the criteria and their verdicts appear at the top of the review, right after the story, with the verdict counts beside the heading and each criterion's reason and ev… ✅ pass live overview-live.html — the extension's real overviewHtml rendered from the live engine result, accepted by the extension's protocol guard; section order story → criteria → unexplained → claims; heading…
Each criterion links to the code that implements it and the tests that cover it (buttons that open the cited head-copy line) and to any manual check the PR reports (a link the reviewer follows to the… ✅ pass live overview-live.html (buttons data-criterion/data-evidence/data-index and button.manual 'description, line 5'); the delivered page script was executed and a manual-check click scrolled the description s…
The criteria-mapping prompt lands with its own evaluation cases and score (ADR 0006): versioned prompt registry entry, hand-labelled cases, its own score set, and the CI evaluation comparing them to t… ✅ pass live evaluation-baseline-check.txt — npm run eval (the real evaluation CLI, offline as CI runs it) over the new criteria-python and criteria-typescript cases: 0 dropped, 0 missing, 0 gained, 95 unchanged;…
  • npm run build
  • Live CLI run: node --import <hook> packages/engine/dist/main.js review https://github.com/example-org/example-repo/pull/9 --token test-token --agent pi against a disposable recorded-GitHub fixture (.live-validation, removed after) with a scripted fake pi on PATH — exit 0, result v15, criteria mapped (evidence: engine-review-cli-result.json, engine-review-cli-stages.txt)
  • Live adversarial CLI run with FAKE_PI_MODE=fallback: mapping answer invalid twice — retried once with problems named, then fell back, every criterion not checked, review still succeeds (evidence: engine-review-fallback.txt)
  • Live protocol run: spawned packages/engine/dist/main.js serve over real stdio, handshake + review request carrying the agent choice — review/stage notifications ending with 'mapping the acceptance criteria with pi', final response carrying the mapped criteria and the request's account label (evidence: engine-serve-protocol-transcript.jsonl)
  • Rendered the extension's real overviewHtml from the live engine result through the extension protocol guard, wrote it as evidence, and executed the delivered page script: manual-check button scrolls the description section into view, cite button posts openEvidence (temporary vitest driver, removed after; evidence: overview-live.html)
  • npx vitest run packages/engine/test/criteria-mapping.test.ts packages/engine/test/review.test.ts packages/engine/test/server.test.ts packages/engine/test/cli.test.ts packages/engine/test/criteria.test.ts — 85 passed
  • npx vitest run packages/extension/test/overview.test.ts packages/extension/test/review-result.test.ts packages/extension/test/engine-client.test.ts packages/evaluation/test/score.test.ts packages/evaluation/test/run.test.ts packages/evaluation/test/cli.test.ts — 175 passed
  • npm run eval (offline, model-free, as CI runs it) — baseline: 0 dropped, 0 missing, 0 gained, 95 unchanged (evidence: evaluation-baseline-check.txt)
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

…at show it

An agent pass, last after the claims are judged, gives each acceptance
criterion read from the linked issues a verdict: met, partly met, not
met, can't tell or needs manual check, with its reason, the lines of
code that implement it and of the tests that cover it, and the manual
checks the description reports. The engine re-reads every citation in
the head copy and finds every manual check in the description; one that
does not match makes the criterion can't tell, as does a met or partly
met verdict that shows nothing. Issue and description text reach the
agent only inside untrusted blocks. The review result (version 15)
carries each criterion's verdict and the criteria's mapping, and the
overview lists them with their counts, each cited line a button that
opens it in the head copy.

The criteria-mapping prompt lands with its cases, two planted cases
whose linked issues ask for every kind of verdict, one with a manual
check its description reports, and planted-typescript's two criteria,
and its scores: accuracy and false-met against the hand verdicts, and
the recall of the labelled code, tests and manual checks.
Pi 0.86.1 with zai-coding-cn/glm-5.3 at its default effort, over the
prompt's three cases: criteria-accuracy 1 over fourteen criteria,
criteria-false-met 0, criteria-code-recall 0.89, and criteria-tests-recall
and criteria-manual-recall 1. The plain rows are rewritten from a full
model-free run; every other agent row stays.
@lbildzinkas
lbildzinkas merged commit 684b3fd into master Oct 6, 2026
9 of 10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Map acceptance criteria to code, tests and manual checks

1 participant