Skip to content

feat: ask the agent to explain a part - #90

Merged
lbildzinkas merged 1 commit into
masterfrom
fm/second-look-ask-explain
Oct 7, 2026
Merged

lbildzinkas merged 1 commit into
masterfrom
fm/second-look-ask-explain

Conversation

@lbildzinkas

@lbildzinkas lbildzinkas commented Oct 7, 2026 •

Copy link
Copy Markdown
Owner

Intent

This pull request delivers the first ask of the v1 milestone and closes #47. Asks are fixed, typed requests a reviewer makes about one part from that part's context menu, answered in the overview panel with the answer's stamp, and never free chat. The first ask, explain, says what a part does and why it matters to the change, citing the part's lines.

To try it by hand, run explain on a must-review part and follow its citations.

The explain prompt lands with its evaluation cases and score, using plain checks: every cited line exists in the part, and the answer names no entity outside the change. The ask mechanism is typed so that a new ask is added in one place.

Validation on the development machine never launches, downloads or installs VS Code or any other tool. Extension tests that need VS Code run only in CI, so local validation goes through the engine protocol, the unit tests and the rendered output.

Closes #47.

What Changed

  • Added a typed ask mechanism: every ask is defined in one place, the ASKS registry in packages/engine/src/asks.ts (its menu title, prompt id and answer), reached through a new ask JSON-RPC request that names the ask by kind and the part by its index in the engine's latest review; the extension registers one part context-menu command per kind, declared in its manifest from the registry.
  • Added the first ask, Explain this part: the agent says what the part does and why it matters to the change, and the engine checks the answer before sending it — every cited line must be one the part shows with its quote on it, and the answer may name no file or code the change does not show — retrying a refused answer once; the overview shows answers at the top, newest first, with their stamp and citations as read-only links into the head or base copy, cleared by a new review.
  • Landed the explain prompt with its evaluation: labelled parts across the canary and library cases scored by two new plain checks, explain-cites-part and explain-names-in-change, with the baseline updated from a fresh model-free run of the plain rows.

Risk Assessment

✅ Low: A well-bounded feature addition that follows every established repo pattern (draft-comment's structure, story's name checks, ADR 0006 evaluation landing), satisfies both acceptance criteria, carries no token or GitHub write, and is covered by behavioral tests at every layer; no substantiated defects.

Testing

Drove the explain ask end-to-end through every surface this machine's boundary allows: the engine's ask protocol live over real process stdio (checked, stamped answers with citations that exist in the part; refused citations and outside names fed back for one retry and then plainly reported; unknown ask kinds and guard paths refused; asks making zero GitHub calls), the evaluation product live over the five recorded cases computing the explain scores with the plain checks biting on bad citations, and the extension's rendered overview HTML showing the answer with its stamp and cited-line buttons. All targeted unit, integration and evaluation tests pass, and the model-free eval run matches the committed baseline. The VS Code host UI itself is CI-only by the ticket's standing boundary, so the context-menu surface, the panel display inside a launched editor, and the VS Code-hosted command registration were not driven against the live product and are reported as untested rather than passed; the hand-label agreement check is by design unit-level. Nothing driven failed.

  • Live validation: ✅ go - 5 of 8 scenarios driven live against the product
Scenario Result Live Evidence
Reviewer makes the explain ask about a part and the engine answers over the protocol with sections, checked citations and a stamp ✅ pass live evidence/engine-ask-transcript.md happy session: ask explain parts 0 and 1 return AskAnswer with partName, both sections, citations that exist in the part (app/fresh.py:1 'def fresh():', web/cart.ts:2…
Plain checks guard: a citation to a line the part does not show and a name the change does not show are refused, the agent is asked again with the problems, and only a fixed answer is returned ✅ pass live evidence/engine-ask-transcript.md retry and always-bad sessions: the retry prompt carries 'the citation app/fresh.py:999 (head) names a line the part does not show' and '"zebra_unrelated_helper" is no…
No free chat: only the typed asks in the registry are accepted, and the ask request's guard paths give plain messages ✅ pass live evidence/engine-ask-transcript.md (ask 'chat' refused with '"ask": "explain"' as the only kind) and evidence/engine-serve-guard-transcript.md (the literal engine binary: handshake-first, no review of…
An ask reads and writes nothing on GitHub and carries no token ✅ pass live evidence/engine-ask-transcript.md GitHub request logs: every request belongs to the review; the asks that followed made zero requests and no POST; extension tests additionally assert the ask params ca…
The explain prompt lands with its evaluation cases and its score (ADR 0006), computed by the real evaluation product ✅ pass live evidence/evaluation-explain-score.md: the five registered cases' six labelled parts each got exactly one explain call, scored explain-cites-part / explain-names-in-change by the plain checks, which re…
The panel shows the answer with its stamp and its cited lines as buttons, and a base-side citation opens the base copy at its line ⏸️ untested no The prior payload did not establish a live result: it recorded live=false, resting on the extension session's rendered webview HTML produced against the repo's fake engine over real stdio and a passin…
The ask mechanism is typed so new asks are added in one place ⏸️ untested no The prior payload did not establish a live result: it recorded live=false, verifying the ASKS registry, the extension's registration and the manifest only by executing the extension's real activation…
The explain prompt's plain checks agree with hand labels before their scores are trusted ⏸️ untested no The prior payload did not establish a live result: it recorded live=false, with only the vitest unit run over thirteen hand-written explanations (26/26 labels agreeing) and no driving of a running pro…
Evidence: Engine ask protocol live transcript (review → explain asks, refused-then-fixed answer, never-passing answer, no-free-chat refusal, GitHub request log, generated explain prompt)
# Driving the engine's ask protocol live

## Session happy (agent mode: happy)

\### → sent

`` `json
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "initialize",
  "params": {
    "protocolVersion": 1
  }
}
`` `
\### ← received

`` `json
{
  "jsonrpc": "2.0",
  "id": 1,
  "result": {
    "protocolVersion": 1
  }
}
`` `
\### → sent

`` `json
{
  "jsonrpc": "2.0",
  "id": 2,
  "method": "review",
  "params": {
    "url": "https://github.com/example-org/example-repo/pull/7",
    "token": "ghp_fixture-only"
  }
}
`` `
\### ← received

`` `json
{
  "jsonrpc": "2.0",
  "id": 2,
  "result": {
    "version": 16,
    "pullRequest": {
      "url": "https://github.com/example-org/example-repo/pull/7",
      "number": 7,
      "title": "Reformat the Python helpers and tax the cart total",
      "author": "reviewer-login",
      "description": "Reformats app/reformat.py, moves the flag out of the discount block, restyles Greeter and taxes the cart total.",
      "base": "master",
      "head": "tidy-and-tax",
      "baseCommit": "5555555555555555555555555555555555555555",
      "headSha": "7777777777777777777777777777777777777777"
    },
    "copies": {
      "base": {
        "commit": "6666666666666666666666666666666666666666",
        "path": "~/.no-mistakes/worktrees/a09930ebf9a5/01M4AWY6BCP007QFYZTAQ9ECYF/.tmp-validation/session-happy-k0viRV/cache/github.com/example-org/example-repo/pull-7/6666666666666666666666666666666666666666",
        "reused": false
      },
      "head": {
        "commit": "7777777777777777777777777777777777777777",
        "path": "~/.no-mistakes/worktrees/a09930ebf9a5/01M4AWY6BCP007QFYZTAQ9ECYF/.tmp-validation/session-happy-k0viRV/cache/github.com/example-org/example-repo/pull-7/7777777777777777777777777777777777777777",
        "reused": false
      }
    },
    "parseTimeMs": 5.654,
    "parts": [
      {
        "path": "app/fresh.py",
        "changeKind": "addition",
        "isBinary": false,
        "oldMissingFinalNewline": false,
        "newMissingFinalNewline": false,
        "hunks": [
          {
            "oldStart": 0,
            "oldLines": 0,
            "newStart": 1,
            "newLines": 2,
            "lines": [
              {
                "kind": "addition",
                "newLineNumber": 1,
                "text": "def fresh():"
              },
              {
                "kind": "addition",
                "newLineNumber": 2,
                "text": "    return 1"
              }
            ],
            "entities": [
              {
                "kind": "function",
                "name": "fresh",
                "public": true,
                "change": "added"
              }
            ]
          }
        ],
        "additions": 2,
        "deletions": 0,
        "syntax": {
          "language": "python",
          "formattingOnly": {
            "status": "structure-changed",
            "reason": "the file is new"
          },
          "checksNotRun": []
        },
        "newMode": "100644",
        "noise": {
          "label": "none",
          "note": "no rule applied"
        },
        "name": "fresh in app/fresh.py",
        "origin": "plain",
        "signals": {
          "novelty": "new",
          "role": "code",
          "changedLines": 2,
          "publicSurface": [
            "fresh"
          ],
          "references": {
            "basis": "name-based",
            "names": [
              "fresh"
            ],
            "files": 1
          }
        },
        "rank": {
          "importance": "must review",
          "reason": "changes the public surface: fresh; code; new code; 2 changed lines",
          "signals": [
            "changes the public surface: fresh",
            "code",
            "new code",
            "2 changed lines"
          ]
        }
      },
      {
        "path": "web/cart.ts",
        "changeKind": "modification",
        "isBinary": false,
        "oldMissingFinalNewline": false,
        "newMissingFinalNewline": false,
        "hunks": [
          {
            "oldStart": 2,
            "oldLines": 7,
            "newStart": 2,
            "newLines": 7,
            "heading": "export class Cart {",
            "lines": [
              {
                "kind": "context",
                "oldLineNumber": 2,
                "newLineNumber": 2,
                "text": "  items: number[] = [];"
              },
              {
                "kind": "context",
                "oldLineNumber": 3,
                "newLineNumber": 3,
                "text": ""
              },
              {
                "kind": "context",
                "oldLineNumber": 4,
                "newLineNumber": 4,
                "text": "  total(): number {"
              },
              {
                "kind": "deletion",
                "oldLineNumber": 5,
                "text": "    return this.items.reduce((sum, item) => sum + item, 0);"
              },
              {
                "kind": "addition",
                "newLineNumber": 5,
                "text": "    return this.items.reduce((sum, item) => sum + item, 0) * 1.2;"
              },
              {
                "kind": "context",
                "oldLineNumber": 6,
                "newLineNumber": 6,
                "text": "  }"
              },
              {
                "kind": "context",
                "oldLineNumber": 7,
                "newLineNumber": 7,
                "text": "}"
              },
              {
                "kind": "context",
                "oldLineNumber": 8,
                "newLineNumber": 8,
                "text": ""
              }
            ],
            "entities": [
              {
                "kind": "method",
                "name": "Cart.total",
                "public": true,
                "change": "body"
              }
            ]
          }
        ],
        "additions": 1,
        "deletions": 1,
        "syntax": {
          "language": "typescript",
          "formattingOnly": {
            "status": "structure-changed",
            "reason": "the syntax tree changes at head line 5"
          },
          "checksNotRun": []
        },
        "noise": {
          "label": "none",
          "note": "no rule applied"
        },
        "name": "Cart.total in web/cart.ts",
        "origin": "plain",
        "signals": {
          "novelty": "changed",
          "role": "code",
          "changedLines": 2,
          "publicSurface": [],
          "references": {
            "basis": "name-based",
            "names": [
              "total"
            ],
            "files": 2
          }
        },
        "rank": {
          "importance": "worth reviewing",
          "reason": "code; changes code named in 2 other files (name-based); 2 changed lines",
          "signals": [
            "code",
            "changes code named in 2 other files (name-based)",
            "2 changed lines"
          ]
        }
      },
      {
        "path": "tests/test_fresh.py",
        "changeKind": "addition",
        "isBinary": false,
        "oldMissingFinalNewline": false,
        "newMissingFinalNewline": false,
        "hunks": [
          {
            "oldStart": 0,
            "oldLines": 0,
            "newStart": 1,
            "newLines": 5,
            "lines": [
              {
                "kind": "addition",
                "newLineNumber": 1,
                "text": "from app.fresh import fresh"
              },
              {
                "kind": "addition",
                "newLineNumber": 2,
                "text": ""
              },
              {
                "kind": "addition",
                "newLineNumber": 3,
                "text": ""
              },
              {
                "kind": "addition",
                "newLineNumber": 4,
                "text": "def test_fresh():"
              },
              {
                "kind": "addition",
                "newLineNumber": 5,
                "text": "    assert fresh() == 1"
              }
            ],
            "entities": [
              {
                "

... [83263 bytes truncated] ...


        },
        "rank": {
          "importance": "context",
          "reason": "formatting only, confirmed by the syntax trees; code; 10 changed lines",
          "signals": [
            "formatting only, confirmed by the syntax trees",
            "code",
            "10 changed lines"
          ]
        }
      }
    ],
    "grouping": {
      "by": "plain",
      "agent": {
        "promptVersion": "2",
        "stamp": {
          "agent": "fake",
          "agentVersion": "1.2.3",
          "model": "fake/model",
          "effort": null,
          "runAt": "2026-10-07T10:17:13.014Z"
        },
        "outcome": "fell back",
        "detail": "the agent gave no usable answer (invalid-answer: the answer was invalid twice: the answer is missing \"parts\")",
        "leftOut": 0
      }
    },
    "ranking": {
      "by": "plain",
      "agent": {
        "promptVersion": "1",
        "outcome": "not tested",
        "detail": "the agent ranking is the default only where its evaluation matched or beat the plain ranking, and fake at its default effort has none"
      }
    },
    "criteria": {
      "outcome": "read",
      "detail": "this pull request links no issue",
      "heading": "Acceptance criteria",
      "issues": [],
      "criteria": []
    },
    "pipeline": {
      "attestation": "missing",
      "detail": "the description carries no no-mistakes attestation",
      "steps": [],
      "findings": []
    },
    "ci": {
      "headSha": "7777777777777777777777777777777777777777",
      "outcome": "read",
      "detail": "0 check runs at the head commit, 0 failed; logs are read only for failed jobs",
      "checks": []
    },
    "story": {
      "promptVersion": "1",
      "stamp": {
        "agent": "fake",
        "agentVersion": "1.2.3",
        "model": "fake/model",
        "effort": null,
        "runAt": "2026-10-07T10:17:13.015Z"
      },
      "outcome": "fell back",
      "detail": "the agent gave no usable answer (invalid-answer: the answer was invalid twice: the answer is missing \"sentences\")",
      "sentences": []
    },
    "unexplained": {
      "promptVersion": "1",
      "parts": [],
      "described": [],
      "outcome": "fell back",
      "detail": "the agent gave no usable answer (invalid-answer: the answer was invalid twice: the answer is missing \"unexplained\"; the answer is missing \"described\")",
      "stamp": {
        "agent": "fake",
        "agentVersion": "1.2.3",
        "model": "fake/model",
        "effort": null,
        "runAt": "2026-10-07T10:17:13.015Z"
      }
    },
    "claims": {
      "promptVersion": "1",
      "stamp": {
        "agent": "fake",
        "agentVersion": "1.2.3",
        "model": "fake/model",
        "effort": null,
        "runAt": "2026-10-07T10:17:13.015Z"
      },
      "outcome": "fell back",
      "detail": "the agent gave no usable answer (invalid-answer: the answer was invalid twice: the answer is missing \"claims\")",
      "claims": []
    }
  }
}
`` `
\### → sent

`` `json
{
  "jsonrpc": "2.0",
  "id": 3,
  "method": "ask",
  "params": {
    "url": "https://github.com/example-org/example-repo/pull/7",
    "ask": "explain",
    "part": 0
  }
}
`` `
\### ← received

`` `json
{
  "jsonrpc": "2.0",
  "id": 3,
  "error": {
    "code": -32002,
    "message": "no answer: the agent gave no usable answer (invalid-answer: the answer was invalid twice: the citation app/fresh.py:999 (head) names a line the part does not show; \"zebra_unrelated_helper\" is not a name the change shows)"
  }
}
`` `
\### GitHub requests the engine made in this session
`` `
GET https://api.github.com/repos/example-org/example-repo/pulls/7
GET https://api.github.com/repos/example-org/example-repo/pulls/7
GET https://api.github.com/repos/example-org/example-repo/contents/.gitattributes?ref=7777777777777777777777777777777777777777
GET https://api.github.com/repos/example-org/example-repo/compare/5555555555555555555555555555555555555555...7777777777777777777777777777777777777777?per_page=1
GET https://api.github.com/repos/example-org/example-repo/commits/7777777777777777777777777777777777777777/check-runs?per_page=100
POST https://api.github.com/graphql
GET https://api.github.com/user
GET https://api.github.com/repos/example-org/example-repo/pulls/7/reviews?per_page=100
GET https://api.github.com/repos/example-org/example-repo/tarball/6666666666666666666666666666666666666666
GET https://api.github.com/repos/example-org/example-repo/tarball/7777777777777777777777777777777777777777
`` `
\### Explain prompt delivered to the agent (explain-prompt-1.txt) — the generated interface, first 60 lines
`` `
The pull request under review:
<untrusted-input id="55a78778c57855e3" source="pull request title">
Reformat the Python helpers and tax the cart total
</untrusted-input id="55a78778c57855e3">
<untrusted-input id="55a78778c57855e3" source="pull request description">
Reformats app/reformat.py, moves the flag out of the discount block, restyles Greeter and taxes the cart total.
</untrusted-input id="55a78778c57855e3">

Explain part p1, must review, one of the change's 7 parts. Its name, why it has
that importance and its diff follow as untrusted text. Each diff line is marked + when the change
adds it, - when it removes it, and blank when it stays, with its side and its line number there.

[p1] must review
<untrusted-input id="55a78778c57855e3" source="part p1">
name: fresh in app/fresh.py
why: changes the public surface: fresh; code; new code; 2 changed lines
file "app/fresh.py" (addition)
+ head 1: def fresh():
+ head 2:     return 1
</untrusted-input id="55a78778c57855e3">

The change's other parts, by name:
<untrusted-input id="55a78778c57855e3" source="other parts">
p2 (worth reviewing): Cart.total in web/cart.ts
p3 (context): test_fresh in tests/test_fresh.py
p4 (context): apply_discount in app/dedent.py
p5 (context): scripts/deploy.rb
p6 (context): load, Store, Store.__init__ and 1 more in app/reformat.py
p7 (context): Greeter, Greeter.Greet in src/Greeter.cs
</untrusted-input id="55a78778c57855e3">

When you have read enough, give your final message as the JSON value alone: start it with { and end it with }, with no summary of what you read before or after it.
`` `
\### Explain prompt delivered to the agent (explain-prompt-2.txt) — the generated interface, first 60 lines
`` `
The pull request under review:
<untrusted-input id="55a78778c57855e3" source="pull request title">
Reformat the Python helpers and tax the cart total
</untrusted-input id="55a78778c57855e3">
<untrusted-input id="55a78778c57855e3" source="pull request description">
Reformats app/reformat.py, moves the flag out of the discount block, restyles Greeter and taxes the cart total.
</untrusted-input id="55a78778c57855e3">

Explain part p1, must review, one of the change's 7 parts. Its name, why it has
that importance and its diff follow as untrusted text. Each diff line is marked + when the change
adds it, - when it removes it, and blank when it stays, with its side and its line number there.

[p1] must review
<untrusted-input id="55a78778c57855e3" source="part p1">
name: fresh in app/fresh.py
why: changes the public surface: fresh; code; new code; 2 changed lines
file "app/fresh.py" (addition)
+ head 1: def fresh():
+ head 2:     return 1
</untrusted-input id="55a78778c57855e3">

The change's other parts, by name:
<untrusted-input id="55a78778c57855e3" source="other parts">
p2 (worth reviewing): Cart.total in web/cart.ts
p3 (context): test_fresh in tests/test_fresh.py
p4 (context): apply_discount in app/dedent.py
p5 (context): scripts/deploy.rb
p6 (context): load, Store, Store.__init__ and 1 more in app/reformat.py
p7 (context): Greeter, Greeter.Greet in src/Greeter.cs
</untrusted-input id="55a78778c57855e3">

When you have read enough, give your final message as the JSON value alone: start it with { and end it with }, with no summary of what you read before or after it.

Your previous answer was rejected because it did not meet the schema and rules:
- the citation app/fresh.py:999 (head) names a line the part does not show
- "zebra_unrelated_helper" is not a name the change shows
Answer again with only the JSON value.
`` `
Evidence: Engine binary (serve) guard-path transcript
# Driving the engine binary (serve) live: the ask request's guard paths

Command: `node packages/engine/dist/main.js serve --cache-dir <disposable>`

\### → sent

`` `json
{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "ask",
  "params": {
    "url": "https://github.com/example-org/example-repo/pull/7",
    "ask": "explain",
    "part": 0
  }
}
`` `
\### ← received

`` `json
{
  "jsonrpc": "2.0",
  "id": 1,
  "error": {
    "code": -32001,
    "message": "the protocol starts with a version handshake: initialize before ask"
  }
}
`` `
\### → sent

`` `json
{
  "jsonrpc": "2.0",
  "id": 3,
  "method": "initialize",
  "params": {
    "protocolVersion": 1
  }
}
`` `
\### ← received

`` `json
{
  "jsonrpc": "2.0",
  "id": 3,
  "result": {
    "protocolVersion": 1
  }
}
`` `
\### → sent

`` `json
{
  "jsonrpc": "2.0",
  "id": 5,
  "method": "ask",
  "params": {
    "url": "https://github.com/example-org/example-repo/pull/7",
    "ask": "explain",
    "part": 0
  }
}
`` `
\### ← received

`` `json
{
  "jsonrpc": "2.0",
  "id": 5,
  "error": {
    "code": -32002,
    "message": "this engine has no part 0 of https://github.com/example-org/example-repo/pull/7 to answer about; wait for the review to finish, or review the pull request again"
  }
}
`` `
\### → sent

`` `json
{
  "jsonrpc": "2.0",
  "id": 7,
  "method": "ask",
  "params": {
    "url": "https://github.com/example-org/example-repo/pull/7",
    "ask": "explain",
    "part": 99
  }
}
`` `
\### ← received

`` `json
{
  "jsonrpc": "2.0",
  "id": 7,
  "error": {
    "code": -32002,
    "message": "this engine has no part 99 of https://github.com/example-org/example-repo/pull/7 to answer about; wait for the review to finish, or review the pull request again"
  }
}
`` `
\### → sent

`` `json
{
  "jsonrpc": "2.0",
  "id": 9,
  "method": "ask",
  "params": {
    "url": "https://github.com/example-org/example-repo/pull/7",
    "ask": "explain",
    "part": -1
  }
}
`` `
\### ← received

`` `json
{
  "jsonrpc": "2.0",
  "id": 9,
  "error": {
    "code": -32602,
    "message": "ask needs params: { \"url\": string, \"ask\": \"explain\", \"part\": number, \"agent\"?: { \"agent\": \"pi\" | \"claude-code\", \"model\"?: string, \"account\"?: string } }"
  }
}
`` `
Evidence: Live evaluation run: the explain prompt's five cases scored by the plain checks
# Driving the evaluation live: the explain prompt's cases and score

Each of the explain prompt's five cases (six labelled parts) was reviewed model-free and each labelled part explained by a scripted agent; the run scored the agent's own answers with the plain checks. Two of the answers were refused by the checks, on purpose in different ways: on `sindresorhus-ky-880` the scripted agent deliberately cited a line its part does not show, and on `encode-httpx-3690` it picked a blank context line of `HTTPServer.wait`, which the check refuses as quoting nothing — so `explain-cites-part` falls below 1 exactly where a citation was refused, while every answer that cites a real line of its part scores 1.

`` `json
[
  {
    "case": "canary-python",
    "name": "explain-cites-part",
    "value": 1,
    "better": "higher",
    "companionVersion": "0.1.0",
    "promptVersions": {
      "grouping": "2",
      "ranking": "1",
      "story": "1",
      "claims": "1",
      "verdicts": "4",
      "library-verdicts": "3",
      "draft-comment": "1",
      "explain": "1"
    },
    "agent": "fake",
    "agentVersion": "1.2.3",
    "model": "fake/model",
    "effort": "default",
    "runDate": "2026-10-07T10:18:01.250Z"
  },
  {
    "case": "canary-python",
    "name": "explain-names-in-change",
    "value": 1,
    "better": "higher",
    "companionVersion": "0.1.0",
    "promptVersions": {
      "grouping": "2",
      "ranking": "1",
      "story": "1",
      "claims": "1",
      "verdicts": "4",
      "library-verdicts": "3",
      "draft-comment": "1",
      "explain": "1"
    },
    "agent": "fake",
    "agentVersion": "1.2.3",
    "model": "fake/model",
    "effort": "default",
    "runDate": "2026-10-07T10:18:01.250Z"
  },
  {
    "case": "canary-csharp",
    "name": "explain-cites-part",
    "value": 1,
    "better": "higher",
    "companionVersion": "0.1.0",
    "promptVersions": {
      "grouping": "2",
      "ranking": "1",
      "story": "1",
      "claims": "1",
      "verdicts": "4",
      "library-verdicts": "3",
      "explain": "1"
    },
    "agent": "fake",
    "agentVersion": "1.2.3",
    "model": "fake/model",
    "effort": "default",
    "runDate": "2026-10-07T10:18:01.250Z"
  },
  {
    "case": "canary-csharp",
    "name": "explain-names-in-change",
    "value": 1,
    "better": "higher",
    "companionVersion": "0.1.0",
    "promptVersions": {
      "grouping": "2",
      "ranking": "1",
      "story": "1",
      "claims": "1",
      "verdicts": "4",
      "library-verdicts": "3",
      "explain": "1"
    },
    "agent": "fake",
    "agentVersion": "1.2.3",
    "model": "fake/model",
    "effort": "default",
    "runDate": "2026-10-07T10:18:01.250Z"
  },
  {
    "case": "criteria-python",
    "name": "explain-cites-part",
    "value": 1,
    "better": "higher",
    "companionVersion": "0.1.0",
    "promptVersions": {
      "criteria-mapping": "1",
      "draft-comment": "1",
      "explain": "1"
    },
    "agent": "fake",
    "agentVersion": "1.2.3",
    "model": "fake/model",
    "effort": "default",
    "runDate": "2026-10-07T10:18:01.250Z"
  },
  {
    "case": "criteria-python",
    "name": "explain-names-in-change",
    "value": 1,
    "better": "higher",
    "companionVersion": "0.1.0",
    "promptVersions": {
      "criteria-mapping": "1",
      "draft-comment": "1",
      "explain": "1"
    },
    "agent": "fake",
    "agentVersion": "1.2.3",
    "model": "fake/model",
    "effort": "default",
    "runDate": "2026-10-07T10:18:01.250Z"
  },
  {
    "case": "encode-httpx-3690",
    "name": "explain-cites-part",
    "value": 0.5,
    "better": "higher",
    "companionVersion": "0.1.0",
    "promptVersions": {
      "grouping": "2",
      "ranking": "1",
      "story": "1",
      "claims": "1",
      "verdicts": "4",
      "unexplained": "1",
      "explain": "1"
    },
    "agent": "fake",
    "agentVersion": "1.2.3",
    "model": "fake/model",
    "effort": "default",
    "runDate": "2026-10-07T10:18:01.250Z"
  },
  {
    "case": "encode-httpx-3690",
    "name": "explain-names-in-change",
    "value": 1,
    "better": "higher",
    "companionVersion": "0.1.0",
    "promptVersions": {
      "grouping": "2",
      "ranking": "1",
      "story": "1",
      "claims": "1",
      "verdicts": "4",
      "unexplained": "1",
      "explain": "1"
    },
    "agent": "fake",
    "agentVersion": "1.2.3",
    "model": "fake/model",
    "effort": "default",
    "runDate": "2026-10-07T10:18:01.250Z"
  },
  {
    "case": "sindresorhus-ky-880",
    "name": "explain-cites-part",
    "value": 0,
    "better": "higher",
    "companionVersion": "0.1.0",
    "promptVersions": {
      "grouping": "2",
      "ranking": "1",
      "story": "1",
      "claims": "1",
      "verdicts": "4",
      "unexplained": "1",
      "explain": "1"
    },
    "agent": "fake",
    "agentVersion": "1.2.3",
    "model": "fake/model",
    "effort": "default",
    "runDate": "2026-10-07T10:18:01.250Z"
  },
  {
    "case": "sindresorhus-ky-880",
    "name": "explain-names-in-change",
    "value": 1,
    "better": "higher",
    "companionVersion": "0.1.0",
    "promptVersions": {
      "grouping": "2",
      "ranking": "1",
      "story": "1",
      "claims": "1",
      "verdicts": "4",
      "unexplained": "1",
      "explain": "1"
    },
    "agent": "fake",
    "agentVersion": "1.2.3",
    "model": "fake/model",
    "effort": "default",
    "runDate": "2026-10-07T10:18:01.250Z"
  },
  {
    "case": "(all)",
    "name": "explain-cites-part",
    "value": 0.6666666666666666,
    "better": "higher",
    "companionVersion": "0.1.0",
    "promptVersions": {
      "explain": "1"
    },
    "agent": "fake",
    "agentVersion": "1.2.3",
    "model": "fake/model",
    "effort": "default",
    "runDate": "2026-10-07T10:18:01.250Z"
  },
  {
    "case": "(all)",
    "name": "explain-names-in-change",
    "value": 1,
    "better": "higher",
    "companionVersion": "0.1.0",
    "promptVersions": {
      "explain": "1"
    },
    "agent": "fake",
    "agentVersion": "1.2.3",
    "model": "fake/model",
    "effort": "default",
    "runDate": "2026-10-07T10:18:01.250Z"
  }
]
`` `

The trace recorded one explain call per part:

`` `
canary-python prompt=explain model=fake/model answered {"does":"It adds a function that returns one, `Fetch`.","matters":"The rest of t…
canary-csharp prompt=explain model=fake/model answered {"does":"It adds a function that returns one, `using`.","matters":"The rest of t…
criteria-python prompt=explain model=fake/model answered {"does":"It adds a function that returns one, `self`.","matters":"The rest of th…
encode-httpx-3690 prompt=explain model=fake/model answered {"does":"It adds a function that returns one, `Handle`.","matters":"The rest of …
encode-httpx-3690 prompt=explain model=fake/model answered {"does":"It adds a function that returns one.","matters":"The rest of the change…
sindresorhus-ky-880 prompt=explain model=fake/model answered {"does":"It merges the options the caller gives.","matters":"The request path bu…
`` `
  • Evidence: Rendered overview webview HTML: the asks section with the explain answer, its stamp and cited-line buttons (local file: ~/.no-mistakes/evidence/01M4AWY6BCP007QFYZTAQ9ECYF/overview-asks-rendered.html)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

✅ **Review** - passed

✅ No issues found.

✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 5 of 8 scenarios driven live against the product
Scenario Result Live Evidence
Reviewer makes the explain ask about a part and the engine answers over the protocol with sections, checked citations and a stamp ✅ pass live evidence/engine-ask-transcript.md happy session: ask explain parts 0 and 1 return AskAnswer with partName, both sections, citations that exist in the part (app/fresh.py:1 'def fresh():', web/cart.ts:2…
Plain checks guard: a citation to a line the part does not show and a name the change does not show are refused, the agent is asked again with the problems, and only a fixed answer is returned ✅ pass live evidence/engine-ask-transcript.md retry and always-bad sessions: the retry prompt carries 'the citation app/fresh.py:999 (head) names a line the part does not show' and '"zebra_unrelated_helper" is no…
No free chat: only the typed asks in the registry are accepted, and the ask request's guard paths give plain messages ✅ pass live evidence/engine-ask-transcript.md (ask 'chat' refused with '"ask": "explain"' as the only kind) and evidence/engine-serve-guard-transcript.md (the literal engine binary: handshake-first, no review of…
An ask reads and writes nothing on GitHub and carries no token ✅ pass live evidence/engine-ask-transcript.md GitHub request logs: every request belongs to the review; the asks that followed made zero requests and no POST; extension tests additionally assert the ask params ca…
The explain prompt lands with its evaluation cases and its score (ADR 0006), computed by the real evaluation product ✅ pass live evidence/evaluation-explain-score.md: the five registered cases' six labelled parts each got exactly one explain call, scored explain-cites-part / explain-names-in-change by the plain checks, which re…
The panel shows the answer with its stamp and its cited lines as buttons, and a base-side citation opens the base copy at its line ⏸️ untested no The prior payload did not establish a live result: it recorded live=false, resting on the extension session's rendered webview HTML produced against the repo's fake engine over real stdio and a passin…
The ask mechanism is typed so new asks are added in one place ⏸️ untested no The prior payload did not establish a live result: it recorded live=false, verifying the ASKS registry, the extension's registration and the manifest only by executing the extension's real activation…
The explain prompt's plain checks agree with hand labels before their scores are trusted ⏸️ untested no The prior payload did not establish a live result: it recorded live=false, with only the vitest unit run over thirteen hand-written explanations (26/26 labels agreeing) and no driving of a running pro…
  • npm ci && npm run build
  • npx vitest run packages/engine/test/explain.test.ts packages/engine/test/server.test.ts packages/engine/test/story.test.ts (67 tests)
  • npx vitest run packages/extension/test/asks.test.ts packages/extension/test/engine-client.test.ts packages/extension/test/overview.test.ts (80 tests)
  • npx vitest run packages/evaluation/test/explanations.test.ts packages/evaluation/test/run.test.ts packages/evaluation/test/score.test.ts packages/evaluation/test/prompts.test.ts packages/evaluation/test/baseline.test.ts (85 tests)
  • npm run test:integration (52 tests, includes 'explains a part from its context menu' and the ask command registration)
  • npm run eval (model-free, baseline comparison: 0 dropped, 0 missing, 0 gained)
  • node .tmp-validation/drive-engine-protocol.mjs — live: initialize → review → ask explain part 0/1 (checked, stamped answers), ask 'chat' refused, retry-after-refused answer accepted, never-passing answer refused with the plain-check problems; transcript at evidence/engine-ask-transcript.md
  • node .tmp-validation/drive-cli-serve.mjs — live: node packages/engine/dist/main.js serve over stdio; ask before handshake, ask with no review, unknown part, malformed part all refused with plain messages; transcript at evidence/engine-serve-guard-transcript.md
  • node .tmp-validation/drive-eval.mjs — live: runEvaluation over disposable copies of the five explain cases with a scripted agent; explain-cites-part/explain-names-in-change computed by the real checks, refusing a citation to a line the part does not show and an empty quote; results at evidence/evaluation-explain-score.md
  • temporary integration test (run then removed) driving the real extension session: ask from the part's tree node, overview opens focused on the answer, base-side citation opens the base copy at its line; rendered webview HTML saved to evidence/overview-asks-rendered.html
  • verified packages/evaluation/baseline.json holds the author's real agent run for the explain prompt (explain-cites-part 1, explain-names-in-change 0.83, pi/glm-5.3)
⚠️ **Document** - 1 info
  • ℹ️ docs/ux/README.md:17 - The reviewing-surface design record (and its mockup reviewing-surface.html) describes "the three asks" as buttons in a banner above the diff, but the shipped surface this change delivers offers asks in the part's context menu in the tree (one ask, Explain this part), as the root README documents. The ux doc is the record of the mockup comparison, so annotating it versus leaving it as the historical design artifact is a human decision; the mockup HTML would need the same treatment to stay in sync.
✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

Asks arrive: fixed, typed requests about one part, made from the part's
context menu in the tree and answered in the overview with their stamp;
no free chat. Every ask is defined in one place, the ASKS registry in
packages/engine/src/asks.ts: the words its menu entry shows, the prompt
that answers it and the sections its answer reads as. The engine's new
`ask` request names the ask by its kind and the part by its index in the
engine's latest review of the pull request, and carries no token; the
extension registers one command per kind, which its manifest declares
with the registry's title, and a test holds the two together.

The first ask, Explain this part, has the agent say what the part does
and why it matters to the change, citing the part's lines, each on the
side it names: head for an added or kept line, base for a removed one.
The part's name, reason and numbered diff, the other parts' names and
the pull request's own text reach the agent only inside untrusted
blocks. The engine checks the answer before sending it: every cited line
must be one the part shows, with its quote on it, and the answer must
name no file or code the change does not show. A refused answer is
retried once. The overview shows each answer at the top, newest first,
with its stamp, each cited line a link that opens it read-only in its
side's copy; a new review clears the answers.

The explain prompt lands with its cases and score: six parts across
canary-python, canary-csharp, criteria-python, encode-httpx-3690 and
sindresorhus-ky-880, scored by the plain checks explain-cites-part and
explain-names-in-change, which agree with all 26 hand labels on 13
hand-written explanations. The baseline records Pi 0.86.1 with
zai-coding-cn/glm-5.3 at its default effort: explain-cites-part 1 and
explain-names-in-change 0.83, the one name outside the change being
ProtocolError, which the agent read in the head copy. The plain rows are
rewritten from a full model-free run; every other agent row stays.
@lbildzinkas
lbildzinkas merged commit d796bc7 into master Oct 7, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Ask: explain this part

1 participant