Repository navigation
feat: ask the agent to explain a part - #90
Merged
Merged
Conversation
Asks arrive: fixed, typed requests about one part, made from the part's context menu in the tree and answered in the overview with their stamp; no free chat. Every ask is defined in one place, the ASKS registry in packages/engine/src/asks.ts: the words its menu entry shows, the prompt that answers it and the sections its answer reads as. The engine's new `ask` request names the ask by its kind and the part by its index in the engine's latest review of the pull request, and carries no token; the extension registers one command per kind, which its manifest declares with the registry's title, and a test holds the two together. The first ask, Explain this part, has the agent say what the part does and why it matters to the change, citing the part's lines, each on the side it names: head for an added or kept line, base for a removed one. The part's name, reason and numbered diff, the other parts' names and the pull request's own text reach the agent only inside untrusted blocks. The engine checks the answer before sending it: every cited line must be one the part shows, with its quote on it, and the answer must name no file or code the change does not show. A refused answer is retried once. The overview shows each answer at the top, newest first, with its stamp, each cited line a link that opens it read-only in its side's copy; a new review clears the answers. The explain prompt lands with its cases and score: six parts across canary-python, canary-csharp, criteria-python, encode-httpx-3690 and sindresorhus-ky-880, scored by the plain checks explain-cites-part and explain-names-in-change, which agree with all 26 hand labels on 13 hand-written explanations. The baseline records Pi 0.86.1 with zai-coding-cn/glm-5.3 at its default effort: explain-cites-part 1 and explain-names-in-change 0.83, the one name outside the change being ProtocolError, which the agent read in the head copy. The plain rows are rewritten from a full model-free run; every other agent row stays.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
This pull request delivers the first ask of the v1 milestone and closes #47. Asks are fixed, typed requests a reviewer makes about one part from that part's context menu, answered in the overview panel with the answer's stamp, and never free chat. The first ask, explain, says what a part does and why it matters to the change, citing the part's lines.
To try it by hand, run explain on a must-review part and follow its citations.
The explain prompt lands with its evaluation cases and score, using plain checks: every cited line exists in the part, and the answer names no entity outside the change. The ask mechanism is typed so that a new ask is added in one place.
Validation on the development machine never launches, downloads or installs VS Code or any other tool. Extension tests that need VS Code run only in CI, so local validation goes through the engine protocol, the unit tests and the rendered output.
Closes #47.
What Changed
ASKSregistry inpackages/engine/src/asks.ts(its menu title, prompt id and answer), reached through a newaskJSON-RPC request that names the ask by kind and the part by its index in the engine's latest review; the extension registers one part context-menu command per kind, declared in its manifest from the registry.explain-cites-partandexplain-names-in-change, with the baseline updated from a fresh model-free run of the plain rows.Risk Assessment
✅ Low: A well-bounded feature addition that follows every established repo pattern (draft-comment's structure, story's name checks, ADR 0006 evaluation landing), satisfies both acceptance criteria, carries no token or GitHub write, and is covered by behavioral tests at every layer; no substantiated defects.
Testing
Drove the explain ask end-to-end through every surface this machine's boundary allows: the engine's ask protocol live over real process stdio (checked, stamped answers with citations that exist in the part; refused citations and outside names fed back for one retry and then plainly reported; unknown ask kinds and guard paths refused; asks making zero GitHub calls), the evaluation product live over the five recorded cases computing the explain scores with the plain checks biting on bad citations, and the extension's rendered overview HTML showing the answer with its stamp and cited-line buttons. All targeted unit, integration and evaluation tests pass, and the model-free eval run matches the committed baseline. The VS Code host UI itself is CI-only by the ticket's standing boundary, so the context-menu surface, the panel display inside a launched editor, and the VS Code-hosted command registration were not driven against the live product and are reported as untested rather than passed; the hand-label agreement check is by design unit-level. Nothing driven failed.
Evidence: Engine ask protocol live transcript (review → explain asks, refused-then-fixed answer, never-passing answer, no-free-chat refusal, GitHub request log, generated explain prompt)
Evidence: Engine binary (serve) guard-path transcript
Evidence: Live evaluation run: the explain prompt's five cases scored by the plain checks
~/.no-mistakes/evidence/01M4AWY6BCP007QFYZTAQ9ECYF/overview-asks-rendered.html)Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
✅ **Review** - passed
✅ No issues found.
✅ **Test** - passed
✅ No issues found.
npm ci && npm run buildnpx vitest run packages/engine/test/explain.test.ts packages/engine/test/server.test.ts packages/engine/test/story.test.ts (67 tests)npx vitest run packages/extension/test/asks.test.ts packages/extension/test/engine-client.test.ts packages/extension/test/overview.test.ts (80 tests)npx vitest run packages/evaluation/test/explanations.test.ts packages/evaluation/test/run.test.ts packages/evaluation/test/score.test.ts packages/evaluation/test/prompts.test.ts packages/evaluation/test/baseline.test.ts (85 tests)npm run test:integration (52 tests, includes 'explains a part from its context menu' and the ask command registration)npm run eval (model-free, baseline comparison: 0 dropped, 0 missing, 0 gained)node .tmp-validation/drive-engine-protocol.mjs — live: initialize → review → ask explain part 0/1 (checked, stamped answers), ask 'chat' refused, retry-after-refused answer accepted, never-passing answer refused with the plain-check problems; transcript at evidence/engine-ask-transcript.mdnode .tmp-validation/drive-cli-serve.mjs — live: node packages/engine/dist/main.js serve over stdio; ask before handshake, ask with no review, unknown part, malformed part all refused with plain messages; transcript at evidence/engine-serve-guard-transcript.mdnode .tmp-validation/drive-eval.mjs — live: runEvaluation over disposable copies of the five explain cases with a scripted agent; explain-cites-part/explain-names-in-change computed by the real checks, refusing a citation to a line the part does not show and an empty quote; results at evidence/evaluation-explain-score.mdtemporary integration test (run then removed) driving the real extension session: ask from the part's tree node, overview opens focused on the answer, base-side citation opens the base copy at its line; rendered webview HTML saved to evidence/overview-asks-rendered.htmlverified packages/evaluation/baseline.json holds the author's real agent run for the explain prompt (explain-cites-part 1, explain-names-in-change 0.83, pi/glm-5.3)docs/ux/README.md:17- The reviewing-surface design record (and its mockup reviewing-surface.html) describes "the three asks" as buttons in a banner above the diff, but the shipped surface this change delivers offers asks in the part's context menu in the tree (one ask, Explain this part), as the root README documents. The ux doc is the record of the mockup comparison, so annotating it versus leaving it as the historical design artifact is a human decision; the mockup HTML would need the same treatment to stay in sync.✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.