Skip to content

feat(engine): agent grouping pass groups related hunks across files - #67

Merged
lbildzinkas merged 17 commits into
masterfrom
fm/second-look-grouping-pass
Oct 4, 2026
Merged

lbildzinkas merged 17 commits into
masterfrom
fm/second-look-grouping-pass

Conversation

@lbildzinkas

@lbildzinkas lbildzinkas commented Oct 2, 2026 •

Copy link
Copy Markdown
Owner

Intent

This change builds the agent grouping pass for the Second Look review companion, issue #28 of the v1 plan. The plain pass groups hunks into parts within one file; the reviewer's installed coding agent now proposes parts that group related hunks across files, such as a function, its caller and its test, named by the entities they touch.

The engine checks every agent answer before showing it. Hunks the agent leaves out go to a part marked "not grouped by the agent", so every changed line still belongs to exactly one part, and an invalid answer falls back to the plain grouping. In VS Code the tree appears first from the plain pass and updates in place when the agent's parts arrive, with a status line naming the stage still running, without losing the reviewer's place.

The grouping prompt is versioned and lands with its evaluation cases: the two canary cases plus three recorded public pull requests with hand-labelled groupings. The evaluation scores coverage as a hard gate at 100% and agreement with the hand labels as pairwise hunk agreement, and records the baseline per agent and model tried. Tests cover the coverage fallback and the invalid-answer fallback.

Validation does not launch the VS Code application locally: the extension integration tests run in CI, and local validation goes through the engine protocol and unit tests. Model calls go only through the locally installed Pi agent with its own sign-in, at modest volume, without reading any credential file and without paid API keys.

Closes #28.

What Changed

  • Engine: Added the agent grouping pass — a versioned prompt (v2) asks the reviewer's installed agent (Pi or Claude Code) to propose parts grouping related hunks across files. Answers are coverage-enforced: hunks the agent leaves out go to a part marked "not grouped by the agent", and a missing, invalid or failed answer falls back to the plain grouping, recorded in the result's new grouping field (schema v4 also gives parts an origin and otherFiles). review and serve gain --agent/--model/--effort/--agent-timeout, plain parts are announced on stderr while the agent works, and the engine stops its agent children on SIGINT/SIGTERM.
  • Extension: The ranked tree shows the plain parts first, then updates in place when the agent's parts arrive — stable node ids keep the reviewer's expanded sections and selected part across the regrouping — with a status line naming the stage still running; the configured agent and model are passed to the spawned engine.
  • Evaluation: Added a grouping-agreement score (pairwise hunk agreement against hand-labelled groups), made coverage a hard gate, stamped agent runs per agent/model with fallbacks listed, and recorded three public pull requests (httpx #3690, click #3781, ky #880) with hand-labelled groupings and licenses, beside the refreshed baseline.

Risk Assessment

✅ Low: Both pipeline-authored fix commits minimally and faithfully implement the recorded human decisions (concurrent request answering in the serve loop; agent children stopped on SIGINT/SIGTERM with tests exercising real processes), and this pass found no reachable defect in them or in the previously reviewed authored change beyond the user-decided items, so the branch is safe to merge pending the pipeline's own test step.

Testing

Drove the real engine process over its JSON-RPC protocol and CLI against a real public pull request (encode/httpx#3690) with a disposable fake-pi PATH shim and temp caches, plus a real-model evaluation through the locally installed Pi (4 calls, within the intent's volume cap): plain parts always arrived first in a review/stage notification, the agent's cross-file parts then replaced them with left-out hunks collected into 'not grouped by the agent', invalid answers retried once and fell back to the plain grouping with the reason stamped, coverage stayed 100% everywhere (offline and agent runs, exit 0), a sendReview was answered while an agent-stage review still ran and the review completed afterwards, SIGTERM stopped the engine's agent child, and the CLI refused agent misconfiguration. The extension tree's place-keeping was exercised only by the repository's unit tests; the intent forbids launching VS Code on this machine, so that surface could not be driven live and is reported untested. No failures; no findings.

  • Live validation: ✅ go - 9 of 10 scenarios driven live against the product
Scenario Result Live Evidence
Reviewer starts a review with the agent: the plain parts arrive first in a review/stage notification naming the running stage and deadline, then the review response carries the agent's grouping ✅ pass live serve-agent-grouping-transcript.txt (stage at 1.20s with grouping.by=plain, final response grouping.by=agent, promptVersion 2)
Agent answer that leaves hunks out: the left-out hunks land in one part marked 'not grouped by the agent', and every changed line still belongs to exactly one part ✅ pass live serve-agent-grouping-transcript.txt and review-cli-agent-transcript.txt: detail '5 hunks the agent left out are in a part marked not grouped by the agent', not-grouped part spans all four files, cover…
Invalid agent answer (wrong schema, then an unoffered hunk id): retried once, then the plain grouping stays with the reason in the result ✅ pass live serve-invalid-answer-transcript.txt: outcome=fell back, detail names 'part 1 names "h999", which was not offered'; fake agent's calls.jsonl recorded exactly 2 runs (the retry); parts keep origin:plain
Evaluation offline run: coverage is a hard gate at 100% and grouping agreement is scored against the hand labels, with the stored baseline compared ✅ pass live eval-offline-report.txt: exit 0, coverage 1 on all 12 rows, grouping-agreement for example-7 and the three public PRs, baseline 0 dropped / 0 missing / 66 unchanged
Agent evaluation through the locally installed Pi: coverage 100% on every agent row, agreement reported, rows stamped per agent and model, baseline holding the tried agent's rows ✅ pass live eval-agent-report.txt + eval-agent-results.json + eval-agent-trace.jsonl: pi 0.86.1 / zai-coding-cn/glm-5.3, coverage 1 on all agent rows, agreement 0.8482 overall, 4 model calls, no fallbacks; baseli…
Reviewer submits while the agent stage still runs: the sendReview request is answered while the review is open, and the review still completes with the agent's grouping ✅ pass live serve-send-during-review-transcript.txt: send id=2 answered at 1.50s (GitHub 401 from a bogus token; no write), review id=1 completed at 13.65s with grouping.by=agent
Engine killed mid-agent-run: the running agent child is stopped with the engine and does not outlive it ✅ pass live serve-sigterm-stops-agent-child.txt: child ignored SIGTERM itself, engine SIGTERM at 3.01s, engine exited, child gone by 5.03s (engine's SIGKILL after grace)
Engine CLI refuses agent misconfiguration: unknown --agent on serve, tuning flags without --agent on review; serve passes the configured agent and model to its engine ✅ pass live engine-cli-guards.txt (exit 1 with plain messages) and serve-send-during-review-transcript.txt run (fake agent's args contained --model concurrency-check-model)
Grouping prompt v2 is the prompt actually delivered: PR title, description, file paths and entity names all sit inside untrusted blocks, only hunk id and @@ range outside ✅ pass live delivered-grouping-prompt-excerpt.txt (engine-recorded stdin) and eval-agent-trace.jsonl (real-Pi run trace): <untrusted-input> wraps title, description and each hunk's header lines
The extension tree updates in place when agent parts arrive and keeps the reviewer's selected part and open diff ⏸️ untested no The prior payload did not establish a live result for this scenario: it recorded only unit tests (npx vitest run packages/extension/test/tree.test.ts) plus a live drive of the engine-side staging the…
Evidence: Serve protocol: plain stage then agent grouping with left-out hunks
[0.10s] INITIALIZE response id=0: protocolVersion=1
[1.20s] STAGE notification for review id=1: running="grouping related hunks with pi" timeoutMs=660000 plainParts=[HTTPParser.wait_ready, HTTPParser in src/ahttpx/_parsers.py | HTTPParser.wait_ready, HTTPParser in src/httpx/_parsers.py | ReadAheadParser.wait_ready, ReadAheadParser in src/ahttpx/_parsers.py | ReadAheadParser.wait_ready, ReadAheadParser in src/httpx/_parsers.py | HTTPConnection.handle_requests in src/ahttpx/_server.py | HTTPConnection.handle_requests in src/httpx/_server.py | HTTPServer.wait in src/httpx/_server.py] grouping.by=plain
[1.45s] REVIEW response id=1: grouping.by=agent agent.outcome=grouped promptVersion=2 detail="5 hunks the agent left out are in a part marked not grouped by the agent" model=fake-provider/fake-model parts=[HTTPParser.wait_ready and its callers in both servers(+3 files) origin:agent | not grouped by the agent(+3 files) origin:not grouped by the agent]
[1.45s] engine exited
Evidence: Serve protocol: invalid answer retried once, falls back to plain grouping
[0.10s] INITIALIZE response id=0: protocolVersion=1
[2.01s] STAGE notification for review id=1: running="grouping related hunks with pi" timeoutMs=660000 plainParts=[HTTPParser.wait_ready, HTTPParser in src/ahttpx/_parsers.py | HTTPParser.wait_ready, HTTPParser in src/httpx/_parsers.py | ReadAheadParser.wait_ready, ReadAheadParser in src/ahttpx/_parsers.py | ReadAheadParser.wait_ready, ReadAheadParser in src/httpx/_parsers.py | HTTPConnection.handle_requests in src/ahttpx/_server.py | HTTPConnection.handle_requests in src/httpx/_server.py | HTTPServer.wait in src/httpx/_server.py] grouping.by=plain
[2.58s] REVIEW response id=1: grouping.by=plain agent.outcome=fell back promptVersion=2 detail="the agent gave no usable answer (invalid-answer: the answer was invalid twice: part 1 names "h999", which was not offered)" model=fake-provider/fake-model parts=[HTTPParser.wait_ready, HTTPParser in src/ahttpx/_parsers.py origin:plain | HTTPParser.wait_ready, HTTPParser in src/httpx/_parsers.py origin:plain | ReadAheadParser.wait_ready, ReadAheadParser in src/ahttpx/_parsers.py origin:plain | ReadAheadParser.wait_ready, ReadAheadParser in src/httpx/_parsers.py origin:plain | HTTPConnection.handle_requests in src/ahttpx/_server.py origin:plain | HTTPConnection.handle_requests in src/httpx/_server.py origin:plain | HTTPServer.wait in src/httpx/_server.py origin:plain]
[2.58s] engine exited
Evidence: Serve protocol: sendReview answered during the agent stage (1.50s vs 13.65s)
[0.09s] INITIALIZE response id=0: protocolVersion=1
[1.37s] STAGE notification for review id=1: running="grouping related hunks with pi" timeoutMs=660000 plainParts=[HTTPParser.wait_ready, HTTPParser in src/ahttpx/_parsers.py | HTTPParser.wait_ready, HTTPParser in src/httpx/_parsers.py | ReadAheadParser.wait_ready, ReadAheadParser in src/ahttpx/_parsers.py | ReadAheadParser.wait_ready, ReadAheadParser in src/httpx/_parsers.py | HTTPConnection.handle_requests in src/ahttpx/_server.py | HTTPConnection.handle_requests in src/httpx/_server.py | HTTPServer.wait in src/httpx/_server.py] grouping.by=plain
[1.37s] SENT sendReview id=2 (bogus token; expecting an error answer, not silence)
[1.50s] SEND response id=2 answered: error code=-32002 message=Bad credentials - https://docs.github.com/rest
[13.65s] REVIEW response id=1: grouping.by=agent agent.outcome=grouped promptVersion=2 detail="5 hunks the agent left out are in a part marked not grouped by the agent" model=fake-provider/fake-model parts=[HTTPParser.wait_ready and its callers in both servers(+3 files) origin:agent | not grouped by the agent(+3 files) origin:not grouped by the agent]
[13.65s] ORDER: sendReview id=2 answered at 1.50s; review id=1 answered at 13.65s
[13.65s] VERDICT send-before-review: PASS
[13.66s] engine exited
Evidence: SIGTERM to the engine stops its running agent child
[1.51s] agent child running: pid=27329, alive=true
[3.01s] before signal: agent child alive=true
[3.01s] sending SIGTERM to the engine
[5.03s] engine exited with signal/code SIGTERM
[5.03s] after engine exit: agent child alive=false
[5.03s] VERDICT child-stopped-with-engine: PASS
Evidence: Review CLI with --agent pi: stderr plain-parts announcement and final agent result
$ ...engine review https://github.com/encode/httpx/pull/3690 --agent pi   (fake pi on PATH)
stderr:
second-look-engine: plain parts ready; grouping related hunks with pi
exit:0 — stdout result:
grouping.by=agent outcome=grouped promptVersion=2
- HTTPParser.wait_ready and its callers in both servers agent src/ahttpx/_parsers.py, src/ahttpx/_server.py, src/httpx/_parsers.py, src/httpx/_server.py
- not grouped by the agent not grouped by the agent src/ahttpx/_parsers.py, src/ahttpx/_server.py, src/httpx/_parsers.py, src/httpx/_server.py
Evidence: Engine CLI agent guards (unknown --agent, tuning without --agent)
$ node packages/engine/dist/main.js serve --agent bogus </dev/null
second-look-engine: unknown agent "bogus": choose pi or claude-code
exit:1

$ node packages/engine/dist/main.js review https://github.com/encode/httpx/pull/3690 --model some-model
second-look-engine: --model tunes the agent; pass --agent to run one
exit:1

$ node packages/engine/dist/main.js review https://github.com/encode/httpx/pull/3690 --effort high
second-look-engine: --effort tunes the agent; pass --agent to run one
exit:1
Evidence: Offline evaluation report: coverage 1 everywhere, baseline 0 drops

> eval
> node packages/evaluation/dist/main.js run --baseline packages/evaluation/baseline.json --cache-dir /var/folders/ng/f_l3tkbs2vd2dsh3lfd97mhh0000gn/T//sl-eval-cache --runs /var/folders/ng/f_l3tkbs2vd2dsh3lfd97mhh0000gn/T//sl-eval-runs

canary-csharp  coverage  1
canary-csharp  noise-precision:none  1
canary-csharp  noise-recall:none  1
canary-csharp  claims-found  0  (expected failure: the review reports no claims)
canary-csharp  claims-verdict:refuted  0  (expected failure: the review reports no claims)
canary-csharp  claims-evidence  0  (expected failure: the review reports no claims)
canary-csharp  claims-fetch-offered  0  (expected failure: the review reports no claims)
canary-python  coverage  1
canary-python  noise-precision:none  1
canary-python  noise-recall:none  1
canary-python  claims-found  0  (expected failure: the review reports no claims)
canary-python  claims-verdict:refuted  0  (expected failure: the review reports no claims)
canary-python  claims-evidence  0  (expected failure: the review reports no claims)
canary-python  claims-fetch-offered  0  (expected failure: the review reports no claims)
encode-httpx-3690  coverage  1
encode-httpx-3690  noise-precision:none  1
encode-httpx-3690  noise-recall:none  1
encode-httpx-3690  rank-median  4
encode-httpx-3690  rank-top-3  0.5
encode-httpx-3690  grouping-agreement  0.2778
example-42  coverage  1
example-42  noise-precision:generated:claimed  1
example-42  noise-recall:generated:claimed  1
example-42  noise-precision:lockfile:claimed  1
example-42  noise-recall:lockfile:claimed  1
example-42  noise-precision:moved or renamed:confirmed  1
example-42  noise-recall:moved or renamed:confirmed  1
example-42  noise-precision:none  1
example-42  noise-recall:none  1
example-42  noise-precision:snapshot:claimed  1
example-42  noise-recall:snapshot:claimed  1
example-42  rank-median  5.5
example-42  rank-top-3  0
example-7  coverage  1
example-7  noise-precision:none  1
example-7  noise-recall:none  1
example-7  rank-median  3
example-7  rank-top-3  0.5
example-7  grouping-agreement  0.9048
pallets-click-3781  coverage  1
pallets-click-3781  noise-precision:none  1
pallets-click-3781  noise-recall:none  1
pallets-click-3781  rank-median  1.5
pallets-click-3781  rank-top-3  1
pallets-click-3781  grouping-agreement  0.3333
seeded-csharp  coverage  1
seeded-csharp  noise-precision:none  1
seeded-csharp  noise-recall:none  1
seeded-csharp  rank-median  1
seeded-csharp  rank-top-3  1
seeded-python  coverage  1
seeded-python  noise-precision:none  1
seeded-python  noise-recall:none  1
seeded-python  rank-median  1
seeded-python  rank-top-3  1
seeded-typescript  coverage  1
seeded-typescript  noise-precision:none  1
seeded-typescript  noise-recall:none  1
seeded-typescript  rank-median  1
seeded-typescript  rank-top-3  1
sindresorhus-ky-880  coverage  1
sindresorhus-ky-880  noise-precision:none  1
sindresorhus-ky-880  noise-recall:none  1
sindresorhus-ky-880  rank-median  3
sindresorhus-ky-880  rank-top-3  1
sindresorhus-ky-880  grouping-agreement  0.3
(all)  coverage  1
(all)  noise-precision:generated:claimed  1
(all)  noise-recall:generated:claimed  1
(all)  noise-precision:lockfile:claimed  1
(all)  noise-recall:lockfile:claimed  1
(all)  noise-precision:moved or renamed:confirmed  1
(all)  noise-recall:moved or renamed:confirmed  1
(all)  noise-precision:none  1
(all)  noise-recall:none  1
(all)  noise-precision:snapshot:claimed  1
(all)  noise-recall:snapshot:claimed  1
(all)  rank-median  2.5
(all)  rank-top-3  0.6667
(all)  grouping-agreement  0.4196
(all)  claims-found  0  (expected failure: the review reports no claims)
(all)  claims-verdict:refuted  0  (expected failure: the review reports no claims)
(all)  claims-evidence  0  (expected failure: the review reports no claims)
(all)  claims-fetch-offered  0  (expected failure: the review reports no claims)
results and trace: /var/folders/ng/f_l3tkbs2vd2dsh3lfd97mhh0000gn/T/sl-eval-runs/2026-10-04T06-01-58-421Z
baseline: 0 dropped, 0 missing, 0 gained, 66 unchanged, 0 without a baseline, 0 unstamped and not compared
Evidence: Live agent evaluation report through Pi (coverage 1, agreement 0.8482)
canary-csharp  coverage  1
canary-csharp  noise-precision:none  1
canary-csharp  noise-recall:none  1
canary-csharp  claims-found  0  (expected failure: the review reports no claims)
canary-csharp  claims-verdict:refuted  0  (expected failure: the review reports no claims)
canary-csharp  claims-evidence  0  (expected failure: the review reports no claims)
canary-csharp  claims-fetch-offered  0  (expected failure: the review reports no claims)
canary-python  coverage  1
canary-python  noise-precision:none  1
canary-python  noise-recall:none  1
canary-python  claims-found  0  (expected failure: the review reports no claims)
canary-python  claims-verdict:refuted  0  (expected failure: the review reports no claims)
canary-python  claims-evidence  0  (expected failure: the review reports no claims)
canary-python  claims-fetch-offered  0  (expected failure: the review reports no claims)
encode-httpx-3690  coverage  1
encode-httpx-3690  noise-precision:none  1
encode-httpx-3690  noise-recall:none  1
encode-httpx-3690  rank-median  4
encode-httpx-3690  rank-top-3  0.5
encode-httpx-3690  grouping-agreement  0.2778
encode-httpx-3690  coverage  1  [pi 0.86.1 zai-coding-cn/glm-5.3]
encode-httpx-3690  grouping-agreement  0.5556  [pi 0.86.1 zai-coding-cn/glm-5.3]
example-42  coverage  1
example-42  noise-precision:generated:claimed  1
example-42  noise-recall:generated:claimed  1
example-42  noise-precision:lockfile:claimed  1
example-42  noise-recall:lockfile:claimed  1
example-42  noise-precision:moved or renamed:confirmed  1
example-42  noise-recall:moved or renamed:confirmed  1
example-42  noise-precision:none  1
example-42  noise-recall:none  1
example-42  noise-precision:snapshot:claimed  1
example-42  noise-recall:snapshot:claimed  1
example-42  rank-median  5.5
example-42  rank-top-3  0
example-7  coverage  1
example-7  noise-precision:none  1
example-7  noise-recall:none  1
example-7  rank-median  3
example-7  rank-top-3  0.5
example-7  grouping-agreement  0.9048
example-7  coverage  1  [pi 0.86.1 zai-coding-cn/glm-5.3]
example-7  grouping-agreement  0.9524  [pi 0.86.1 zai-coding-cn/glm-5.3]
pallets-click-3781  coverage  1
pallets-click-3781  noise-precision:none  1
pallets-click-3781  noise-recall:none  1
pallets-click-3781  rank-median  1.5
pallets-click-3781  rank-top-3  1
pallets-click-3781  grouping-agreement  0.3333
pallets-click-3781  coverage  1  [pi 0.86.1 zai-coding-cn/glm-5.3]
pallets-click-3781  grouping-agreement  1  [pi 0.86.1 zai-coding-cn/glm-5.3]
seeded-csharp  coverage  1
seeded-csharp  noise-precision:none  1
seeded-csharp  noise-recall:none  1
seeded-csharp  rank-median  1
seeded-csharp  rank-top-3  1
seeded-python  coverage  1
seeded-python  noise-precision:none  1
seeded-python  noise-recall:none  1
seeded-python  rank-median  1
seeded-python  rank-top-3  1
seeded-typescript  coverage  1
seeded-typescript  noise-precision:none  1
seeded-typescript  noise-recall:none  1
seeded-typescript  rank-median  1
seeded-typescript  rank-top-3  1
sindresorhus-ky-880  coverage  1
sindresorhus-ky-880  noise-precision:none  1
sindresorhus-ky-880  noise-recall:none  1
sindresorhus-ky-880  rank-median  3
sindresorhus-ky-880  rank-top-3  1
sindresorhus-ky-880  grouping-agreement  0.3
sindresorhus-ky-880  coverage  1  [pi 0.86.1 zai-coding-cn/glm-5.3]
sindresorhus-ky-880  grouping-agreement  1  [pi 0.86.1 zai-coding-cn/glm-5.3]
(all)  coverage  1
(all)  noise-precision:generated:claimed  1
(all)  noise-recall:generated:claimed  1
(all)  noise-precision:lockfile:claimed  1
(all)  noise-recall:lockfile:claimed  1
(all)  noise-precision:moved or renamed:confirmed  1
(all)  noise-recall:moved or renamed:confirmed  1
(all)  noise-precision:none  1
(all)  noise-recall:none  1
(all)  noise-precision:snapshot:claimed  1
(all)  noise-recall:snapshot:claimed  1
(all)  rank-median  2.5
(all)  rank-top-3  0.6667
(all)  grouping-agreement  0.4196
(all)  claims-found  0  (expected failure: the review reports no claims)
(all)  claims-verdict:refuted  0  (expected failure: the review reports no claims)
(all)  claims-evidence  0  (expected failure: the review reports no claims)
(all)  claims-fetch-offered  0  (expected failure: the review reports no claims)
(all)  coverage  1  [pi 0.86.1 zai-coding-cn/glm-5.3]
(all)  grouping-agreement  0.8482  [pi 0.86.1 zai-coding-cn/glm-5.3]
results and trace: /var/folders/ng/f_l3tkbs2vd2dsh3lfd97mhh0000gn/T/sl-agent-runs/2026-10-04T06-02-13-235Z
Evidence: Agent evaluation results.json (stamped rows per agent and model)
{
  "rows": [
    {
      "case": "canary-csharp",
      "name": "coverage",
      "value": 1,
      "better": "higher",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "canary-csharp",
      "name": "noise-precision:none",
      "value": 1,
      "better": "higher",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "canary-csharp",
      "name": "noise-recall:none",
      "value": 1,
      "better": "higher",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "canary-csharp",
      "name": "claims-found",
      "value": 0,
      "better": "higher",
      "note": "expected failure: the review reports no claims",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "canary-csharp",
      "name": "claims-verdict:refuted",
      "value": 0,
      "better": "higher",
      "note": "expected failure: the review reports no claims",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "canary-csharp",
      "name": "claims-evidence",
      "value": 0,
      "better": "higher",
      "note": "expected failure: the review reports no claims",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "canary-csharp",
      "name": "claims-fetch-offered",
      "value": 0,
      "better": "higher",
      "note": "expected failure: the review reports no claims",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "canary-python",
      "name": "coverage",
      "value": 1,
      "better": "higher",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "canary-python",
      "name": "noise-precision:none",
      "value": 1,
      "better": "higher",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "canary-python",
      "name": "noise-recall:none",
      "value": 1,
      "better": "higher",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "canary-python",
      "name": "claims-found",
      "value": 0,
      "better": "higher",
      "note": "expected failure: the review reports no claims",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "canary-python",
      "name": "claims-verdict:refuted",
      "value": 0,
      "better": "higher",
      "note": "expected failure: the review reports no claims",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "canary-python",
      "name": "claims-evidence",
      "value": 0,
      "better": "higher",
      "note": "expected failure: the review reports no claims",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "canary-python",
      "name": "claims-fetch-offered",
      "value": 0,
      "better": "higher",
      "note": "expected failure: the review reports no claims",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "encode-httpx-3690",
      "name": "coverage",
      "value": 1,
      "better": "higher",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "encode-httpx-3690",
      "name": "noise-precision:none",
      "value": 1,
      "better": "higher",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "encode-httpx-3690",
      "name": "noise-recall:none",
      "value": 1,
      "better": "higher",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "encode-httpx-3690",
      "name": "rank-median",
      "value": 4,
      "better": "lower",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "encode-httpx-3690",
      "name": "rank-top-3",
      "value": 0.5,
      "better": "higher",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "encode-httpx-3690",
      "name": "grouping-agreement",
      "value": 0.2777777777777778,
      "better": "higher",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "encode-httpx-3690",
      "name": "coverage",
      "value": 1,
      "better": "higher",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "pi",
      "agentVersion": "0.86.1",
      "model": "zai-coding-cn/glm-5.3",
      "effort": "default",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "encode-httpx-3690",
      "name": "

... [17800 bytes truncated] ...

ersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "pi",
      "agentVersion": "0.86.1",
      "model": "zai-coding-cn/glm-5.3",
      "effort": "default",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "sindresorhus-ky-880",
      "name": "grouping-agreement",
      "value": 1,
      "better": "higher",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "pi",
      "agentVersion": "0.86.1",
      "model": "zai-coding-cn/glm-5.3",
      "effort": "default",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "(all)",
      "name": "coverage",
      "value": 1,
      "better": "higher",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "(all)",
      "name": "noise-precision:generated:claimed",
      "value": 1,
      "better": "higher",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "(all)",
      "name": "noise-recall:generated:claimed",
      "value": 1,
      "better": "higher",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "(all)",
      "name": "noise-precision:lockfile:claimed",
      "value": 1,
      "better": "higher",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "(all)",
      "name": "noise-recall:lockfile:claimed",
      "value": 1,
      "better": "higher",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "(all)",
      "name": "noise-precision:moved or renamed:confirmed",
      "value": 1,
      "better": "higher",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "(all)",
      "name": "noise-recall:moved or renamed:confirmed",
      "value": 1,
      "better": "higher",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "(all)",
      "name": "noise-precision:none",
      "value": 1,
      "better": "higher",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "(all)",
      "name": "noise-recall:none",
      "value": 1,
      "better": "higher",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "(all)",
      "name": "noise-precision:snapshot:claimed",
      "value": 1,
      "better": "higher",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "(all)",
      "name": "noise-recall:snapshot:claimed",
      "value": 1,
      "better": "higher",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "(all)",
      "name": "rank-median",
      "value": 2.5,
      "better": "lower",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "(all)",
      "name": "rank-top-3",
      "value": 0.6666666666666666,
      "better": "higher",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "(all)",
      "name": "grouping-agreement",
      "value": 0.41964285714285715,
      "better": "higher",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "(all)",
      "name": "claims-found",
      "value": 0,
      "better": "higher",
      "note": "expected failure: the review reports no claims",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "(all)",
      "name": "claims-verdict:refuted",
      "value": 0,
      "better": "higher",
      "note": "expected failure: the review reports no claims",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "(all)",
      "name": "claims-evidence",
      "value": 0,
      "better": "higher",
      "note": "expected failure: the review reports no claims",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "(all)",
      "name": "claims-fetch-offered",
      "value": 0,
      "better": "higher",
      "note": "expected failure: the review reports no claims",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "none",
      "agentVersion": "none",
      "model": "none",
      "effort": "none",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "(all)",
      "name": "coverage",
      "value": 1,
      "better": "higher",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "pi",
      "agentVersion": "0.86.1",
      "model": "zai-coding-cn/glm-5.3",
      "effort": "default",
      "runDate": "2026-10-04T06:02:13.235Z"
    },
    {
      "case": "(all)",
      "name": "grouping-agreement",
      "value": 0.8482142857142857,
      "better": "higher",
      "companionVersion": "0.1.0",
      "promptVersions": {
        "grouping": "2"
      },
      "agent": "pi",
      "agentVersion": "0.86.1",
      "model": "zai-coding-cn/glm-5.3",
      "effort": "default",
      "runDate": "2026-10-04T06:02:13.235Z"
    }
  ],
  "failures": [],
  "fallbacks": []
}
  • Evidence: Agent evaluation trace (4 calls, promptVersion 2, delivered prompts) (local file: ~/.no-mistakes/evidence/01M42NEJEXTP669SNQX5PPM2ME/eval-agent-trace.jsonl)
Evidence: Delivered grouping prompt v2 excerpt: hunk headers inside untrusted blocks
cc95730" source="pull request title">
Add `.wait_ready` to parser for clean server disconnects
</untrusted-input id="1f6a11491cc95730">
<untrusted-input id="1f6a11491cc95730" source="pull request description">
Add `.wait_ready()` to `HTTPParser`...

We need this in order to differentiate between clean disconnects at the start of a new request/response cycle, rather than a `ProtocolError` while calling `recv_method_line()`.

</untrusted-input id="1f6a11491cc95730">

Group these 9 hunks. Each has its id and its range; its file, the entities it
touches and its lines follow as untrusted text.

[h1] @@ -224,6 +224,13 @@
<untrusted-input id="1f6a11491cc95730" source="hunk h1">
"src/ahttpx/_parsers.py" (modified) touches method HTTPParser.wait_ready (added), class HTTPParser (body)
             # Handle body close
             self.send_state = State.DONE
 
+    async def wait_ready(self) -> bool:
+        """
+        Wait until read data starts arriving, and return `True`.
+        Return `False` if the stream closes.
+        """
+        return await self.parser.wait_ready()
+
     async def recv_method_line(self) -> tuple[bytes, bytes, bytes]:
         """
         Receive the initial request method line:
</untrusted-input id="1f6a11491cc95730">
[h2] @@ -453,6 +460,15 @@
<untrusted-input id="1f6a11491cc95730" source="hunk h2">
"src/ahttpx/_parsers.py" (modified) touches method ReadAheadParser.wait_ready (added), class ReadAheadParser (body)
         assert self._buffer == b'

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 2 issues found → auto-fixed (2) ✅
  • ⚠️ packages/engine/src/server.ts:104 - The serve loop handles one request at a time (await review(...) before reading the next line), and the agent grouping stage now keeps a review request open for minutes (groupingStageTimeoutMs = 2×5 min + 60 s) by design — the reviewer is invited to read the plain tree and write comments during it. Concrete sequence: reviewer starts a review (engine spawned with --agent), the plain stage arrives, they write comments and press Submit while the agent stage still runs; submitReview → engineSend (extension.ts:463) reuses the same engine, the sendReview line is buffered unread behind the running review, and after SEND_REVIEW_TIMEOUT_MS = 60 s (engine-client.ts:80) expire() rejects the send and calls this.dispose() (engine-client.ts:259), which SIGTERMs the engine and fails the in-flight review request too — the reviewer gets 'the engine did not answer in time', the agent's grouping is lost, and only the comments survive. Remedy needs a product/protocol decision, hence ask-user: either the engine answers a send while a review runs (concurrent request handling at this shared boundary), or the extension holds a send until the running review settles (serialize in ReviewSession). Sibling paths that must hold the same invariant: a second review request on the same engine (the extension side-steps it by disposing first at extension.ts:239), and the stage-notification deadline restart (engine-client.ts onResponse) which correctly extends only the review's own deadline.
  • ℹ️ packages/extension/src/extension.ts:239 - Killing the engine mid-agent-run orphans the running agent child. The new replaced-review flow disposes the engine while the grouping agent works (this.engine?.dispose() at extension.ts:239, and the same mid-run kill is reachable via the send-timeout dispose at engine-client.ts:259 and deactivate at extension.ts:515), but the engine installs no SIGTERM handler (packages/engine/src/main.ts) and the adapters kill their child only through the engine's own timeout timer (pi.ts:246, claude-code.ts:298), so the pi/claude process survives its parent and runs to completion on the reviewer's subscription after the engine died. Bounded impact: the one-shot run finishes by itself; the cost is one wasted agent run per killed review. Remedy (mechanical): the engine kills its running agent children on SIGTERM/SIGINT, or the adapters tie the child to the parent's exit.

🔧 Fix applied.
2 issues (1 warning, 1 info) still open:

  • ⚠️ packages/engine/src/server.ts:104 - The serve loop handles one request at a time (await review(...) before reading the next line), and the agent grouping stage now keeps a review request open for minutes (groupingStageTimeoutMs = 2×5 min + 60 s) by design — the reviewer is invited to read the plain tree and write comments during it. Concrete sequence: reviewer starts a review (engine spawned with --agent), the plain stage arrives, they write comments and press Submit while the agent stage still runs; submitReview → engineSend (extension.ts:463) reuses the same engine, the sendReview line is buffered unread behind the running review, and after SEND_REVIEW_TIMEOUT_MS = 60 s (engine-client.ts:80) expire() rejects the send and calls this.dispose() (engine-client.ts:259), which SIGTERMs the engine and fails the in-flight review request too — the reviewer gets 'the engine did not answer in time', the agent's grouping is lost, and only the comments survive. Remedy needs a product/protocol decision, hence ask-user: either the engine answers a send while a review runs (concurrent request handling at this shared boundary), or the extension holds a send until the running review settles (serialize in ReviewSession). Sibling paths that must hold the same invariant: a second review request on the same engine (the extension side-steps it by disposing first at extension.ts:239), and the stage-notification deadline restart (engine-client.ts onResponse) which correctly extends only the review's own deadline.
  • ℹ️ packages/extension/src/extension.ts:239 - Killing the engine mid-agent-run orphans the running agent child. The new replaced-review flow disposes the engine while the grouping agent works (this.engine?.dispose() at extension.ts:239, and the same mid-run kill is reachable via the send-timeout dispose at engine-client.ts:259 and deactivate at extension.ts:515), but the engine installs no SIGTERM handler (packages/engine/src/main.ts) and the adapters kill their child only through the engine's own timeout timer (pi.ts:246, claude-code.ts:298), so the pi/claude process survives its parent and runs to completion on the reviewer's subscription after the engine died. Bounded impact: the one-shot run finishes by itself; the cost is one wasted agent run per killed review. Remedy (mechanical): the engine kills its running agent children on SIGTERM/SIGINT, or the adapters tie the child to the parent's exit.

🔧 Fix applied.
✅ Re-checked - no issues remain.

✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 9 of 10 scenarios driven live against the product
Scenario Result Live Evidence
Reviewer starts a review with the agent: the plain parts arrive first in a review/stage notification naming the running stage and deadline, then the review response carries the agent's grouping ✅ pass live serve-agent-grouping-transcript.txt (stage at 1.20s with grouping.by=plain, final response grouping.by=agent, promptVersion 2)
Agent answer that leaves hunks out: the left-out hunks land in one part marked 'not grouped by the agent', and every changed line still belongs to exactly one part ✅ pass live serve-agent-grouping-transcript.txt and review-cli-agent-transcript.txt: detail '5 hunks the agent left out are in a part marked not grouped by the agent', not-grouped part spans all four files, cover…
Invalid agent answer (wrong schema, then an unoffered hunk id): retried once, then the plain grouping stays with the reason in the result ✅ pass live serve-invalid-answer-transcript.txt: outcome=fell back, detail names 'part 1 names "h999", which was not offered'; fake agent's calls.jsonl recorded exactly 2 runs (the retry); parts keep origin:plain
Evaluation offline run: coverage is a hard gate at 100% and grouping agreement is scored against the hand labels, with the stored baseline compared ✅ pass live eval-offline-report.txt: exit 0, coverage 1 on all 12 rows, grouping-agreement for example-7 and the three public PRs, baseline 0 dropped / 0 missing / 66 unchanged
Agent evaluation through the locally installed Pi: coverage 100% on every agent row, agreement reported, rows stamped per agent and model, baseline holding the tried agent's rows ✅ pass live eval-agent-report.txt + eval-agent-results.json + eval-agent-trace.jsonl: pi 0.86.1 / zai-coding-cn/glm-5.3, coverage 1 on all agent rows, agreement 0.8482 overall, 4 model calls, no fallbacks; baseli…
Reviewer submits while the agent stage still runs: the sendReview request is answered while the review is open, and the review still completes with the agent's grouping ✅ pass live serve-send-during-review-transcript.txt: send id=2 answered at 1.50s (GitHub 401 from a bogus token; no write), review id=1 completed at 13.65s with grouping.by=agent
Engine killed mid-agent-run: the running agent child is stopped with the engine and does not outlive it ✅ pass live serve-sigterm-stops-agent-child.txt: child ignored SIGTERM itself, engine SIGTERM at 3.01s, engine exited, child gone by 5.03s (engine's SIGKILL after grace)
Engine CLI refuses agent misconfiguration: unknown --agent on serve, tuning flags without --agent on review; serve passes the configured agent and model to its engine ✅ pass live engine-cli-guards.txt (exit 1 with plain messages) and serve-send-during-review-transcript.txt run (fake agent's args contained --model concurrency-check-model)
Grouping prompt v2 is the prompt actually delivered: PR title, description, file paths and entity names all sit inside untrusted blocks, only hunk id and @@ range outside ✅ pass live delivered-grouping-prompt-excerpt.txt (engine-recorded stdin) and eval-agent-trace.jsonl (real-Pi run trace): <untrusted-input> wraps title, description and each hunk's header lines
The extension tree updates in place when agent parts arrive and keeps the reviewer's selected part and open diff ⏸️ untested no The prior payload did not establish a live result for this scenario: it recorded only unit tests (npx vitest run packages/extension/test/tree.test.ts) plus a live drive of the engine-side staging the…
  • npx vitest run packages/engine/test/grouping.test.ts packages/engine/test/review.test.ts packages/engine/test/server.test.ts packages/engine/test/agent-children.test.ts packages/engine/test/cli.test.ts (64 tests pass: coverage/invalid-answer fallbacks, concurrent serve, signal handling, CLI guards)
  • npx vitest run packages/extension/test/tree.test.ts packages/extension/test/review-result.test.ts packages/extension/test/engine-client.test.ts (52 tests pass, incl. 'finds the part that now holds the first hunk of the part the reviewer was on' and groupingStatus tests)
  • npm run eval -- --cache-dir/--runs in temp (offline evaluation over all 10 cases: exit 0, coverage 1 on every row, grouping-agreement computed for example-7 + 3 public PRs, baseline 0 dropped/0 missing/66 unchanged)
  • node packages/evaluation/dist/main.js run --agent pi (live evaluation through locally installed Pi 0.86.1 / zai-coding-cn/glm-5.3: 4 model calls, coverage 1 on all agent rows, agreement 0.8482 overall, no fallbacks; trace stamped promptVersion 2)
  • Live serve-protocol drives against node packages/engine/dist/main.js serve on https://github.com/encode/httpx/pull/3690 with a PATH-shimmed fake pi and temp cache: stage notification then agent grouping; left-out hunks -> 'not grouped by the agent'; two invalid answers -> retry once -> plain fallback; sendReview answered at 1.50s while the review finished at 13.65s; SIGTERM to the engine stopped its SIGTERM-ignoring agent child
  • node packages/engine/dist/main.js review .../pull/3690 --agent pi (fake pi): stderr announces 'plain parts ready; grouping related hunks with pi', stdout result grouping.by=agent with cross-file part + not-grouped part
  • CLI guards: serve --agent bogus -> 'unknown agent' exit 1; review --model/--effort without --agent -> 'tunes the agent; pass --agent' exit 1; serve --model reached the agent child's args
  • Verified the delivered grouping prompt (recorded in the real-Pi trace and the fake agent's stdin) carries PR title, description, file paths and entity names inside <untrusted-input> blocks with only hunk id and @@ range outside
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

@lbildzinkas
lbildzinkas force-pushed the fm/second-look-grouping-pass branch from 8e9a192 to 89ca8fd Compare October 2, 2026 20:17
@lbildzinkas lbildzinkas changed the title feat(engine): agent grouping pass groups related hunks across files into named parts feat(engine): agent grouping pass with staged tree updates Oct 2, 2026
The reviewer's installed agent groups related hunks across files into parts named by the entities they touch, behind a versioned grouping prompt. The engine checks every answer: an invalid one is retried once and then the plain grouping stays, and hunks a valid answer leaves out go to a part marked not grouped by the agent, so coverage holds. A part can now span files (review result version 4), and a review arrives in stages over the protocol: the plain result first in a review/stage notification, then the agent's parts. The extension shows the plain tree first with a status line naming the running stage, then regroups in place and keeps the reviewer's selected part.

The evaluation scores grouping by pairwise hunk agreement with hand labels, gates on 100% coverage, runs the prompt through Pi with --agent pi, traces each call, and keeps one baseline per agent and model. The prompt's cases are the two canaries, example-7 and three recorded public pull requests.
@lbildzinkas
lbildzinkas force-pushed the fm/second-look-grouping-pass branch from 89ca8fd to a5d7311 Compare October 2, 2026 22:10
@lbildzinkas lbildzinkas changed the title feat(engine): agent grouping pass with staged tree updates feat(engine): agent grouping pass with staged review updates Oct 2, 2026
The reviewer's installed agent groups related hunks across files into parts named by the entities they touch, behind a versioned grouping prompt. The engine checks every answer: an invalid one is retried once and then the plain grouping stays, and hunks a valid answer leaves out go to a part marked not grouped by the agent, so coverage holds. A part can now span files (review result version 4), and a review arrives in stages over the protocol: the plain result first in a review/stage notification, then the agent's parts. The extension shows the plain tree first with a status line naming the running stage, then regroups in place and keeps the reviewer's selected part.

The evaluation scores grouping by pairwise hunk agreement with hand labels, gates on 100% coverage, runs the prompt through Pi with --agent pi, traces each call, and keeps one baseline per agent and model. The prompt's cases are the two canaries, example-7 and three recorded public pull requests.
…ping-pass

# Conflicts:
#	README.md
#	packages/engine/src/rpc.ts
#	packages/engine/src/server.ts
#	packages/engine/test/server.test.ts
#	packages/extension/src/engine-client.ts
#	packages/extension/src/extension.ts
#	packages/extension/src/tree.ts
#	packages/extension/test/integration/extension.test.ts
@lbildzinkas lbildzinkas changed the title feat(engine): agent grouping pass with staged review updates feat(engine): agent grouping pass groups related hunks across files Oct 4, 2026
@lbildzinkas
lbildzinkas merged commit 385070d into master Oct 4, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Agent grouping pass, arriving in stages

1 participant