Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 12 additions & 0 deletions .perry/events.jsonl
Original file line number Diff line number Diff line change
Expand Up @@ -579,3 +579,15 @@
{"ts": "2026-08-20T23:35:10", "event": "evidence", "id": "TASK-122", "title": "the repair path the tools advertise leaves the file needing a whitespace fix", "track": "main", "actor": "agent", "from": "—", "to": "evidence/2026-08/TASK-122-spec.md"}
{"ts": "2026-08-20T23:37:40", "event": "status", "id": "TASK-140", "title": "every mode contract slot is assigned to an axis, and the spine-to-unit map is written down", "track": "main", "actor": "agent", "depends_on": [], "from": "in_progress", "to": "review", "reason": ""}
{"ts": "2026-08-20T23:37:40", "event": "evidence", "id": "TASK-140", "title": "every mode contract slot is assigned to an axis, and the spine-to-unit map is written down", "track": "main", "actor": "agent", "from": "evidence/2026-08/TASK-140-dispatch-2026-08-21.md", "to": "evidence/2026-08/TASK-140-dispatch-2026-08-21-result.md"}
{"ts": "2026-08-20T23:38:54", "event": "start", "id": "TASK-141", "title": "a row stays blocked after its blockers close, because the stored status masks the computed one", "track": "intake", "actor": "agent", "from": "not_started", "to": "in_progress"}
{"ts": "2026-08-20T23:38:54", "event": "evidence", "id": "TASK-141", "title": "a row stays blocked after its blockers close, because the stored status masks the computed one", "track": "intake", "actor": "agent", "from": "—", "to": "evidence/2026-08/TASK-141-spec.md"}
{"ts": "2026-08-20T23:52:06", "event": "status", "id": "TASK-120", "title": "the linkage edges are read but never folded into KR progress", "track": "main", "actor": "agent", "depends_on": [], "from": "in_progress", "to": "review", "reason": ""}
{"ts": "2026-08-20T23:52:06", "event": "evidence", "id": "TASK-120", "title": "the linkage edges are read but never folded into KR progress", "track": "main", "actor": "agent", "from": "evidence/2026-08/TASK-120-spec.md", "to": "evidence/2026-08/TASK-120-dispatch-2026-08-21-result.md"}
{"ts": "2026-08-20T23:52:24", "event": "add", "id": "TASK-144", "title": "the event log timestamp has no zone and the register has one, so ordering them is a guess", "track": "intake", "mode": "queue", "priority": "P1", "actor": "agent", "summary": "", "depends_on": [], "from": null, "to": "not_started"}
{"ts": "2026-08-20T23:52:24", "event": "add", "id": "TASK-145", "title": "the contract shape baseline is stale against its own recorder", "track": "intake", "mode": "queue", "priority": "P2", "actor": "agent", "summary": "", "depends_on": [], "from": null, "to": "not_started"}
{"ts": "2026-08-20T23:52:24", "event": "add", "id": "TASK-146", "title": "the viewer renders a KR current with no provenance because it does not go through the shared derivation", "track": "intake", "mode": "queue", "priority": "P2", "actor": "agent", "summary": "", "depends_on": [], "from": null, "to": "not_started"}
{"ts": "2026-08-20T23:53:36", "event": "start", "id": "TASK-143", "title": "two PRs each green on their own base merged into a red tree, and nothing checked the pair", "track": "intake", "actor": "agent", "from": "not_started", "to": "in_progress"}
{"ts": "2026-08-20T23:53:36", "event": "evidence", "id": "TASK-143", "title": "two PRs each green on their own base merged into a red tree, and nothing checked the pair", "track": "intake", "actor": "agent", "from": "—", "to": "evidence/2026-08/TASK-143-spec.md"}
{"ts": "2026-08-21T00:03:14", "event": "status", "id": "TASK-122", "title": "the repair path the tools advertise leaves the file needing a whitespace fix", "track": "main", "actor": "agent", "depends_on": [], "from": "in_progress", "to": "review", "reason": ""}
{"ts": "2026-08-21T00:03:14", "event": "evidence", "id": "TASK-122", "title": "the repair path the tools advertise leaves the file needing a whitespace fix", "track": "main", "actor": "agent", "from": "evidence/2026-08/TASK-122-spec.md", "to": "evidence/2026-08/TASK-122-dispatch-2026-08-21-result.md"}
{"ts": "2026-08-21T00:03:52", "event": "add", "id": "TASK-147", "title": "nothing outside describe_cell proves the table and bullet paths stay separated", "track": "intake", "mode": "queue", "priority": "P2", "actor": "agent", "summary": "", "depends_on": [], "from": null, "to": "not_started"}
12 changes: 8 additions & 4 deletions perry/BOARD.md
Original file line number Diff line number Diff line change
Expand Up @@ -36,18 +36,19 @@
| TASK-102 | Evidence becomes a typed relation: {path, kind, round}, not one prose cell | Coding Agent | not_started | — | — | V4 | TASK-090, TASK-092 | main | | | | | | |
| TASK-114 | aiMark reads Perry through the current contracts instead of a pin nine versions old | Coding Agent | in_progress | delegated to an aiMark coding agent; awaiting paste-back | evidence/2026-08/TASK-114-delegation-prompt.md | V4 | — | main | | | | | | |
| TASK-119 | the linkage graph is documented as machine-written and no tool writes it | Coding Agent | not_started | — | — | V3 | | main | | | | | | |
| TASK-120 | the linkage edges are read but never folded into KR progress | Coding Agent | in_progress | dispatched to claude-subagent; worktree pinned to 7c0bb99; state-schema.json scoped out so the gate passes without a release | evidence/2026-08/TASK-120-spec.md | V3 | — | main | | | | | | |
| TASK-120 | the linkage edges are read but never folded into KR progress | Coding Agent | review | PR #24 — contract 2.1; four findings handed back, two worth their own rows | evidence/2026-08/TASK-120-dispatch-2026-08-21-result.md | V3 | — | main | | | | | | |
| TASK-121 | the sweep that found four more live-state assertions runs once and then is thrown away | Coding Agent | not_started | — | — | V3 | | main | | | | | | |
| TASK-122 | the repair path the tools advertise leaves the file needing a whitespace fix | Coding Agent | in_progress | dispatched to claude-subagent; worktree pinned to 6c01b93; spec carries a live reproduction | evidence/2026-08/TASK-122-spec.md | V3 | — | main | | | | | | |
| TASK-122 | the repair path the tools advertise leaves the file needing a whitespace fix | Coding Agent | review | PR #25 — item 2 proved both ways on one run; two questions handed back | evidence/2026-08/TASK-122-dispatch-2026-08-21-result.md | V3 | — | main | | | | | | |
| TASK-123 | the goals writer takes the file as truth and derives the store, which is the opposite direction from the KR | Coding Agent | not_started | — | — | V4 | | main | | | | | | |
| TASK-126 | closing the dangling-id row requires writing the record that re-dangles it | Coding Agent | review | PR #22 — the suite is fully green; verify the strong anti-vacuity case survives review, then close at V3 | evidence/2026-08/TASK-126-dispatch-2026-08-21-result.md | V3 | TASK-112 | main | | | | | | |
| TASK-129 | Agent is five strings that do not join, and role has never once been written | Coding Agent | not_started | unblocked: work owns .perry/agents.jsonl → .perry/roles/ as of the 2026-08-20 signature; needs a spec, then dispatch | — | V3 | TASK-128 | main | | | | | | |
| TASK-135 | a track can be declared but no existing row can be moved onto it | Coding Agent | not_started | — | — | V3 | | main | | | | | | |
| TASK-136 | a queue track SLA is parsed, stored and never measured against anything | Coding Agent | not_started | — | — | V3 | | main | | | | | | |
| TASK-140 | every mode contract slot is assigned to an axis, and the spine-to-unit map is written down | Coding Agent | review | PR #23 — three open questions for the user, incl. whether an empty illegal-pair list discharges § 7 risk 2 | evidence/2026-08/TASK-140-dispatch-2026-08-21-result.md | V3 | — | main | | | | | | |
| TASK-141 | a row stays blocked after its blockers close, because the stored status masks the computed one | Coding Agent | not_started | — | — | V3 | | intake | triaged | | 2026-08-20 | | | |
| TASK-141 | a row stays blocked after its blockers close, because the stored status masks the computed one | Coding Agent | in_progress | dispatched to claude-subagent; worktree pinned to f42e84b | evidence/2026-08/TASK-141-spec.md | V3 | | intake | triaged | | 2026-08-20 | | | |
| TASK-142 | triage has no check for a row stranded by a process bug, and the one signal that fired was read as prose hygiene | Coding Agent | not_started | design question answered 2026-08-20: it belongs in conformance, which triage already reads at step 0.5 — not as a new triage feature | — | V3 | — | intake | triaged | | 2026-08-20 | | | |
| TASK-143 | two PRs each green on their own base merged into a red tree, and nothing checked the pair | Coding Agent | not_started | — | — | V3 | | intake | triaged | | 2026-08-20 | | | |
| TASK-143 | two PRs each green on their own base merged into a red tree, and nothing checked the pair | Coding Agent | in_progress | dispatched to claude-subagent; worktree pinned to a10f897 | evidence/2026-08/TASK-143-spec.md | V3 | — | intake | triaged | | 2026-08-20 | | | |
| TASK-144 | the event log timestamp has no zone and the register has one, so ordering them is a guess | Coding Agent | not_started | — | — | V3 | | intake | triaged | | 2026-08-20 | | | |

## P2

Expand All @@ -68,6 +69,9 @@
| TASK-132 | the parity check cannot see 23 keys because Perry own state leaves four collections empty | Coding Agent | not_started | — | — | V3 | | main | | |
| TASK-137 | a new queue row is born in the second stage, not the first | Coding Agent | not_started | — | — | V2 | | main | | |
| TASK-139 | a design back-reference lives in a cell the close path clears, so a finished design reports as never handed off | Coding Agent | not_started | — | — | V3 | TASK-102 | intake | triaged | 2026-08-20 |
| TASK-145 | the contract shape baseline is stale against its own recorder | Coding Agent | not_started | — | — | V2 | | intake | triaged | 2026-08-20 |
| TASK-146 | the viewer renders a KR current with no provenance because it does not go through the shared derivation | Coding Agent | not_started | — | — | V3 | | intake | triaged | 2026-08-20 |
| TASK-147 | nothing outside describe_cell proves the table and bullet paths stay separated | Coding Agent | not_started | — | — | V3 | | intake | triaged | 2026-08-21 |

## Cadence (recurring; doesn't consume P0 slots)

Expand Down
106 changes: 106 additions & 0 deletions perry/evidence/2026-08/TASK-120-dispatch-2026-08-21-result.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,106 @@
# TASK-120 — result

> Date: 2026-08-21 · Executor: claude-subagent · PR: https://github.com/ranjiao/Perry/pull/24
> Branch: `coding/task-120-kr-progress-provenance` · Cycle time: ~35 min
> `perry-goals/list` **2.0 → 2.1**, additive: four keys added, none removed or
> retyped. `schema/state-schema.json` **untouched** — verified by diff, and the
> hard stop was live, so this is a real avoidance rather than luck.

## The shape: Perry exposes the contradiction, it does not resolve it

One derivation in `bin/lib/__init__.py`, three sibling keys per KR, emitted by
both payloads. Verified on this repository:

```
P-O1.1 current 0.0 target 1.0
provenance state=asserted measured=false source=linkage-register
completion total 4 done 4 open 0
P-O2.2 current 0.0 target 0.0
provenance state=asserted measured=false
completion total 2 done 0 open 2
```

`P-O1.1` no longer reads as 0-of-1 progress: it reads as **an author's assertion
of 0 against four closed tasks**, and Perry resolves neither. `P-O2.2` can no
longer be read as met, because nothing claims the zero was measured.

**No `met` / `achieved` / `progress` / `ratio` key exists** — asserted as an
absence in the tests. No percentage is emitted anywhere: counts only, in their
own unit, so the tally cannot be misread as a metric value. That is the line the
spec drew and it held.

## What it rejected, and one of the rejections is the interesting one

- any ratio or percentage;
- a conformance entry for *"`current` disagrees with its tasks"* — that would be
Perry inferring the metric from the edges, which is the forbidden move one
step removed;
- **any new authored field in the register.** `asserted_at` reuses the
register's existing top-level `updated`, and `asserted_scope: "register"` is
emitted so a reader cannot mistake it for a per-KR date. **That choice is what
kept `schema/state-schema.json` out of the change** — the agent found the
cheap path around a gate rather than asking for a release.

## The tally is what flips P-O1.1, not the staleness check

Worth recording because it is counter-intuitive: P-O1.1 is **not stale by any
timestamp test** — all four of its tasks closed *before* the register's
`updated`. The contradiction is visible only because the completion tally sits
beside the number. A design that shipped staleness alone would have left that
KR reading exactly as wrongly as before.

## Staleness, both directions on one fixture

```
not stale "no linked task has changed state since 2026-08-15T12:00:00"
stale "1 linked task changed state after 2026-08-15T12:00:00:
TASK-003 (in_progress → done)"
moved_tasks: [{id, from, to, at}] · only that KR goes stale
```

The fixture also carries a `next` event dated after the assertion, **so a check
keyed on the event name rather than on whether `to` is a status would redden.**

## Parity, stated separately as instructed

`perry-goals/list` **before**: 0 documented-not-emitted, **5** emitted-not-documented.
**After**: 0 and **5** — the same five, byte-identical, TASK-131's, not hidden
inside. Documented 54 → 77, emitted 59 → 78. Repo-wide total unchanged at 17.

## Four findings handed back

1. **The `Z` problem, unresolved and documented in the contract.** The register
writes `updated` as ISO with a `Z`; `.perry/events.jsonl` writes `ts` as naive
local time. There is no honest conversion, so the `Z` is **stripped rather
than applied**, and a register written within a few hours of a task move can
order wrongly. Fixing it means deciding what the event log's `ts` means — a
row of its own, touching every consumer.
2. **`tests/fixtures/contract-shapes.json` is stale w.r.t. its own recorder.**
`--record` wants to add an `empty_lists` block to two contracts and drop a
trailing newline. The agent refused to re-record rather than hide unrelated
drift in this row. **Someone will eventually re-record it into an unrelated
diff.**
3. **`viewer/serve.py` renders the chain view from `viewer/parsers.py`**, not
through `bin/lib`, so the viewer still shows `current` with no provenance.
4. **`measured` is `false` everywhere, by construction, until something re-runs
a metric.** Honest — and it means `stale` is Perry's only mechanical opinion
about a KR number.

## Two notes on this project's own rules

- The `current: 0` default the spec targeted lives in
`goals/state/linkage_TEMPLATE.md:13`, which writes it into every new register.
That is **authoring**, belongs to TASK-119's writer, and was correctly not
touched. `_num()` already returned `None` for an unwritten value, so V3 item 2
was true but untested; it is now locked by four tests.
- **TASK-091's Definition of Done makes `bin/` a place where the history of that
defect cannot be written down.** The agent's first draft of a comment named the
deleted symbol and reddened `test_goals_writer`; it rewrote the comment to
describe the symbol without naming it. Same family as TASK-126 and TASK-112 —
a guard that forbids describing the thing it guards.

## Process error, mine

The worktree was cut from `feat/work-modes` **before** I committed the spec, so
the agent had to fetch it with `git checkout e71d7c0 -- <path>`. Corrected for
TASK-122 and TASK-141: commit the spec, *then* cut the worktree.
85 changes: 85 additions & 0 deletions perry/evidence/2026-08/TASK-121-spec.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,85 @@
# TASK-121 — the sweep that finds checks reading live state runs once and is thrown away

> Source: `perry/evidence/2026-08/TASK-113-dispatch-2026-08-20-1813.md`
> Dispatch mode: auto
> Executor: claude-subagent
> Estimated cycle: medium
> Subjective verification: no
> Touches architecture: no — it adds a guard over the test suite
> Deployed: no

## Schema

- **Owner**: Coding Agent
- **Priority**: P1
- **Attribution**: unlinked

## The class, and its eight known instances

A check that reads **the project living around it** as its expected value goes
red on ordinary progress, and green again for reasons that have nothing to do
with what it measures. Every instance below is real and dated within four days:

| # | check | what moved under it |
|---|---|---|
| 1–3 | `test_diagnose`, `test_one_line_break_rule`, `test_v5_signoff` | TASK-113 found and fixed three |
| 4 | a fourth, handed over mid-run | same row |
| 5 | a fifth the agent found itself — `DESIGN-900` | same row |
| 6 | `test_md_store § test_config_including_its_prose_section` | asserted every config record is a `setting`; **declaring one track reddened it** |
| 7 | `test_track_attribution § TestPerrysOwnProjectIsUnmoved` | asserted Perry itself has no track register; same declaration reddened it |
| 8 | `test_state_cost` ×2 | asserted `perry/tasks.jsonl` is unclaimed and `.perry/events.jsonl` rolls up under `.perry/` — **both true until PR #14 declared the two store files owned** |

TASK-113 fixed instances 1–5 **by hand, in one pass, and the pass was thrown
away.** Instances 6–8 arrived afterwards. There is no mechanism; there is a
memory of having looked.

## Deliverable

A guard that finds this class **mechanically**, so the next instance is reported
rather than discovered by a human running the suite after a merge.

**What "this class" is, precisely, is the hard part of this row** — and getting
it wrong in either direction makes the guard worthless:

- too broad, and it flags every test that reads a fixture, which is all of them;
- too narrow, and it is a list of the eight above wearing a regex.

The instances give you the shape to generalise from: each one asserted a
**literal about the project's current state** — a count, an id, a set membership,
a filename — where the *property* being tested was true independently of that
literal. Note that instance 8's literals were about **which paths the schema declares
Perry owns**, not about a board row — so a guard keyed only on `BOARD.md` or the
task store would have missed it.

Report what you decided the class is, in the guard's own docstring, in the voice
of the surrounding modules — and **name what it deliberately does not catch.**

## Verification — V3

1. **It finds instances it was not shown.** Reconstruct at least three of the
eight from git history — `test_md_store` and `test_track_attribution` before
their 2026-08-21 fixes, and one of TASK-113's — and show the guard flags them.
Reconstruct, do not hand-write an approximation.
2. **It does not flag the fixes.** The same three, after their repairs, are
clean. A guard that still flags the repaired form is measuring the wrong
thing.
3. **False-positive floor, stated as a number.** Run it over the whole suite as
it stands and report **every** hit. If the count is not zero, each survivor
is either a real instance — open a row for it — or a false positive you must
name and explain. **Do not silence one to reach zero.**
4. **Reverting the guard reddens its own test.**
5. `python3 tests/parallel -j 4`, `bash tests/run`, `python3 bin/perry-lint`,
`git diff --check`.

## Files in scope

- the guard, as a new test module or a check under `tests/`
- its own tests and fixtures

## Out of scope

- **Fixing any instance you find.** Report them; each is its own row. This row
ships the mechanism, not the repairs.
- `bin/perry-diagnose` and `tests/test_diagnose.py` — an unmerged branch (PR #22)
is editing both. Cutting across it would conflict.
- `perry/` — no project state changes; `git diff -- perry/` must end empty.
Loading
Loading