Skip to content

Truthful send states, hard channel contract, and outbox safeguards - #166

Open
MaxGhenis wants to merge 13 commits into
mainfrom
fix/truthful-send-pipeline
Open

MaxGhenis wants to merge 13 commits into
mainfrom
fix/truthful-send-pipeline

Conversation

@MaxGhenis

Copy link
Copy Markdown
Owner

Rebuilds the send pipeline's reporting and safeguards after the 2026-08-05→06 incident: send_message returned {ok:true, settled:true, state:"confirmed"} for a message that did not reach the recipient until ~15 hours later — seconds behind a manual day-of retry, double-texting the recipient — while send_to_conversation on a whatsapp:* JID returned a bare HTTP 404 even though get_status showed whatsapp connected:true and v2_send:true, and resolve_contact_routes said sendable:false. Three surfaces, three answers, and a success receipt that was a claim about the local ledger, not the world.

1. Truthful send states

Every durable send result now reports transport_state:

transport_state meaning
queued has not left this machine (queued/dispatching/not_dispatched)
transmitted the platform transport acknowledged it (remote message ID). Not proof of delivery
delivered a delivery/read receipt for it was observed in the local store
uncertain outcome unknown; reported as settled:false + uncertain:true (previously counted as settled)
failed / canceled terminal, nothing was sent

settled and transmitted are true only on transport acknowledgment (confirmed/store_failed). The "Message delivery confirmed" wording is gone everywhere (MCP text, CLI); transmitted results carry explicit "transport acceptance is not delivery" guidance. Results include the platform actually used (sms/whatsapp/signal — RCS is not distinguishable at this layer and is deliberately not guessed) and the conversation_id written to. The daemon's v1 delivery responses carry account_id/conversation_id/platform/expires_at_ms/expired so the MCP client can relay them.

2. Hard channel contract

  • The requested platform is enforced at send time against per-platform send capability (see §3): a hard-down platform (unpaired, adapter unregistered/receive-only, auth revoked) fails with the reason and queues nothing — there is never a fallback to another channel. A transient disconnect (queueable) still queues, truthfully reported and bounded by the TTL.
  • send_to_conversation / send_media_to_conversation gain an optional platform assertion that fails on mismatch (platform_mismatch) instead of sending on an unintended channel.
  • platform: "imessage" gets an explicit "import/read-only, cannot send" refusal.
  • The WhatsApp 404 is now self-explanatory: "the app could not resolve conversation X in its serving store … the connection can be up for receiving while this send path has no usable conversation record", instead of a bare HTTP 404: not found.

3. Status honesty

/api/status publishes a send block — send.{sms,whatsapp,signal} with available / queueable / reason — computed in the new internal/sendcap package from transport snapshots (paired, connected, needs_repair, auth_expired, phone_responding, needs_reauth) plus the v2 adapter registry (TextSend). get_status renders it in both serve modes and its text now states that connected/v2_send alone never imply a platform can send. resolve_contact_routes sendability, get_status, and send-time enforcement all read the same source (daemon truth in client mode — previously the MCP client judged sendability from its own transportless, always-disconnected process state), and routes carry a sendable_reason.

4. Outbox safeguards

  • Send window (TTL). New expires_at_ms on the outbox (migration 0011): an intent still queued when its window closes is canceled (error_class:"ttl", surfaced as expired:true / "NOT SENT") instead of transmitting stale. The lease query independently excludes expired rows, so a sweep/dispatch race can never transmit one. MCP sends default to 10 minutes (ttl_seconds per call, OPENMESSAGES_SEND_TTL_SECONDS per install, 0 = never expire); app/UI sends are unchanged (no TTL unless requested). Scheduled sends measure the window from their not_before time.
  • Near-duplicate guard. A text ≥0.75-similar (normalized Levenshtein; catches the incident's "tomorrow"→"today" edit, leaves "ok"/"ok!" alone) to one submitted to the same conversation within 10 minutes is refused (HTTP 409 / near_duplicate_blocked) naming the prior outbox item — unless force:true. Same-key replays (the documented lost-response retry) bypass the guard and hit idempotent dedup, and SendAgain (an explicit user action) is untouched.
  • list_outbox / cancel_outbox MCP tools (both serve modes): see everything still queued/retrying/uncertain, and stop a send before it crosses the transport boundary. During the incident there was no way to see or stop the overnight-queued message from the MCP surface.

5. MCP ergonomics

wait_for_transmit: true holds the call through auto-retrying states until transport acknowledgment or terminal failure, bounded by wait_seconds (default 25, max 120), so an agent can report truthfully in one call. Interrupted waits re-read the durable state on a detached context and keep the do-not-resend guidance.

Tests

83 new/updated assertions across the stack, including the three called out in the spec:

  • queued-not-transmitted must never report settled — TestDaemonQueuedSendNeverReportsSettledOrConfirmed (fake daemon stuck in queued; asserts settled:false, ok:false, transmitted:false, no "delivery confirmed" text) plus TestSendPayloadSettledMatrix pinning settled/transmitted/transport_state for all 8 outbox states.
  • duplicate guard — service-level matrix (blocked → forced; scope: different body / different conversation / short repeats / outside window; same-key replay bypass; similarity threshold table) and end-to-end through the MCP tool.
  • platform mismatch — TestSendToConversationPlatformAssertionMismatch, plus send-time enforcement (TestDaemonSendBlockedWhenPlatformCannotSend: hard-down refuses without submitting; queueable outage still queues).
  • TTL: stamping, scheduled-send windows, expiry sweep state scope (uncertain/post-transport rows never touched), lease-race exclusion, 15-hour incident replay (TestExpiredQueuedSendIsCanceledAndNeverDispatched).
  • sendcap tier matrix, /api/status send block, localapi wire round-trips, migration-gate update (staged-store pin now schema 11 + 0011 checksum).

go test ./... green locally (including the R5 migration integration suite); gofmt clean on all touched files.

Notes for review

  • settled semantics changed for uncertain (was true, now false + uncertain:true): an unknown outcome is not a settled one; reporting it settled is what invited "treat as done" during the incident. not_dispatched keeps settled:false/auto_retry:true.
  • Queueable vs hard-down: unconditional send-time blocking on any unavailability would make the durable outbox useless for its core purpose (transient disconnects). The tiering — refuse what won't self-heal, queue-with-TTL what will — is the incident-shaped compromise: the overnight message would have canceled at ~11:10pm instead of double-sending at 2:18pm, and the 404'd WhatsApp path now refuses up front with the real reason.
  • The in-process (standalone daemon) send path now bounds its wait at 25s like the client path did, instead of waiting indefinitely on the request context; both paths share one result builder (internal/tools/send_result.go).
  • Older daemons without the send block are treated as unknown capability (never blocked on), and their missing delivery fields degrade gracefully.

🤖 Generated with Claude Code

After the 2026-08-05 incident — a send reported ok/settled/confirmed sat
~15 hours before transmitting, a manual retry double-texted the
recipient, and WhatsApp sends 404ed while status said connected:

- transport_state queued/transmitted/delivered/uncertain/failed/canceled
  in every durable send result; settled and transmitted are true only on
  transport acknowledgment; results carry the platform actually used and
  the conversation_id written to; "delivery confirmed" wording removed
- per-platform send capability (new internal/sendcap) published at
  /api/status "send", rendered by get_status in both serve modes, read
  by resolve_contact_routes, and enforced at send time: hard-down
  platforms (unpaired, adapter unregistered, auth revoked) refuse
  without queuing; transient disconnects queue with truthful reporting
- outbox send window (migration 0011, expires_at_ms): a send still
  queued when its window closes cancels as expired instead of
  transmitting stale; MCP sends default to 10 minutes (ttl_seconds,
  OPENMESSAGES_SEND_TTL_SECONDS); the lease query excludes expired rows
  so a race can never transmit one; cutover carries windows forward
- near-duplicate guard: a text >= 0.75-similar (normalized Levenshtein)
  to one submitted to the same conversation within 10 minutes is
  refused, naming the prior intent, unless force=true; same-key replays
  keep idempotent dedup; SendAgain untouched
- list_outbox / cancel_outbox MCP tools in both serve modes
- wait_for_transmit + wait_seconds hold the call through auto-retrying
  states until transport acknowledgment
- explanatory 404s (conversation unresolvable in the serving store),
  optional platform assertion on send_to_conversation, and an explicit
  imessage cannot-send refusal

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@MaxGhenis
MaxGhenis force-pushed the fix/truthful-send-pipeline branch from 964844b to 3d942e2 Compare September 25, 2026 14:56
MaxGhenis and others added 12 commits September 29, 2026 05:22
The Signal status only carried needs_reauth / upgrade_required booleans,
so nothing outside the supervisor could tell the self-retesting
signal_account_unreadable park (ambiguous local evidence) from a
server-confirmed signal_account_invalid park or an upgrade gate. Send
capability therefore had to treat every needs_reauth as hard-down with
re-pair advice, even for the park the app re-probes on its own.

- StatusSnapshot gains park_fingerprint: the terminal exit fingerprint
  behind needs_reauth or upgrade_required, set at every park site (probe
  classification, receive account-invalid, version gate, poison envelope,
  ApplyPollerFailure) and cleared wherever both flags clear. An empty exit
  fingerprint records signal_park_unspecified, so the field is non-empty
  exactly when a park flag is set. parkUpgradeRequired now takes the
  fingerprint its caller returns. Lifecycle transitions are otherwise
  byte-identical.
- signallive exports ParkRetestedAutomatically and ParkRetestInterval;
  the cmd park retest uses both, so "re-checked automatically" in a send
  reason and the retest itself read one predicate and one cadence.

Tests: a differential test drives every park through StartPoller and
checks status.park_fingerprint == the PollerExit fingerprint, that a
retest generation in flight shows no park, and that a healthy generation
clears it; a testing/quick property checks the fingerprint/flag invariant
and a reference model over arbitrary ApplyPollerFailure sequences; the
supervisor retest test is table-driven over every fingerprint and class.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
sendcap turned every Signal needs_reauth into hard-down with re-pair
advice, including the ambiguous signal_account_unreadable park the
supervisor re-probes on its own, while an upgrade_required park fell
through to queueable "disconnected" although it never heals without an
upgrade. Separately, route discovery against a daemon without the send
block invented SMS availability and called connected WhatsApp/Signal
unsendable, disagreeing with send-time enforcement (which treats that
daemon as unknown), and the typed information any fix adds was dropped
on the daemon hop because the client copied three fields by hand.

- sendcap: a typed Condition on every non-available capability, a static
  condition-to-tier table, TierOf, and Classify(capability, known) as the
  one classifier behind enforcement and route discovery. Signal order:
  adapter_missing, not_paired, upgrade_required (hard; upgrade then
  reconnect or restart), needs_reauth whose park fingerprint
  signallive.ParkRetestedAutomatically accepts -> account_recheck
  (queueable; the reason says the park may be transient or may need a
  re-link, that the app re-checks about every 15 minutes, and that a send
  waits until Signal recovers or its send window closes, if it has one),
  any other needs_reauth including empty/unknown fingerprints ->
  relink_required (hard, fail-safe), disconnected (queueable), available.
- localapi.PlatformSendCapability mirrors condition; the MCP client
  converts daemon entries with capabilityFromDaemon so nothing is dropped;
  refusals carry condition in the structured payload.
- resolve_contact_routes: every route reports send_capability
  (available/queueable/unavailable/unknown/read_only). An older daemon
  yields unknown for every send platform, never invented availability;
  the preferred route falls back to an unknown SMS route only when no
  route is known-available. Text labels distinguish queues-only,
  unavailable, unknown, and history.
- get_status (in-process and daemon-routed) prints needs_reauth,
  upgrade_required, park_fingerprint, and each capability's condition,
  and says so explicitly when the app predates the send block.
- tools.Options.TransportsEnabled: serve sets it from its transports flag
  so in-process capability agrees with the process's own /api/status
  under --no-transports; Register and option-less handlers keep the
  historical transports-on assumption.

Invariants (executed by tests): well-formed tiers and condition/tier
agreement for every input; exact key set and determinism; per-platform
isolation; Signal tier mapping as above; fault monotonicity;
account_recheck iff the supervisor retests the park; route unavailable
iff enforcement refuses, and unknown never reported available; the
/api/status send block round-trips losslessly. Exhaustive enumeration
covers 516,096 inputs; testing/quick covers random fingerprints, daemon
statuses, and malformed entries.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
LeaseDue refuses rows whose send window has closed, but DispatchDue
handles a leased batch one item at a time and re-checked only the lease
deadline. A later item whose window closed while an earlier item sat in
its transport call (or during a slow bridge acquisition) was still handed
to the transport after expiry. The minimized case: two sends, the first
takes 2s, the second has a 1s window; its transport call started 1000 ms
late. The expiry sweep cannot help, because it deliberately leaves
dispatching rows to their lease owner.

MarkTransportCalled is now the send-window gate. In one transaction it
commits transport_called_at_ms only while the lease is live and
expires_at_ms IS NULL OR expires_at_ms > now (the same instant LeaseDue
and CancelExpired use). If the caller still owns the uncalled row but
its window has closed, the same transaction cancels it as expired
(error_class ttl / send_window_expired, lease cleared, attempt_count
unchanged) and returns the new sqlite.ErrSendWindowExpired.

The four identical marker blocks in the dispatcher (text, media,
reaction, read receipt) become one helper, crossTransportBoundary. It
treats the sentinel as a normal per-item outcome, so the batch continues
and DispatchDue returns no error. It also fixes an adjacent pre-existing
liveness bug: a lease that expired between DispatchDue's per-item check
and the marker returned ErrLeaseLost as a dispatch error, which stops
Run, and nothing restarts the dispatcher. When the clock shows the
lease expired, the helper now recovers leases the way the per-item check
does and continues; a lease lost while still live stays an error.

Reactions and read receipts now stamp expires_at_ms from
CommonCommand.TTL, which is documented for every kind but was silently
ignored for these two (no production caller sets it today). SendAgain
keeps no window, now documented and pinned by a test: it is an explicit
UI outbox-tray action, UI sends carry no window, and copying the
predecessor's absolute expiry would usually give a row born expired.

Invariants, each executed by a test (property tests use testing/quick
with fixed seeds and boundary-biased generators):
- I1 transport_called_at_ms commits at t only if the window is open at
  t, and the transport call follows that commit
  (TestQuickMarkTransportCalledGateOutcomes,
  TestQuickNoTransportCallAtOrAfterExpiry[WithSlowAcquire]).
- I2 gate totality: nil iff token matches, lease live and window open;
  ErrSendWindowExpired iff token matches, row uncalled and window closed
  (row canceled/ttl, lease cleared, attempts unchanged); otherwise
  ErrLeaseLost with the row byte-for-byte unchanged
  (TestQuickMarkTransportCalledGateOutcomes).
- I3 expiry never modifies a row that may have been sent: a called row
  or one in uncertain/confirmed/store_failed/rejected
  (TestQuickExpiryNeverTouchesPossiblySentRows, TestCancelExpiredScope).
- I4 a closed window, or a lease expiring at the boundary, never makes
  DispatchDue return an error (TestLeaseExpiryAtTransportBoundaryIsNotFatal
  and the slow-acquire property).
- I5 each windowed message ends confirmed after exactly one transport
  call started before its window closed, or canceled as expired with
  zero calls (the batch properties).
- I6 LeaseDue, the gate and CancelExpired split time at the same instant
  (TestSendWindowBoundaryIsConsistentAcrossLeaseGateAndSweep, exhaustive
  over expires_at-3..+3 ms).

With the base MarkTransportCalled swapped back in, the batch property
fails after 9 generated cases (a call 2430 ms late) and the slow-acquire
variant after 5 (8403 ms late); the gate property, the boundary test and
the per-kind tests fail too. Without the liveness branch, the
slow-acquire property fails I4 after 29 cases.

No schema or migration change: expires_at_ms already exists (0011).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Each TTL surface converted before it checked, so out-of-range input
could silently become a different window:

- HTTP ttl_ms checked only the sign before multiplying into a Duration.
  2^58 ms times 10^6 ns/ms is 0 mod 2^64, so ttl_ms=288230376151711744
  became "never expire"; 2^58+1 became a 1ms window; larger values
  became negative windows rejected with a misleading "TTL is negative".
- MCP ttl_seconds and OPENMESSAGES_SEND_TTL_SECONDS converted the float
  to a Duration before the 24h check. Out-of-range float-to-integer
  conversion is implementation-defined; observed on this Mac: arm64
  saturates (NaN gives 0, so a NaN env value disabled the default
  window), and amd64 (Rosetta) gives math.MinInt64 ns for 1e10, 1e300,
  NaN and Inf, which the daemon client then drops (ttl > 0 is false), so
  the send never expires.
- Positive sub-millisecond windows were born expired in-process but
  never expired via the daemon (ttl_ms truncated to 0), and fractional
  milliseconds were truncated only on the daemon path, so the two paths
  stamped different windows.

internal/messaging/ttl.go adds MaxTTL (24h), ValidateTTL, and
TTLFromMilliseconds / TTLFromSeconds, which reject NaN, +-Inf, negative,
over-24h, and positive sub-millisecond input before any conversion or
multiplication, and round an accepted seconds value to whole
milliseconds. Rejections are *InvalidTTLError wrapping ErrInvalidCommand,
with a reason each surface prefixes with its field name.

- HTTP ttl_ms (JSON and multipart routes) uses TTLFromMilliseconds and
  answers 400 "ttl_ms must be between 0 (no expiry) and 86400000".
- MCP ttl_seconds and the env override use TTLFromSeconds; the agent
  wording is kept ("ttl_seconds must not exceed 86400 (24 hours)") and
  the tool description states the range, millisecond rounding, and
  enforcement up to the hand-off to the transport.
- expiryMilliseconds runs ValidateTTL as defense in depth for every
  kind, so "never born expired" now holds; SendMedia validates before
  it writes the blob.
- parseSendWaitOptions had the same bug class: wait_seconds=1e10 gave a
  negative wait on amd64. It now rejects NaN/Inf and clamps in float64
  before converting.

Invariants, each executed by a test (testing/quick, fixed seeds,
generators biased to 0, 0.001, 86400, NaN, +-Inf, 1e300, subnormals,
2^58 and random bit patterns):
- I7 any accepted positive ttl_ms or ttl_seconds is 1ms..24h on every
  GOARCH; seconds results are ms-aligned and within 0.5ms of the
  request; 0 means no expiry; rejection is exactly NaN/Inf/negative/
  over-max/positive-sub-ms (TestQuickTTLFromMillisecondsBounded,
  TestQuickTTLFromSecondsBounded, TestQuickParseSendTTLBoundedOnEveryPlatform).
- I8 an accepted TTL > 0 gives expires_at_ms >= scheduled_for_ms + 1 for
  every sub-ms phase of the schedule (TestQuickAcceptedTTLNeverBornExpired).
- I9 differential: the daemon path (Milliseconds() -> ttl_ms ->
  TTLFromMilliseconds) and the in-process path stamp the same window
  (TestQuickDaemonAndInProcessTTLAgree in messaging and tools, plus
  TestDaemonSendCarriesMillisecondExactTTL end to end).

With the base parsing swapped back in, the tools and web tests fail on
arm64 (0.0005 s reached the daemon as ttl_ms=0; 1e-10 and NaN parsed as
0; 2^58 ms parsed as 0) and under GOARCH=amd64 (1e10 s parsed as
-2562047h; wait_seconds=1e10 gave the same negative wait). The new tests
pass on both architectures.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
sqlite.Open is the owner's open: it always takes SQLite's write
reservation and applies every pending migration. Clients that only read a
store another process owns (the transportless MCP client, openmessage read
and status) used it too, so a client built with migration 0011 migrated a
schema-10 daemon's live store to 11, and the older daemon then refused the
store on restart. The same BEGIN IMMEDIATE also made client startup fail
with SQLITE_BUSY whenever the daemon held a write transaction past the 5 s
busy timeout.

sqlite.OpenReadOnly attaches with mode=ro plus query_only and no _txlock,
never sets journal_mode, never creates a missing file, and never writes
the main database file. Inside one read snapshot it checks the ledger with
classifyLedger and serves stores from MinClientReadSchemaVersion (10, the
reactions table the client reads needs) through the build's latest
version. A newer store fails with ErrSchemaNewer, an older one with
ErrSchemaTooOld, and a ledger no build writes with ErrLedgerMismatch. When
SQLite cannot create the WAL index in a read-only directory, the error is
ErrReadOnlyAttach with remediation text. IsReadOnlyError recognizes the
SQLITE_READONLY refusal every write through such a handle gets.

validateDatabaseState now delegates to the same pure classifyLedger (with
a minimum of 1), so the owner and client rules cannot drift. The owner's
error texts are unchanged; they additionally satisfy errors.Is for the new
class sentinels. Store gains ReadOnly() and SchemaVersion().
testsupport.go adds BuildStoreAtVersion (a new store at an older schema,
refusing existing paths) and ReadStoreFingerprint (hash, pragmas, ledger,
logical dump) for fixtures in other packages.

Invariants, each executed by a test:
- A read-only session leaves user_version, the ledger, PRAGMA
  schema_version, every row, and the at-rest main file unchanged; the
  only files it may add are the -wal/-shm sidecars
  (TestOpenReadOnlyNeverMutatesStore, testing/quick, plus example tests).
- Every write through a read-only handle fails with IsReadOnlyError, and
  each write in that table succeeds on a writable handle
  (TestReadOnlyWritesSucceedOnWritableHandle keeps that non-vacuous).
- classifyLedger accepts exactly the clean prefixes in [min, len(known)]
  and otherwise returns exactly one predicted class sentinel
  (TestClassifyLedgerProperty, testing/quick).
- The client path and the owner's validateDatabaseState agree on every
  stored ledger (TestClassifyLedgerAgreesWithValidateDatabaseState,
  differential).
- Every version in [MinClientReadSchemaVersion, latest] serves the full
  client read inventory, and version min-1 does not
  (TestClientReadInventoryAcrossSupportedVersions); the inventory equals
  v2read's actual store calls (TestClientReadInventoryCoversV2Read).
- Reads through a read-only attach equal reads through the owner's handle
  on a byte copy, including an older store against its migrated copy
  (TestReadOnlyAttachMatchesWritableReads, differential over every
  ReadSource method).
- The attach never waits for the owner's write lock and sees its later
  commits; it recovers a crashed WAL without touching the main file
  (TestOpenReadOnlyBesideLiveWriter, TestOpenReadOnlyAfterCrashRecoversWAL).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The transportless MCP client (serve --mcp-stdio) and openmessage
read/status opened the daemon-owned v2 store with the migrating
sqlite.Open. Both now go through openV2ReadStore, which uses
sqlite.OpenReadOnly. The daemon becomes the only process that migrates
the v2 store:
- A newer client serves an older store (down to schema 10) and leaves
  it unchanged.
- An older client refuses a newer store at startup. The error names the
  executable to update, since MCP hosts often run a different openmessage
  than the app.
- Client startup no longer waits behind the daemon's write lock.
The client logs store_schema_version and build_schema_version. The
existing refusal when the v2 store file is missing is unchanged.

runServeMCPClient is split into newMCPClientServer (builds the server and
returns a cleanup) plus ServeStdio, so cmd tests can drive the client's
tools in-process.

The legacy messages.db open (app.NewClient) is deliberately unchanged in
this commit and still runs db.New's schema migrate step; the comments now
say so instead of implying that open is write-free.

Regression tests (cmd):
- TestMCPClientOpeningOlderV2DoesNotMigrate: a schema-10 fixture serves
  list_conversations, get_conversation, get_messages, search_messages,
  get_person_messages and list_outbox through the client's MCP tools.
  After the session, user_version (10), the 10-row ledger, PRAGMA
  schema_version, the logical dump and the main-file hash are unchanged;
  the fake daemon saw only the expected GETs. A contrast leg shows
  sqlite.Open does migrate a copy. With the base behavior (sqlite.Open)
  restored, this test fails with user_version 10 -> 11.
- TestRunServeMCPClientRefusesNewerV2Store and the newer-store leg of
  TestOpenCommandReadSourceV2DoesNotMigrate: refusal with
  sqlite.ErrSchemaNewer and "update this openmessage binary", store
  unchanged.
- TestOpenCommandReadSourceV2DoesNotMigrate: read/status on schema 10
  leave the store unchanged (fails on the base behavior).
- TestMCPClientStartsWhileDaemonHoldsWriteLock: startup completes while
  another connection holds BEGIN IMMEDIATE (the base behavior fails after
  the 5 s busy timeout with SQLITE_BUSY).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The guard used to read recent intents on the connection pool before the
enqueue transaction began, skipping only candidates that shared the
submission's idempotency key. Two consequences, both reproduced:

- A replay of an existing key could be refused by a sibling under another
  key (a forced resend, a SendAgain successor, or a same-body send accepted
  after the window), and MCP rendered that refusal as "NOT QUEUED" for a
  send that was durably queued.
- Concurrent near-duplicates with new keys all passed the guard before any
  of them inserted (8 of 8 accepted in one process; both accepted in 40 of
  40 rounds across two store handles on one file).

The check now runs inside enqueue's write transaction, after the
idempotency key is resolved and only when the insert created a new row
(sqlite.NearDuplicateCheck via EnqueueOutgoingMessageGuarded). A replay
returns the stored intent, or conflicts on a changed payload, before the
check is consulted. Every store connection begins writes with BEGIN
IMMEDIATE, so the candidate read and the insert hold SQLite's write
reservation together and at most one of a set of concurrent near-duplicates
is accepted, across goroutines and processes. A refusal rolls back and
writes nothing, so the refused key stays usable with force.

Guard scope is now an explicit per-command bool,
messaging.CommonCommand.GuardNearDuplicates (zero value unguarded, so a
caller that does not decide keeps the pre-#166 behavior rather than failing
every send). Entry points follow d472 Variant A: MCP in-process and MCP
daemon mode guard; `openmessage send` guards and gains --force (agents shell
out to it); HTTP guards only when the request carries
guard_near_duplicates, so the web UI never blocks a person's "ok" or a
"there at 7" -> "there at 8" correction. httpDuplicateGuard is the single
switch for a later Variant B. SendAgain, media, reactions and read receipts
are never guarded; media captions are never candidates.

The daemon's near-duplicate 409 is now structured: error (keeps the
"near-duplicate" wording older clients match on), error_kind
near_duplicate_blocked, duplicate_of_outbox_id, duplicate_state and
duplicate_age_ms. localapi.AsNearDuplicateRejection trusts error_kind and
falls back to the wording for older daemons; MCP daemon mode now returns the
same fields as the in-process path. The CLI no longer tells the operator to
use a new key for a near-duplicate 409 (a new key would be refused again);
it names the prior outbox and says to rerun with --force. "Use a new key"
remains only for idempotency conflicts.

Similarity is factored into guardSimilarity; textsNearDuplicate keeps its
verdicts exactly and adds a length prefilter (the edit distance is at least
the length difference) to cut work done under the write lock. The d472
defaults are unchanged: 10-minute window, 0.75 threshold, 8 candidates.

Invariants, each executed by a test:
- Replay-first: replaying a stored key with an identical payload returns
  the stored intent (same outbox ID, deduplicated) and never
  ErrDuplicateSend, whatever else was submitted.
- Conflict preserved: a stored key with a different payload is
  ErrIdempotencyConflict and writes nothing.
- Refusal is side-effect free: outbox and messages rows are unchanged and
  the key stays unused.
- Mutual exclusion: of concurrent guarded near-duplicates with new keys,
  exactly one is accepted and the rest name it.
- Scope: the guard runs iff GuardNearDuplicates && !Force (window > 0).
- Candidate set: same account and conversation, text only, created within
  the window (inclusive), not rejected or canceled, newest 8, never itself.
- Similarity: symmetric, in [0, 1], 1 on equal non-empty normalized input,
  0 on empty input, invariant under whitespace and ASCII case; Levenshtein
  is a metric within its length bounds; the prefilter is exact.

Tests: testing/quick properties with fixed seeds (arbitrary-history replay,
a differential against a reference model of the guard, similarity and
Levenshtein), deterministic concurrency tests (a barrier clock that lines
all parties up at BEGIN, and two store handles), storage-level candidate and
rollback tests, HTTP scope and 409-body tests, localapi wire tests, MCP
in-process and daemon-mode tests, CLI --force/409 tests, and a cmd-level
scope matrix driving MCP daemon mode, the CLI, the UI request shape and
media through their real client code against one real daemon. The replay,
conflict and concurrency tests fail against an emulation of the previous
pre-enqueue guard.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Area A: enforce the send window at the transport boundary and validate every TTL input through one bounded converter.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Area D: read-only, migration-free v2 store open for the MCP client and read/status.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Area E: run the near-duplicate guard inside the enqueue transaction.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Area G: publish the Signal park fingerprint, tier Signal send capability by park, and report unknown capability for older apps.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Area A's send-window gate (MarkTransportCalled) cancels a leased, uncalled
text whose window closed; area E's enqueue-time near-duplicate check skips
rejected and canceled candidates. Neither branch tested the two together.

Add two storage tests at that seam. A text canceled at the gate was
provably never sent, so a guarded resend of the same body is accepted and
the check sees no candidates. A text that crossed the boundary
(transport_called_at_ms committed) may have been sent, so it keeps blocking
a guarded near-duplicate after its window closes: CancelExpired leaves it
dispatching and the resend is refused with the prior named.

Test-only; no behavior change. Both tests fail under a mutation that lets
canceled rows be candidates or lets the expiry sweep take a called row.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant