Skip to content

Declare what every operation destroys, sends and reads from strangers - #230

Merged
robzolkos merged 3 commits into
mainfrom
behavior-traits
Oct 2, 2026
Merged

robzolkos merged 3 commits into
mainfrom
behavior-traits

Conversation

@jeremy

@jeremy jeremy commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Today no operation in behavior-model.json carries a destructive trait, and nothing tells a consumer which operations send mail or return text written by someone else. So the MCP toolkit (basecamp/mcp) guesses from action names (catalog.BridgeDestructive), hey-mcp-server hand-curates empty_spam and empty_trash, and every HEY tool is annotated openWorldHint: false. That last one is wrong for sends, and it matters for the App Security review: private mail, plus sender-authored content reaching the agent, plus an outbound send, is the "lethal trifecta". The toolkit can only gate the send if it can tell which calls send.

This PR has every operation declare three things, each classified from haystack's controllers rather than from the verb or the name.

The traits (spec/hey-traits.smithy)

Smithy trait On OpenAPI behavior-model.json Meaning
@heyDestructive(bool) every write x-hey-destructive destructive (reads: false) Some path destroys data, or the caller's own access to it, and the caller can't get it back
@heyOpenWorld(bool) every write x-hey-open-world open_world (reads: false) The call can reach people outside the mailbox: it delivers mail, publishes to HEY World, or sends calendar cancellations
@heyDraftWhen([...]) open-world writes that can save a draft x-hey-draft-when draft_when: [{pointer, equals | not_equals}] Conditions on the request body under which the call delivers nothing. They're sufficient, not necessary
@heyUntrustedContent(bool) every operation x-hey-untrusted-content untrusted_content The response can carry text written by someone other than the caller

Naming. basecamp-sdk has no equivalent to reuse. Its scripts/gen-catalog/main.go records the destructive trait as a known gap ("That trait does not exist in the Smithy model yet"), and it has nothing for open-world or provenance. basecampSensitive is a member-level redaction trait and a different concept. So these follow basecamp-sdk's conventions instead: <product>X trait → x-<product>-x extension → a flat behavior-model key. The destructive key is exactly what mcp/catalog and basecamp-sdk's gen-catalog already read as a tri-state. If basecamp-sdk adopts basecampDestructive, basecampOpenWorld and basecampUntrustedContent, both SDKs emit the same keys.

Tripwire. EmitEachSelector validators at severity DANGER make each declaration mandatory. An operation added without one fails smithy validate, and with it make check and the Smithy CI job. One more validator forbids resending an open-world operation, because a retry after an ambiguous first attempt could deliver twice. It follows the rule every generator uses to decide a resend: x-hey-idempotent.natural when set, otherwise @readonly or @idempotent, otherwise the verb. Without natural: false, an open-world operation is refused if @idempotent, natural: true, or a GET/HEAD/PUT verb would resend it. DELETE is exempt because the cancellations HEY sends on a delete are keyed to the record it destroys, so a resend gets a 404.

A validator whose selector stops matching passes silently. To catch that, scripts/test-behavior-traits (make behavior-traits-test, in check-mvp/check-full and CI) breaks the model one declaration at a time and requires a refusal under the expected validator id and shape. It also checks that the unbroken model validates. generate-behavior-model independently refuses to emit an undeclared operation, so if a validator is edited away, a missing declaration still can't turn into false.

@heyDraftWhen requires at least one condition, because an empty list would hold vacuously and pass every send as a draft.

One existing value changes. generate-behavior-model used to count any @heyIdempotent as idempotent, so UpdateMessage (natural: false, it delivers) was advertised as idempotent: true. It now reads natural, and UpdateMessage is the only operation that flips. No generated SDK code changes: the generators read x-hey-idempotent.natural first and ignore the new keys and extensions.

Classification: 131 operations (76 writes, 55 reads)

Evidence is pinned to haystack 49cbde18.

Destructive (13):

  • EmptySpam, EmptyTrash: Topic#deleted! on every thread (spam, trash).
  • DeleteCalendarEvent, DeleteCalendarEventOccurrence: destroy / destroy! (event, occurrence). A shared event is removed for every member.
  • DeleteHabit (L49), DeleteTimeTrack (L48), DeleteCalendarTodo (L31), DeleteSticky (L19): hard destroy.
  • UpdateJournalEntry: empty content destroys the entry (Calendar::Days::JournalEntriesController#update).
  • DeleteContactNote: note: nil, and the text is gone (L22).
  • DeleteExtenzion: destroy_contactable_and_erect_tombstone plus access teardown (L85).
  • TrashPostings: judgment call, see below.
  • PuntClearances: judgment call, see below.

Open world (6):

  • CreateMessage, UpdateMessage, CreateReply: these deliver unless drafted (messages, replies). CreateMessage can also publish to HEY World (L163). All three carry draft_when.
  • CreateBulkReply: always delivers (L18).
  • DeleteCalendarEvent, DeleteCalendarEventOccurrence: when the caller organizes the event, these email a cancellation to every attendee who hasn't declined (notifier).

Untrusted content (45): mailbox and box reads, topics, entries, messages, drafts, search, folders, collections-with-postings, contacts and clearances (and the clearance/contact writes that echo them), calendar periods and recordings, calendars, clips, the reply/forward/bulk-reply compose prefills, and the workflow stage page.

Not destructive, noted: these are reversible.

  • TrashTopic and DeleteDraft: trashed!, restorable for 30 days (restore, window).
  • MarkPostingsSpam, MarkEntrySpam: MarkTopicHam undoes them. They do train rspamd.
  • HideContact (DELETE) ↔ RevealContact.
  • DeleteBoxGroup: moves the group's mail back to the Imbox (L16).
  • DeleteBoxDesignation: a rule you can recreate.
  • Clearance updates, including screening out: re-approving restores the threads.
  • Every un-/re- toggle.
  • Ordinary edits (UpdateMessage, UpdateContact, ...).

Judgment calls worth a look

  1. TrashPostings is destructive. For JSON the server treats the removal decision as made. On a shared thread, that revokes the caller's own access (accesses.destroy_by) instead of trashing it (L15, L38), and the caller can't undo that. Non-shared threads go to the restorable trash. MCP's hint means "may", so true.
  2. PuntClearances is destructive for the same reason. It trashes every pending sender's threads, but revokes access on shared ones (L12).
  3. Trash is not destructive here, and basecamp-sdk disagrees. The comment in basecamp-sdk's gen-catalog calls TrashRecording "genuinely destructive", and the toolkit's bridge treats a trash prefix as destructive. Once basecamp-sdk declares its own trait, the two SDKs should agree on this definition.
  4. draft_when requires /entry/scheduled_delivery ≠ "true" as well as /entry/status = "drafted". A drafted entry with a schedule is still delivered at the scheduled hour with no further call (L19), so a gate that trusted status alone would wave through a timed send.
  5. Some untrusted_content calls are conservative:
    • ListDrafts and GetMessageEdit: a reply draft carries the thread's subject and quoted text.
    • ListCalendars: subscribed feeds and shared calendars are named by others.
    • Contact writes (CreateContact, UpdateContact, RevealContact): they return a contact whose name and address its owner may have declared.
    • Own-record calendar reads and writes (habits, todos, journal, time tracks) are false. The shared Recording schema's third-party fields (organizer, attendances, attached_entry) belong to the Calendar::Event variant only.
  6. Not modeled, so not classified here. The SDK has no calendar event create/update, RSVP, or forward-send operation. Each of those emails others in haystack, and each will be refused by the validators until someone classifies it.

How the toolkit should consume this (not changed here)

In basecamp/mcp:

  • Read the new keys in catalog.behaviorTraits, tri-state like destructive:

    • OpenWorld *bool (open_world)
    • UntrustedContent *bool (untrusted_content)
    • DraftWhen []struct{ Pointer, Equals, NotEquals string } (draft_when, with not_equals on the wire)

    Absent means undeclared. That covers basecamp-sdk today, which should keep its current behavior.

  • destructive: nothing new to read. Every HEY operation now declares it, so BridgeDestructive never fires for HEY. hey-mcp-server must delete its DestructiveActions override for empty_spam/empty_trash when it vendors this model. The catalog loader refuses an override of a declared trait, by design. Some actions flip from bridged-true to declared-false: delete_draft, delete_box_group, delete_box_designation, trash_topic, remove_postings_from_box_group.

  • openWorldHint: set it per tool as AnyOpenWorld() over the served actions, instead of the hard-coded false in gateway.BuildMCPServer. This puts it on messages, entries, bulk reply and calendar events.

  • Delivery gate: in dispatch, treat an open_world action as a send unless every draft_when condition holds on the call's body. equals holds when the value is present and equal; not_equals holds when the value is absent or different. Hold sends to an explicit policy: off unless the operator enables sending (the way read-only filtering already trims writes), or behind a per-call confirmation or elicitation. Drafts pass, which keeps "the agent drafts, the human sends" frictionless.

  • Provenance: mark untrusted_content results as untrusted input before the model reads them, for example by delimiting the text and adding a _meta flag, and taint the session. Once a session is tainted, an open-world call needs the gate even where policy would otherwise allow it. That's the trifecta rule: untrusted content may be read, but it can't carry data out on its own.

Verification

  • make -k check on Linux: everything passes except cargo deny, which isn't installed on that host. CI's Rust job runs cargo deny.
  • scripts/test-behavior-traits: all eight cases, including the unbroken control.
  • Every generator re-run (smithy-build, url-routes, fingerprint, coverage, Go, Rust, TypeScript, Kotlin, Swift): only openapi.json and behavior-model.json change.

Copilot AI balanced review requested due to automatic review settings September 30, 2026 02:33
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-02T20:48:08.022110Z c323084 New commits
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Retry validation and generated idempotency metadata are inconsistent, and one destructive journal path is misclassified.

Review effort: Balanced
Findings: 2 High severity · 1 Medium severity

Open (3)
What changed in this PR

Adds explicit operation metadata for destructive effects, external delivery, draft conditions, and untrusted content, enabling safer MCP policy decisions.

Changes:

  • Defines and applies behavior/provenance traits across all operations.
  • Generates corresponding OpenAPI and behavior-model metadata.
  • Adds validation tripwire tests and CI enforcement.

[!TIP]
If you aren't ready for review, convert to a draft PR.
Click "Convert to draft" or run gh pr ready --undo.
Click "Ready for review" or run gh pr ready to reengage.

File Description
spec/​hey.smithy Classifies operation behavior and provenance.
spec/​hey-traits.smithy Defines traits and validation rules.
scripts/​test-behavior-traits Tests validator tripwires.
scripts/​generate-behavior-model Emits new behavior fields.
openapi.json Adds generated trait extensions.
behavior-model.json Adds generated policy metadata.
Makefile Integrates tripwire tests into checks.
.github/​workflows/​smithy-verify.yml Runs tripwire tests in CI.
AGENTS.md Documents trait conventions and workflow.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread scripts/generate-behavior-model
Comment thread spec/hey-traits.smithy Outdated
Comment thread spec/hey.smithy Outdated

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 38d9298891

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread spec/hey-traits.smithy
Comment thread spec/hey-traits.smithy Outdated
@jeremy

jeremy commented Sep 30, 2026

Copy link
Copy Markdown
Member Author

Review status at a2ad894

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved. Cursor Security Agent completed with no findings that need human review; Cursor Bugbot was not present, so that signal was skipped. No prior approval from this automation.

Open in Web View Automation 

Sent by Cursor Approval Agent: Pull Request Approver

@jeremy

jeremy commented Oct 2, 2026

Copy link
Copy Markdown
Member Author

Review status at c323084 (2026-10-02)

  • Rebased onto main (3 commits). Both commits on this branch replayed unchanged (git range-diff shows them identical). Main had gained ListAddressableContacts (Model the composer's recipient autocomplete: ListAddressableContacts #231) after this branch's HeyUntrustedContentUndeclared validator was written. GitHub showed the PR as mergeable, but the merged spec failed smithy-validate. c323084 declares that operation @heyUntrustedContent(true), because its labels are the display names correspondents chose for themselves (the same content ListContacts and GetContact already declare). openapi.json and behavior-model.json were regenerated.
  • Local make -k check: everything passes except one TypeScript release-workflow test, which fails on macOS only (it calls /bin/true, which macOS doesn't have). The Go, Rust, Kotlin and Swift checks, drift checks and the 196 conformance tests all pass.
  • CI: everything passes except GitHub Actions audit. It fails the same way on main: zizmor ref-version-mismatch on dtolnay/rust-toolchain at test.yml:138 and :172, because upstream moved v1 again (now 7e38f4b). Dependabot Bump the github-actions group with 3 updates #239 is the fix on main. I haven't pulled it into this PR, to keep the PR on scope.
  • Reviews: Codex reviewed c323084 and found nothing. Cursor approved. Copilot last reviewed 38d9298, and only a human can request a re-review.
  • Threads: all 5 are resolved.
  • For a human: the trash-semantics question from the earlier status comment is now item 7 of the Decisions comment on hey-mcp-server#4: https://github.com/basecamp/hey-mcp-server/pull/4#issuecomment-5961322847. The DestructiveActions override drop is an implementation step for hey-mcp-server when it vendors this model.

jeremy added 3 commits October 2, 2026 17:56
Three explicit declarations on every operation, emitted into openapi.json as
x-hey-* extensions and into behavior-model.json for consumers such as the MCP
toolkit:

- @heyDestructive (writes) -> destructive: a path that destroys data, or the
  caller's access to it, with no way back for the caller.
- @heyOpenWorld (writes) -> open_world: the call can deliver mail, publish to
  HEY World, or send calendar cancellations. @heyDraftWhen -> draft_when names
  the request-body conditions under which a send saves a draft instead.
- @heyUntrustedContent (all) -> untrusted_content: the response can carry text
  someone other than the caller wrote.

EmitEachSelector validators make each declaration mandatory and forbid
resending an open-world POST/PUT; scripts/test-behavior-traits breaks the model
on purpose to prove they fire, and runs in make check and the Smithy CI job.
…asure destructive

- behavior-model.json counted any @heyIdempotent as idempotent, so UpdateMessage,
  which opts out with natural: false, was advertised as safe to repeat. Only
  natural decides now; UpdateMessage is the one operation that changes.
- HeyOpenWorldRetried now mirrors how every generator decides a resend: an
  open-world operation that does not say natural: false is refused when
  @idempotent, natural: true, or a GET/HEAD/PUT verb would resend it.
- @heyDraftWhen needs at least one condition; an empty list holds vacuously.
- UpdateJournalEntry with empty content destroys the entry: destructive.

The tripwire test gains the dropped-opt-out and empty-conditions cases.
Main gained ListAddressableContacts (#231) after this branch's
HeyUntrustedContentUndeclared validator was written, so the rebased spec
failed smithy-validate. Its labels are the display names correspondents
declared for themselves, the same content ListContacts and GetContact
already mark untrusted.
@robzolkos
robzolkos merged commit 3ffb45b into main Oct 2, 2026
37 checks passed
@robzolkos
robzolkos deleted the behavior-traits branch October 2, 2026 22:09
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants