You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Report when sandbox provider changes are applied #3390
As an operator or automation author managing providers attached to running sandboxes, I want to know when a specific provider change has reached installed credentials, effective policy, and the environment for new processes, so that I can launch a client or finish a detach without guessed delays.
Problem Statement
A successful provider attach, update, or detach response confirms that the gateway saved a mutation. It does not establish that the running sandbox installed the matching credentials, activated the matching policy, and published the matching environment at its workload process boundary. Callers need a completion result tied to the exact change they requested.
Impact / Why This Matters
Operators currently combine mutation responses, logs, sleeps, and application retries. A process launched too early may receive a missing or previous credential reference. A request immediately after detach may precede actual revocation. Longer sleeps add delay without distinguishing a slow installation from a failed one.
One provider update can affect multiple sandboxes that install it at different times. A single success response hides pending or disconnected targets, while application retries can hide a failed first request after apparent completion.
Proposed Design
Extend the shared configuration-operation model owned by #1731 with provider-specific targets and completion evidence. Follow the structured-error conventions owned by #3051, while keeping mutation admission/replay separate from installation observation. The provider extension confirms installed credentials, active policy, and the environment for new processes for the exact requested change.
The operator workflow is:
Attach a provider, receive the saved change identity, and wait within a chosen deadline. Launch client A after the matching installation completes.
Update the provider once and wait for that exact change. Launch client B after completion. B's first request uses the new credential.
Detach the provider and wait for acknowledged revocation. Previously issued references can no longer resolve through the baseline credential path, and future process environments omit the detached credential.
Client A retains the baseline's revision-scoped reference semantics after an update; it is not silently restarted. Readiness does not promise that A's old reference resolves the replacement key or prove that an old upstream key can be retired. Requests already forwarded may finish. Detach does not revoke credentials at the upstream service.
Attach/update completion requires matching evidence from the current authenticated supervisor for installed credentials and active policy, plus acknowledgment from the workload boundary for the environment used by future launches. Saving desired state alone cannot establish completion. A separate supervisor deployment must satisfy the same conditions.
For shared providers, freeze the affected sandbox identities for the mutation and return individual results under one overall deadline. A newer mutation cannot satisfy an older wait. Missing capabilities, stale or missing observations, disconnection, failed installation, rejected policy, and missing boundary acknowledgment remain incomplete with safe reason categories. A timeout leaves the saved mutation in place and does not replay it. If a mutation was saved but recording its result fails, report uncertainty explicitly without claiming rollback or safe retry.
Distinguish immutable operation history from current readiness. An operation may retain an applied result after a later disconnect or credential expiry, while a current readiness query reports that the target is no longer ready. Historical success cannot override current session, policy, credential, or launch-environment checks. Repeated status queries for the same unchanged target should recover the same observation without continually adding durable records.
CLI and Rust SDK callers can inspect an exact change and wait for it with a bounded deadline. Current authorization applies to status lookup, and operation records contain only safe identities and closed reason categories. This extension adds no external-stable profile field, new reference format, resolver retargeting, or dependency on #3339.
Acceptance Criteria
A secret-free result identifies the exact requested change and distinguishes saved intent from installation.
Attach/update completion requires matching credential state, active policy, and an acknowledged environment for new processes, including separate supervisor deployments.
With an ordinary static provider: attach → wait → launch A; update once → wait → launch B; B's first request uses the updated credential. A retains baseline reference semantics.
Detach → wait confirms baseline reference revocation and removal from future process environments; an independent reachable backend control excludes network outage as the reason a credentialed request fails.
Replacement, concurrent updates, reconnects, stale acknowledgments, same-revision repair, missing capabilities, and expired observations cannot falsely complete a change.
Shared-provider results preserve the original target identities; every selected target receives an outcome under one deadline that includes RPC time. Timeout does not replay the mutation.
Current authorization applies to status lookup. Later resource state cannot be mislabeled as the original operation's result, and historical applied state cannot override a current incomplete result.
Repeated status queries for an unchanged target preserve the original observation identity; receipt-less queries for unknown providers create no operation record.
Diagnostics and evidence expose no credential values, references, headers, or secret-bearing exception text.
Alternatives Considered
Fixed delays and application retries: they cannot prove installation or identify the failed component. Restart every client after every mutation: this disrupts workloads and still requires readiness before launching the replacement. Require stable credential references: this couples installation observation to a separate credential-lifecycle design. Create a separate general operation engine: this duplicates #1731 and #3051 and creates incompatible completion and retry semantics.
Agent Investigation
No response
Checklist
I've reviewed existing issues and the architecture docs
This is a design proposal, not a "please build this" request
User Story
As an operator or automation author managing providers attached to running sandboxes, I want to know when a specific provider change has reached installed credentials, effective policy, and the environment for new processes, so that I can launch a client or finish a detach without guessed delays.
Problem Statement
A successful provider attach, update, or detach response confirms that the gateway saved a mutation. It does not establish that the running sandbox installed the matching credentials, activated the matching policy, and published the matching environment at its workload process boundary. Callers need a completion result tied to the exact change they requested.
Impact / Why This Matters
Operators currently combine mutation responses, logs, sleeps, and application retries. A process launched too early may receive a missing or previous credential reference. A request immediately after detach may precede actual revocation. Longer sleeps add delay without distinguishing a slow installation from a failed one.
One provider update can affect multiple sandboxes that install it at different times. A single success response hides pending or disconnected targets, while application retries can hide a failed first request after apparent completion.
Proposed Design
Extend the shared configuration-operation model owned by #1731 with provider-specific targets and completion evidence. Follow the structured-error conventions owned by #3051, while keeping mutation admission/replay separate from installation observation. The provider extension confirms installed credentials, active policy, and the environment for new processes for the exact requested change.
The operator workflow is:
Client A retains the baseline's revision-scoped reference semantics after an update; it is not silently restarted. Readiness does not promise that A's old reference resolves the replacement key or prove that an old upstream key can be retired. Requests already forwarded may finish. Detach does not revoke credentials at the upstream service.
Attach/update completion requires matching evidence from the current authenticated supervisor for installed credentials and active policy, plus acknowledgment from the workload boundary for the environment used by future launches. Saving desired state alone cannot establish completion. A separate supervisor deployment must satisfy the same conditions.
For shared providers, freeze the affected sandbox identities for the mutation and return individual results under one overall deadline. A newer mutation cannot satisfy an older wait. Missing capabilities, stale or missing observations, disconnection, failed installation, rejected policy, and missing boundary acknowledgment remain incomplete with safe reason categories. A timeout leaves the saved mutation in place and does not replay it. If a mutation was saved but recording its result fails, report uncertainty explicitly without claiming rollback or safe retry.
Distinguish immutable operation history from current readiness. An operation may retain an applied result after a later disconnect or credential expiry, while a current readiness query reports that the target is no longer ready. Historical success cannot override current session, policy, credential, or launch-environment checks. Repeated status queries for the same unchanged target should recover the same observation without continually adding durable records.
CLI and Rust SDK callers can inspect an exact change and wait for it with a bounded deadline. Current authorization applies to status lookup, and operation records contain only safe identities and closed reason categories. This extension adds no external-stable profile field, new reference format, resolver retargeting, or dependency on #3339.
Acceptance Criteria
Alternatives Considered
Fixed delays and application retries: they cannot prove installation or identify the failed component. Restart every client after every mutation: this disrupts workloads and still requires readiness before launching the replacement. Require stable credential references: this couples installation observation to a separate credential-lifecycle design. Create a separate general operation engine: this duplicates #1731 and #3051 and creates incompatible completion and retry semantics.
Agent Investigation
No response
Checklist