Skip to content

docs(rfc): propose OpenShell testing strategy - #3460

Open
elezar wants to merge 7 commits into
mainfrom
codex/testing-target-state/el
Open

elezar wants to merge 7 commits into
mainfrom
codex/testing-target-state/el

Conversation

@elezar

@elezar elezar commented Sep 18, 2026 •

Copy link
Copy Markdown
Member

Summary

Propose a testing strategy for OpenShell that verifies required public behavior across drivers and deployment environments. Separate test requirements, target preparation, and execution policy.

Related Issue

Part of #3954. Implementation remains follow-up work.

Changes

  • Define test families and conformance, feature-specific, and driver-specific e2e types.
  • Describe conformance areas, test requirements, and failure handling without automatic retries.
  • Propose opt-in PR runs, release qualification, and labels for selecting e2e streams.
  • Place suites under e2e/suites/ and separate provisioning from test execution.
  • Plan incremental migration, with disruption tests in a later phase and standalone conformance packaging as a follow-up.

Testing

  • mise run pre-commit passed, including Markdown and Mermaid validation.
  • Whitespace checks passed.
  • Only the RFC changed; runtime tests are not applicable.

Checklist

  • Conventional commit with DCO sign-off.
  • Documentation-only proposal; implementation and CI changes follow separately.

@copy-pr-bot

copy-pr-bot Bot commented Sep 18, 2026

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@elezar elezar added the rfc label Oct 1, 2026
@elezar elezar changed the title docs(testing): define Nix and tmachine target state docs(rfc): propose OpenShell testing strategy Oct 1, 2026
elezar added 4 commits October 1, 2026 14:59
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
@elezar
elezar force-pushed the codex/testing-target-state/el branch from 5309c7e to 9c98a7f Compare October 1, 2026 13:08
elezar added 2 commits October 1, 2026 17:23
Signed-off-by: Evan Lezar <elezar@nvidia.com>
Signed-off-by: Evan Lezar <elezar@nvidia.com>
@elezar
elezar marked this pull request as ready for review October 1, 2026 16:05
@elezar
elezar requested review from a team, derekwaynecarr, mrunalp and sjenning as code owners October 1, 2026 16:05
Comment thread rfc/0016-testing-strategy/README.md Outdated
| --- | --- |
| Unit and component integration | Internal logic and implementation mechanics, at the lowest effective layer. |
| General conformance | Public behavioral contracts across drivers and environments. |
| Feature-specific | Features requiring configured external integration or currently implemented on only one driver. |

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If something is only implemented on one driver, I'd assume that should be part of "driver specific"

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think that's generally true, but only if it doesn't rely on some external service. For example, something that only works on podman, but requires some addtional service to work could also be a feature-specific test.

Comment thread rfc/0016-testing-strategy/README.md Outdated

| Family | Purpose |
| --- | --- |
| Unit and component integration | Internal logic and implementation mechanics, at the lowest effective layer. |

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is this only unit and integration? is everything else an e2e test then?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, the conformance, feature-specific, and driver-specific tests are end-to-end tests. I'll update the tables as per @drew's suggestion to clarify things. We'll use "integration" for component-integration tests and end-to-end for tests that run CLI commands against a real OpenShell installation.

Comment thread rfc/0016-testing-strategy/README.md Outdated
| Family | Purpose |
| --- | --- |
| Unit and component integration | Internal logic and implementation mechanics, at the lowest effective layer. |
| General conformance | Public behavioral contracts across drivers and environments. |

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is general conformance different than feature-specific where all drivers support it? I'm imagining that there's a set of features w/ venn diagrams including specific drivers, and in a world where all drivers are included, it just becomes a circle and is now deemed a "conformance test"

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

effectively wondering if conformance is just a special-case version of feature-specific tests where all drivers are included, or if you see it as something different

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The way I have tried to reason about it, "conformance" is table stakes. These are the things that every OpenShell installation should be able to do. Think: "I can start a sandbox and stop it again", "A running sandbox is restarted when the gateway restarts and keeps its data", "The sandbox can't communicate with the outside world by default".

For feature-specific tests, the driver-dependence is secondary. The factor here is that the functionality requires some external service or property of the system. The example that's currently included in the repo is using keycloak to implement provider refresh. This requires a functional keycloak deployment to work. These tests MAY work on all drivers, by may not be applicable to all OpenShell installations becase they require opt-in behaviour from the user.

Does that help / clarify things?

Comment thread rfc/0016-testing-strategy/README.md Outdated
Tests must not change gateway startup configuration. They may mutate public
API-managed state, using unique names and cleanup and avoiding conflicting
global-setting changes within a run. Document global effects; restoring prior
global settings is recommended, not mandatory.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

at a practical level I don't see any issues with global-setting changes if we're isolating tests, but I find this statement a little confusing. is it implying that multiple tests are sharing the same environment, so a good test-citizen should restore the env how they found it?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think the point is that if a test is a conformance test it should be able to run against any OpenShell installation. This means that it MUST not updated gateway config to get a test to pass. The discussion around global state comes from my agent looking at some of the tests and determined that running them modifies some globabl state that isn't the CAS store (I think it had something to do with OCSF). I need to improve the wording to focus on not modifying a running Gateway in ways that a user with CLI / API access would not be able to do.

Comment thread rfc/0016-testing-strategy/README.md Outdated
distinguish test corrections from changes to promised behavior, including
withdrawal of advertised support.

### 3. Make applicability and coverage explicit

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

making sure I understand this section:

Rather than encoding in our test-suite which drivers support features a/b/c, that should be encoded as an API contract that consumers can see. The the test suite just becomes one consumer of that API, and decides which tests are worth running against this given openshell deploy?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No, I don't think that's correct. (Which shows that I need to rework this).

The goal of the conformance test suite is that it can run against any OpenShell installation. In cases where this is not possible (or is not yet possible), my thinking was that one should have a programatic way of indicating that a failure is expected and indicates a conformance gap for a particular configuration. For example, the MXC driver does not support openshell sandbox exec which means that a conformance test that checks: "I can run a sandbox and exec something in it" will not work there. A conformance test that fails (as expected) against this driver will then allow us to explicitly capture the gap and also give a strong signal of the change when we finally add support.

The alternative is to have the tests discover what a driver supports and only run a subset of tests against a driver (or positive and negative tests depending on the reported capability). As @drew mentions later this is probably not scalable.

I think with a couple of concrete examples, we can tighten the definitions a bit.

Comment thread rfc/0016-testing-strategy/README.md Outdated

| Family | Purpose |
| --- | --- |
| Unit and component integration | Internal logic and implementation mechanics, at the lowest effective layer. |

@drew drew Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

its a little confusing to put unit and integration in the same sentence. these are usually distinct types of tests and integration has become very overloaded. we call current conformance tests, "conformance integration", and current feature specific tests "feature specific integration"

Image

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

They were grouped together in the doc because they are both product internal. That is to say that they interact with Rust primatives at a software level and do not require "real" testing infrastructure.

You're correct that "integration" is overloaded here. In the cases you mention these should be end-to-end tests of which conformance and feature-specific are two classes. As part of this work, we can update the CI jobs too.

Comment thread rfc/0016-testing-strategy/README.md Outdated
and a suite groups related cases. For example, deleting a sandbox must remove
it from the sandbox list; an assertion checks that its identifier is absent.

| Family | Purpose |

@drew drew Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I would suggest the following taxonomy

Test family Description
Lint Check formatting, style, and static rules.
Unit Verify one component in isolation. Place tests inline in its Rust module or in an adjacent test.rs file under src/.
Integration Verify interactions between components. Place Rust integration tests in the crate’s top-level tests/ directory.
End-to-end Verify a configured OpenShell target through a client interface. Place suites in the repository-level e2e/ directory; conformance is a class of end-to-end test for portable public contracts.
Benchmark Measure performance or scale. Place crate benchmarks in benchmarks/ and full-system benchmarks in a root benchmarks/.

Types of e2e tests

End-to-end type What it verifies
Conformance Portable public contracts across applicable drivers, including advertised optional capabilities. Includes recovery and continuity after an induced failure. It should be possible to run these tests out of tree.
Feature Behavior requiring a named external service or special gateway configuration.
Driver Behavior specific to a driver, its host integration, or its configuration.
Installation Installing, upgrading, and uninstalling candidate artifacts.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I like the suggestion.

I also think I've convinced myself that the tests that we're writing are user-facing e2e tests and not system-level integration tests. I'll update naming to address that.

Comment thread rfc/0016-testing-strategy/README.md Outdated
Comment thread rfc/0016-testing-strategy/README.md Outdated
suites. Installation, upgrade, and uninstall behavior need separate assertions;
successful conformance alone does not validate packaging.

## Implementation plan

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It would be good if we could start to define what conformance tests we're going to build. I would propose the following suites. Here's a list to get us started

  • Sandbox lifecycle: Create, inspect, execute, stop, restart, and delete.
  • Policy: Validate, apply, update, and report effective policy.
  • Sandbox enforcement: Enforce filesystem, process, and network rules.
  • Providers: Manage providers and verify credential delivery, isolation, rotation, and protection from exposure.
  • Identity and authorization: Authenticate and enforce access boundaries.
  • Middleware: Verify selection, ordering, transformation, and failure handling.
  • Interceptors: Verify request transformation, rejection, and preservation of authorization.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, adding categories / focus areas for the conformance tests makes sense (and already exists to some extend). I'll update the doc to formalize it a bit.

Comment thread rfc/0016-testing-strategy/README.md Outdated
├── config.nix # tmachine definitions
├── artifacts.nix # Artifact construction
├── ansible/ # Provisioning and execution
├── CONFORMANCE.md # Proposed: agreed policy

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why is conformance at the top instead of inside conformance/?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this was pulled from k8s, but should be updated. I'll assess the layout again a bit more critically.

Comment thread rfc/0016-testing-strategy/README.md Outdated
specialized infrastructure. Performance thresholds belong in load/scale unless
the deadline is itself a public contract.

### 2. Define conformance through public behavior

@drew drew Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this whole section is pretty dense. i'm not sure i understand it. conformance tests should

  • use public apis
  • be parametrized by gateway endpoint
  • have the ability to run out of tree so third parties can test for conformance
  • run as part of nightly qualification
  • run on branch checks when manually triggered.
  • it would be great if we can granularly trigger conformance checks by suite on ci

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think I agree with your points, but want to be careful to not mix WHAT the tests are with HOW and WHEN they are run.

Q: What are conformance tests?
A: Conformance tests are end-to-end tests that run against any OpenShell installation (i.e. parameterized endpoint). They use public APIs (e.g. the openshell CLI or SDKs) to test required behaviour against the configured installation.

Q: How are conformance tests run?
A: (Still in progress) They are run as a nextest archive against the configured installation. The archive can be downloaded and run against any OpenShell installation and do not required the OpenShell CI machinery.

Q: When are they run?
A: Whenever feasible.

This last answer could also be "it depends". We should definitely run them as part of automated qualification. It should also be possible to run them locally for development and trigger them in CI. Providing more granular selection should be possible and could also be considered, but then we would have to determine what the selection criteria are. Why are we NOT running all the tests?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(I will work on rewording this).

Comment thread rfc/0016-testing-strategy/README.md Outdated
configuration, or equivalent target identity. Record mock targets as mocks,
not evidence for production drivers.

### 4. Separate target preparation from test execution

@drew drew Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can we please detail what tests run where. for example, unit tests always run on ci, e2e tests run when manually triggered or on release qualification, etc. we also need to cleanup the current labelling approach.

i would also like to see us be more efficient in what tests run on branch checks. @SDAChess mentioned work to dynamically figure out what tests to run per pr. i think we should consider this. i've seen interesting approaches that use llms or jev to figure out what tests should be run based on changes. might be interesting to experiment with.

we don't have to build this all at once, but since this rfc is broadly titled "propose OpenShell testing strategy" and talks about execution stragies, i think we should detail where we're headed with things. current branch checks on prs are too slow.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I have updated the contents to give a better description of what runs where as well as defining concrete labels for the different "streams" of tests. These also allow for futher sub-categories.

Although dynamic change-based selection is a goal to strive to, I don't want to make that the focus of the first iteration of this RFC. It should lay the groundwork to enable that as a follow-up. I'm happy to reword the title and update the issue to narrow the scope.

Comment thread rfc/0016-testing-strategy/README.md Outdated
Comment on lines +119 to +120
| Unsupported | Optional support was not advertised. |
| Skipped | The test was deliberately excluded. |

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

i'm assuming this is just for conformance tests. how do we specify a test to be unsupported or skipped?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This has been updated. We no longer have these different states. We either have success of failure, with failure in conformance tests indicating a gap in driver support for expected behaviour.

Comment thread rfc/0016-testing-strategy/README.md Outdated
Comment on lines +104 to +108
The gateway must report effective capabilities for its running configuration
through the public API and machine-readable CLI output. Discovery failure aborts
conformance. Start with flat, namespaced booleans; defer hierarchy, parameters,
and profiles. Capabilities describe product behavior, not test selectors,
credentials, or external-service prerequisites.

@drew drew Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not sure how scalable this is. For example, what if down the road we support loading more than one compute driver. Some compute drivers might enable different capabilities. This is also going to cause options on the compute driver to explode with every product capability a driver might or might not support. We've already started to see this and it creates quite a change amplification that I'd like to avoid.

As part of RFC-0012 we proposed using validation to assert capabilities. For example, if you create a sandbox with a file system policy, and no filesystem policy exists, that sandbox should fail to create will a validation error. I think this is a more scalable approach.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, gateway-reported capabilities might not be the right approach here. As I mentioned to MG above, we may be able to get a better idea of what things look like once we add some concrete examples.

Comment thread rfc/0016-testing-strategy/README.md Outdated
Comment on lines +97 to +100
Use normal PR review, without a soak period or separate promotion PR. Resolve
known flakiness rather than hiding it with retries. Changes or removals must
distinguish test corrections from changes to promised behavior, including
withdrawal of advertised support.

@drew drew Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should handle some level of flaky-ness and include infra for retries. If a test is identified as flaky there should be a report that we monitor and fix out of band. If we hard fail on flakes, we're going to be fighting builds and wind up just manually retrying tests anyways (we already see this today).

Down the road, if there's a report we can have some agent iterate on the report to reduce flakes.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The point here was to simplify the process for calling a test a conformance test. As I understand it, in the k8s case, the tests have to be considered stable for a period of time before being allowed to be labeled as such.

I do agree that we should include some flakiness analysis of our tests in general.

@politerealism

Copy link
Copy Markdown
Contributor

Research pass (no code changes) on the "temporary e2e-podman feature, exclusion list" the implementation plan calls out for retirement (item 4), done in support of this RFC's general-conformance/driver-specific categorization. Posting findings here since this seems like the concrete migration example the open questions section asks for.

Ground truth: e2e/rust/tests/ has 54 test files. The e2e-podman Cargo feature (which pulls in e2e, e2e-host-gateway, e2e-local-container-driver) makes 37 of them reachable under the Podman driver. PODMAN_CI_TESTS in e2e/rust/e2e-podman.sh — the required-CI inclusion list — only runs 20 of those 37. That leaves 17 tests that run under Podman today but are excluded from required CI, which I read in full against this RFC's category definitions.

Only 3 of 17 are genuinely driver-specific

File Why it can't generalize
podman_oci_identity.rs Shells directly to podman inspect/podman ps --filter label=..., asserts on Podman-specific container state (EffectiveCaps, HostConfig.NetworkMode, Mounts). Tests the Podman driver's own OCI-build→inspect→launch implementation, not a public contract.
podman_userns.rs Tests Podman's userns=keep-id/auto/private modes — a Podman-only gateway config concept — by editing the [openshell.drivers.podman] TOML section directly and inspecting /proc/self/uid_map. No equivalent exists for other drivers.
transparent_tcp.rs's rootless_podman_musl_getaddrinfo_uses_udp_policy_dns test Explicitly gated if !is_e2e_driver("podman"); its own doc comment says it isolates rootless Podman's UDP policy-DNS path specifically.

11 of 17 are already driver-independent in practice — general-conformance candidates

port_forward.rs, provider_auto_create.rs, provider_refresh_handles.rs, proxy_egress_pipeline.rs, sandbox_labels.rs (not even feature-gated today — already compiles unconditionally), sandbox_templates.rs, settings_management.rs, sync.rs, upload_create.rs, websocket_conformance.rs, workspace_lifecycle.rs. These exercise only the public CLI/API and generic process/network assertions — no engine shell-outs, no runtime inspection. Their current e2e-podman/e2e-host-gateway gating looks like an artifact of CI convenience, not an actual dependency.

2 of 17 are mostly general with one driver-coupled corner

  • sandbox_lifecycle.rs — almost entirely CLI-driven, but a few tests walk /proc directly for PID-signal assertions and one diagnostics helper shells to docker. Those sub-tests assume a local host-process model (fine for Docker/Podman, not Kubernetes/VM) — worth splitting out rather than excluding the whole file.
  • transparent_tcp.rs's other test (local_container_native_tcp_uses_policy_dns_and_fails_closed) — gated Docker-or-Podman, but assertions are identical either way; only fixture setup branches. This reads as feature-specific, not driver-specific — the harness needs extending to a third driver, not new assertions.

2 of 17 aren't conformance tests at all

internet_network_perf.rs and live_internet_traffic_perf.rs are opt-in (#[ignore]) performance benchmarks. Per this RFC's categories these belong in load/scale, not general conformance or driver-specific — they shouldn't have been part of this exclusion-list question in the first place.

Pattern worth flagging directly

Several files are gated as e2e-podman/e2e-host-gateway purely by historical convenience of how they're currently launched in CI, not by any real Podman dependency in the test body — provider_refresh_handles.rs and websocket_conformance.rs are the clearest examples. That matches this RFC's framing exactly: "Rootful and rootless Podman are two environments of one driver, not two-driver evidence" — several of these 11 general-conformance candidates are currently only validated against one driver's environment, not actually coupled to it.

Net: 11 of 17 are ready-now general-conformance candidates (pending a second driver's evidence per the admission rule), 1 more is feature-specific pending harness work, 1 needs splitting rather than wholesale reclassification, 3 are legitimately driver-specific, and 2 aren't conformance tests at all (load/scale).

Happy to go deeper on any of these (e.g. draft the actual split for sandbox_lifecycle.rs, or run the 11 candidates against Docker to confirm the "two drivers" admission bar) if useful.

Signed-off-by: Evan Lezar <elezar@nvidia.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants