Repository navigation
bug: Podman supervisor cannot trust operator-supplied CA for upstream TLS after v0.1.2 architecture change #3781
Description
Activity
- addedstate:triage-neededOpened without agent diagnostics and needs triageOpened without agent diagnostics and needs triage
on Sep 28, 2026 📋 triage-agent
Triage Assessment
Classification: validated-bug
Summary
Confirmed. The report's code references and reproduction narrative match current
main: in the Podman driver, the supervisor runs as its own container usinginput.config.supervisor_image(crates/openshell-driver-podman/src/container.rs:1588), separate from the workload container.build_upstream_root_store(crates/openshell-supervisor-network/src/l7/tls.rs:218-249) overlays only the system CA bundle found at fixedSYSTEM_CA_PATHS(tls.rs:25-31) inside whichever image the supervisor process itself runs in — the stock supervisor image, not the workload image that trusts the private CA. This is a real regression versus v0.0.116, where the supervisor binary ran inside the workload container and inherited its CA bundle.Investigation
proxy_ca_bundle(Podman/VM only) is confirmed gated onhttps_proxybeing set (crates/openshell-driver-podman/src/config.rs:331-334), so it cannot be repurposed for direct-egress private-CA trust as the report states.tls: skipavoids the upstream verification failure but disables L7 inspection, consistent with the report.- This is the same underlying gap as feat: trust additional destination CAs for sandbox egress TLS #3295 ("feat: trust additional destination CAs for sandbox egress TLS"), which is already
state:validatedwith a candidate implementation (PR feat(network): support additional destination CAs #3292) and a full cross-driver design write-up. bug: Podman supervisor cannot trust operator-supplied CA for upstream TLS after v0.1.2 architecture change #3781 is not a pure duplicate: it documents a concrete regression (working in v0.0.116, broken in v0.1.2 specifically due to the workload/supervisor container split) with a verified workaround (customsupervisor_imagebuild), rather than a general capability request. The reporter already cross-linked feat: trust additional destination CAs for sandbox egress TLS #3295 and did not claim duplication. - No other open or closed issue duplicates the regression framing.
Impact Signals
- Affected users/scope: Any Podman-driver operator using inspected (non-
tls: skip) HTTPS egress to a privately-signed internal service, on the current architecture (v0.1.2+, separate supervisor container). Likely also affects Docker/VM/Kubernetes for the same root-store reason, though this report only reproduces Podman. - Regression: Yes — confirmed behavior change from v0.0.116 to v0.1.2 caused by the workload/supervisor container split, not a pre-existing limitation.
- Workaround: Available but costly — build and maintain a version-matched custom
supervisor_imagecontaining the private root; must be rebuilt on every OpenShell upgrade and CA rotation. - Evidence quality: High — reporter cites exact line ranges across two tagged versions, includes OCSF log excerpts (
NET:OPEN ... ALLOWEDfollowed byNET:FAIL ... Upstream TLS establishment failed), and confirmed the root cause empirically via the custom-image workaround.
Human Decision Required
Decide whether OpenShell should address this issue, and whether to track it separately from #3295 or fold it in as regression evidence for that broader feature. If yes, apply
state:accepted, associate it with a roadmap item, or do both, and decide
whether the work remains human-owned. Either action records acceptance;
roadmap placement additionally records sequencing.
To queue investigation or planning for an unattended agent, also apply
agent:plan-requested. You can instead directly ask an agent to use
create-spikeorbuild-from-issueon this issue; the agent will warn about
missing expected workflow labels and continue without changing them. If no,
close it as not planned and record the rationale.I'm also seeing this error and would love a fix.
Filing a note on the plan here before starting work: we're going to land a small, Podman-scoped fix targeting exactly this regression — restoring the ability to trust an operator-supplied CA for direct, inspected HTTPS egress from the Podman supervisor, without requiring a custom
supervisor_imagerebuild per upgrade/rotation.@jhjaggars's #3292 is a larger, more complete cross-driver implementation of this (Docker/Podman/Kubernetes/VM delivery, gateway-level config schema, lifecycle-safe rotation reconciliation) and is the direction we'd like to work toward and eventually adopt. We're not landing that broader surface area right now — there are a few other things in flight on our side that need to settle before taking on a change of that size — so this fix is intentionally scoped down to just what #3781 needs, not a replacement for or competing design against #3292.
Reacted by Tomas Mlcoch
User Story
As an operator running Podman sandboxes against private-CA HTTPS services, I want inspected egress to trust an operator-supplied CA, so that agents can reach internal services without rebuilding the supervisor image for each OpenShell upgrade or CA rotation.
Problem Statement
After upgrading OpenShell from v0.0.116 to v0.1.2, a Podman sandbox whose workload image trusts a private CA no longer receives HTTP responses from services signed by that CA. The sandbox client completes TLS with OpenShell's interception certificate and sends its request, then receives
curl: (52) Empty reply from server. The supervisor allows the connection but fails during upstream TLS establishment, before emitting an HTTP request event. The same service responds from the host, and other HTTPS destinations work.This is a regression in the trust path: v0.0.116 ran the supervisor binary inside the workload container, where it could read that image's CA bundle. v0.1.2 runs a separate container from
supervisor_image, whose CA bundle does not inherit workload-image additions. The supervisor reads its own system bundle for upstream TLS verification. v0.0.116 Podman spec, v0.1.2 supervisor spec, upstream root-store construction.Impact / Why This Matters
Policy-allowed, inspected HTTPS requests to internal services stop working after the upgrade. The confirmed workaround is to build a version-matched custom supervisor image containing the private root, configure
supervisor_image, restart the gateway, and create a fresh sandbox. That image must be rebuilt for OpenShell upgrades and CA rotation.tls: skipcan avoid supervisor upstream verification but gives up HTTP inspection and credential rewriting, so it is insufficient for inspected endpoints. Podman'sproxy_ca_bundlerequireshttps_proxyand cannot configure trust for direct egress. Podman validation.Acceptance Criteria
https_proxyor rebuilding the supervisor image.Reproduction Steps
https_proxy). Request the same destination.curl: (52) Empty reply from serverand a supervisorNET:FAILevent reportingUpstream TLS establishment failedafterNET:OPEN ... ALLOWED. Verify the destination still responds from the host.supervisor_imageto it, restart the gateway, and create a fresh sandbox. The inspected request now returns an application-level response.Environment
Logs
The OCSF event does not include the nested TLS library error. The custom-supervisor-image test confirms the missing private root in the stock supervisor's trust store as the cause in this environment.
Related: #3295 requests operator-managed additional destination CA trust across drivers. This report records the v0.1.2 Podman regression and its confirmed workaround.