Skip to content

bug: Podman supervisor cannot trust operator-supplied CA for upstream TLS after v0.1.2 architecture change #3781

Description

@Tojaj

User Story

As an operator running Podman sandboxes against private-CA HTTPS services, I want inspected egress to trust an operator-supplied CA, so that agents can reach internal services without rebuilding the supervisor image for each OpenShell upgrade or CA rotation.

Problem Statement

After upgrading OpenShell from v0.0.116 to v0.1.2, a Podman sandbox whose workload image trusts a private CA no longer receives HTTP responses from services signed by that CA. The sandbox client completes TLS with OpenShell's interception certificate and sends its request, then receives curl: (52) Empty reply from server. The supervisor allows the connection but fails during upstream TLS establishment, before emitting an HTTP request event. The same service responds from the host, and other HTTPS destinations work.

This is a regression in the trust path: v0.0.116 ran the supervisor binary inside the workload container, where it could read that image's CA bundle. v0.1.2 runs a separate container from supervisor_image, whose CA bundle does not inherit workload-image additions. The supervisor reads its own system bundle for upstream TLS verification. v0.0.116 Podman spec, v0.1.2 supervisor spec, upstream root-store construction.

Impact / Why This Matters

Policy-allowed, inspected HTTPS requests to internal services stop working after the upgrade. The confirmed workaround is to build a version-matched custom supervisor image containing the private root, configure supervisor_image, restart the gateway, and create a fresh sandbox. That image must be rebuilt for OpenShell upgrades and CA rotation. tls: skip can avoid supervisor upstream verification but gives up HTTP inspection and credential rewriting, so it is insufficient for inspected endpoints. Podman's proxy_ca_bundle requires https_proxy and cannot configure trust for direct egress. Podman validation.

Acceptance Criteria

  • Podman sandboxes can use an operator-owned additional CA bundle for direct, inspected HTTPS egress without configuring https_proxy or rebuilding the supervisor image.
  • A policy-allowed private-CA HTTPS endpoint returns its application response while public-CA endpoints continue to work.
  • Upstream certificate and hostname verification remain enabled; malformed or unreadable configured CA material produces an actionable error.
  • The operator-provided trust material reaches the supervisor's upstream TLS client and remains separate from gateway authentication trust.

Reproduction Steps

  1. On v0.0.116, run a Podman sandbox with a workload image whose system CA bundle contains a private root. Allow an HTTPS destination signed by that root and request it from the sandbox. Observe an HTTP response.
  2. Upgrade to v0.1.2 and create a fresh sandbox using the same workload image and policy, with the stock supervisor image and direct egress (no https_proxy). Request the same destination.
  3. Observe curl: (52) Empty reply from server and a supervisor NET:FAIL event reporting Upstream TLS establishment failed after NET:OPEN ... ALLOWED. Verify the destination still responds from the host.
  4. Build a custom v0.1.2 supervisor image with that root CA, set Podman's supervisor_image to it, restart the gateway, and create a fresh sandbox. The inspected request now returns an application-level response.

Environment

  • OpenShell: worked on v0.0.116; regressed on v0.1.2 (latest release at filing).
  • Runtime: Podman with the native Podman compute driver and separate workload/supervisor containers in v0.1.2.
  • Egress: direct HTTPS to a service signed by a private CA; no corporate forward proxy.
  • Host OS and Podman version: not recorded in the source report.

Logs

NET:OPEN ... ALLOWED ... -> <private-CA-host>:443
NET:FAIL ... <private-CA-host>:443
"message":"Upstream TLS establishment failed"
curl: (52) Empty reply from server

The OCSF event does not include the nested TLS library error. The custom-supervisor-image test confirms the missing private root in the stock supervisor's trust store as the cause in this environment.

Related: #3295 requests operator-managed additional destination CA trust across drivers. This report records the v0.1.2 Podman regression and its confirmed workaround.

Activity

  1. politerealism commented on Sep 28, 2026

    @politerealism
    Contributor

    📋 triage-agent

    Triage Assessment

    Classification: validated-bug

    Summary

    Confirmed. The report's code references and reproduction narrative match current main: in the Podman driver, the supervisor runs as its own container using input.config.supervisor_image (crates/openshell-driver-podman/src/container.rs:1588), separate from the workload container. build_upstream_root_store (crates/openshell-supervisor-network/src/l7/tls.rs:218-249) overlays only the system CA bundle found at fixed SYSTEM_CA_PATHS (tls.rs:25-31) inside whichever image the supervisor process itself runs in — the stock supervisor image, not the workload image that trusts the private CA. This is a real regression versus v0.0.116, where the supervisor binary ran inside the workload container and inherited its CA bundle.

    Investigation

    Impact Signals

    • Affected users/scope: Any Podman-driver operator using inspected (non-tls: skip) HTTPS egress to a privately-signed internal service, on the current architecture (v0.1.2+, separate supervisor container). Likely also affects Docker/VM/Kubernetes for the same root-store reason, though this report only reproduces Podman.
    • Regression: Yes — confirmed behavior change from v0.0.116 to v0.1.2 caused by the workload/supervisor container split, not a pre-existing limitation.
    • Workaround: Available but costly — build and maintain a version-matched custom supervisor_image containing the private root; must be rebuilt on every OpenShell upgrade and CA rotation.
    • Evidence quality: High — reporter cites exact line ranges across two tagged versions, includes OCSF log excerpts (NET:OPEN ... ALLOWED followed by NET:FAIL ... Upstream TLS establishment failed), and confirmed the root cause empirically via the custom-image workaround.

    Human Decision Required

    Decide whether OpenShell should address this issue, and whether to track it separately from #3295 or fold it in as regression evidence for that broader feature. If yes, apply
    state:accepted, associate it with a roadmap item, or do both, and decide
    whether the work remains human-owned. Either action records acceptance;
    roadmap placement additionally records sequencing.
    To queue investigation or planning for an unattended agent, also apply
    agent:plan-requested. You can instead directly ask an agent to use
    create-spike or build-from-issue on this issue; the agent will warn about
    missing expected workflow labels and continue without changing them. If no,
    close it as not planned and record the rationale.

  2. jpiersol commented on Sep 29, 2026

    @jpiersol

    I'm also seeing this error and would love a fix.

  3. politerealism commented on Sep 30, 2026

    @politerealism
    Contributor

    Filing a note on the plan here before starting work: we're going to land a small, Podman-scoped fix targeting exactly this regression — restoring the ability to trust an operator-supplied CA for direct, inspected HTTPS egress from the Podman supervisor, without requiring a custom supervisor_image rebuild per upgrade/rotation.

    @jhjaggars's #3292 is a larger, more complete cross-driver implementation of this (Docker/Podman/Kubernetes/VM delivery, gateway-level config schema, lifecycle-safe rotation reconciliation) and is the direction we'd like to work toward and eventually adopt. We're not landing that broader surface area right now — there are a few other things in flight on our side that need to settle before taking on a change of that size — so this fix is intentionally scoped down to just what #3781 needs, not a replacement for or competing design against #3292.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    state:triage-neededOpened without agent diagnostics and needs triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions