Repository navigation
bug(network): port-80 plaintext rest endpoints stall behind CONNECT-only upstream proxy (no deny, no relay) #4278
Description
Activity
- addedstate:triage-neededOpened without agent diagnostics and needs triageOpened without agent diagnostics and needs triage
on Oct 7, 2026 - addedstate:needs-infoAssessment needs specific evidence or reproduction detailsAssessment needs specific evidence or reproduction details
on Oct 7, 2026 @piflepaf1-eng can you provide a proxy configuration that reproduces this? I think I understand that in theory our supervisor would need to know that the upstream corp proxy only allows CONNECT before it tries to dial plain HTTP. I don't want to guess on how to do that and I'm not sure we want to expose another config knob to have to specify that as well. So having a known env that matches yours would be good.
Thanks — here's a repro that doesn't require our internal config. Any forward proxy that accepts CONNECT tunnels but refuses absolute-form plaintext HTTP reproduces it. Minimal squid:
squid.conf:
http_port 3128 acl SSL_ports port 443 http_access deny CONNECT !SSL_ports http_access allow CONNECT http_access deny alldocker-compose.yml:
services: squid: image: ubuntu/squid:latest ports: ["3128:3128"] volumes: ["./squid.conf:/etc/squid/squid.conf:ro"]Point the gateway driver config at it (https_proxy = "http://:3128"), create a sandbox with one
protocol: restendpoint on port 80 and one on 443 for the same host, then curl both from the workload: port 80 connects at L4 and then stalls (0 bytes relayed until timeout, no HTTP:* audit event); 443 completes normally — same asymmetry as in the issue body.Either behavior would unblock us: relaying plaintext HTTP through the tunnel, or a fast, clearly-reasoned deny at policy validation / connect time instead of an open-ended stall.
Follow-up: we reproduced this end-to-end with the squid repro above, and there is a detail that may help diagnose it.
With squid instrumented (access.log), the proxy is NOT actually silent in our synthetic setup: it answers the plaintext-CONNECT dial with a fast
403 TCP_DENIED(0.6 ms). The workload still sees the exact phenomenology from the issue — 60 s silent stall, 0 bytes,NET:FAIL, noHTTP:*audit event. So in our repro the stall is produced by the supervisor swallowing the proxy's error response: it neither relays the 403 to the workload nor closes the connection.In our production network the upstream proxy never answers at all (stall at the source), so both paths end in the same observable state — but it means "fail fast" has two distinct layers to consider:
- the supervisor relaying/closing on upstream proxy error responses (would have turned our squid repro into an immediate, clearly-reasoned failure), and
- a policy-validation-time reject for port-80 plaintext endpoints, as in the original report.
Either would remove the silent-hang debugging cost.
One note for anyone reproducing on a network without direct egress (explicit-proxy-only environments): the squid container as posted assumes it can reach the internet directly. If it can't, chain it to your existing proxy with
cache_peer <existing-proxy> parent <port> 0 no-query+never_direct allow all— the client-side ACL semantics (deny CONNECT !SSL_portsetc.) are unchanged.
Summary
On a network where the only egress is an explicit corporate proxy that supports CONNECT/TLS only, any
protocol: restendpoint on port 80 (plain HTTP) is unusable: the sandbox L4 layer accepts the connection (synthetic IP), the supervisor dials upstream through the forward proxy, and the relay then stalls indefinitely — 0 bytes relayed, noHTTP:*audit event, workload-sidecurleventually times out. The same host on port 443 works in the same sandbox with the same policy, so the failure is specific to plaintext HTTP through the CONNECT-only proxy chain.Two things could be improved here, either one would unblock us:
port: 80external endpoint at policy validation (sandbox create) with a distinct reason code, instead of stalling at connect time.Environment
https_proxyto the corporate explicit proxy (port 3128) — the proxy chain accepts CONNECT/TLS and does not complete plaintext HTTP dials;no_proxycovers internal hostsnone(all egress via supervisor relay, as designed)Repro
Task policy (
task-policy-port80.yaml, minimal — same host on both ports):Command:
Observed — port 80 (synthetic IP resolves and connects at L4, then nothing arrives):
Observed — port 443, same sandbox, same policy, same host, immediately after: full TLS 1.3 handshake,
HTTP/1.1 200 OK, curl exit 0.OCSF audit log (
openshell logs <name>) shows the asymmetry — no ALLOW event for the 80 leg, only the failure; the 443 leg gets the normalNET:OPEN+HTTP:GETpair:Expected
Either plaintext HTTP to policy-approved port-80 endpoints relays through the CONNECT-only upstream proxy (HTTP-in-tunnel), or the endpoint is rejected up front:
sandbox create/ policy validation, with a distinct deny reason code (e.g. "plaintext endpoint unreachable through CONNECT-only proxy"), orNET:FAIL/deny at connect instead of an open-ended stall with zero feedback to the workload (it looks like a hung network, not a policy/topology incompatibility).Impact
http://:80endpoints — still common for package mirrors and internal/legacy services — are unusable on CONNECT-only-proxy networks.Notes / related issues
protocol: tcp; this one isprotocol: restthrough the configured upstream proxy.NET:FAILafter successful requests; ours is deterministic and port-80-specific.restendpoints.