Skip to content

feat(services): continuously observe HTTP application readiness #4229

Description

@drew

User Story

I need continuous application readiness for exposed sandbox services, including OpenShell's Codex app-server example, so I can see whether the application is ready before connecting and while it runs. The requested workflow is deliberately simple, with HTTP checks and fixed timings.

Problem Statement

OpenShell exposes HTTP and WebSocket services but service reads currently report endpoint configuration and URLs without an ongoing application readiness observation. Sandbox lifecycle readiness does not establish that an exposed application is responding successfully.

Impact / Why This Matters

Users must poll each application's endpoint separately. That polling does not appear in OpenShell's service API or CLI and must be recreated by each client.

Proposed Design

Continuously check exposed services through their existing loopback target port. By default, GET / establishes HTTP responsiveness. An optional readiness path requires a 2xx response. Check every five seconds with a one-second timeout; three consecutive failures mark unhealthy and one success restores health. Return cached health through service Get/List, invalidate observations after runtime or configuration changes, and report unknown for stale or unavailable observations. Health is observational and does not change routing, sandbox lifecycle, or restart policy.

Suggested UX

openshell sandbox create --name codex-app-server --from openshell/codex-app-server:local \
  --expose 4500 --expose-readiness-path /readyz --detach -- start-codex-app-server
openshell service list codex-app-server --output json
openshell service expose my-sandbox 8080 web --readiness-path /readyz

Acceptance Criteria

  • Service Get/List return current cached health with check time, HTTP status when available, and a concise observation message.
  • Optional readiness paths are validated and persist across endpoint updates that omit them.
  • HTTP 2xx responses pass application readiness; any HTTP response passes default listener responsiveness.
  • Checks continue while the sandbox is ready, use fixed timings and thresholds, and expire stale observations.
  • Stopped, restarted, reconfigured, and unavailable sandboxes/services cannot retain misleading healthy observations.
  • Codex example configures /readyz and documents how to inspect readiness.
  • Tests cover failure/recovery, stale observations, API configuration, and the Docker service relay.

Alternatives Considered

TCP connectivity establishes that a port accepts connections but does not establish HTTP application readiness. A startup-only probe does not detect failures while a service runs. External polling can perform the check but does not provide a shared observation in OpenShell's service API. Supervisor middleware and gateway interceptors apply to traffic processing; using them alone would not generate a continuous observation when there is no client traffic.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions