The names, shapes and slot positions every observability component shares (ADR-0041 revision 3, ADR-0042). Permanent document: it outlives the implementation task board and is what a dashboard author or a new emitter reads.
Implemented once, in packages/runtime/src/telemetry/, exported as
@handsontable/demo-runtime/telemetry — the subpath-export pattern monitor-inject.ts
already uses for @handsontable/demo-runtime/monitor. The module is pure (no DOM, no
Cloudflare imports), so the API worker, the o11y worker, the authoring app and
pipeline/ tests import the same definitions. pipeline/telemetry-contract.test.mjs
parses the tables below and fails when the module disagrees with this file.
Changing this file: in its own small PR, together with the module, before any code relies on the change. Metric rows and slots are append-only: never reuse, move or rename a slot or a metric name. Analytics Engine columns are positional, so a moved slot silently corrupts every stored row and no type checks it.
| Name | Path | Notes |
|---|---|---|
handsontable-demos-o11y |
workers/o11y/ |
Worker + InboxWriter DO + GrafanaBox Container class; workers_dev: false, preview_urls: false |
| Grafana box image | containers/o11y/ |
Loki + Grafana only |
Routes on demos.handsontable.com, owned by the o11y worker, passed as --routes flags
in its deploy script (ADR-0020), never in wrangler.jsonc:
| Route | Purpose | Gate (ADR-0041 §B.5) |
|---|---|---|
POST /telemetry/collect |
Faro payloads from the authoring app | host/env, bot filter, caps, kind allowlist, rate limit, server scrub |
POST /telemetry/lite |
lite beacon from /d and /embed |
same as collect |
POST /telemetry/v1/logs |
Cloudflare OTLP log export | x-o11y-secret |
POST /telemetry/deploy |
deploy event from CI | GitHub OIDC token, x-o11y-secret fallback |
POST /telemetry/hooks/sentry |
Sentry issue-alert webhook | sentry-hook-signature HMAC |
GET /grafana/_o11y/login |
start a broker sign-in | per-IP rate limit; mints the __Host-o11y_login nonce cookie |
GET /grafana/_o11y/callback |
broker return; static page, no server-side check of its own | none (fix round M6: the __Host-o11y_login cookie + the live /broker/userinfo call happen on the session row below, not here — this route only serves a hash-pinned static page) |
POST /grafana/_o11y/session |
mint the __Host-o11y_session cookie |
per-IP rate limit; __Host-o11y_login cookie (nonce bound), same-origin Origin, @handsontable.com broker identity |
GET /grafana/_o11y/logout |
a same-origin sign-out page (fix round M5) | none; the page itself fires the POST below |
POST /grafana/_o11y/logout |
clear both cookies | same-origin Origin |
/grafana/* |
Grafana UI, waking page | the Worker's own session cookie (O11Y_SESSION_SECRET, K1 — not Access) |
POST /grafana/_o11y/reopen |
manual ledger re-open | session cookie, same-origin Origin |
GET /grafana/_o11y/admin/<name> |
ADR-0043 read forwarder (after launch) | session cookie, name allowlist, GET only |
There is no trace route: traces are not exported (ADR-0041 §C.4).
Ports inside the Grafana box, reached only through GrafanaBox.containerFetch:
| Port | Service |
|---|---|
| 3000 | Grafana (served from sub-path /grafana/) |
| 3100 | Loki HTTP (/otlp/v1/logs, /ready, /metrics) |
o11y worker (workers/o11y/wrangler.jsonc):
| Name | Kind | Value / purpose |
|---|---|---|
INBOX_WRITER |
Durable Object | class InboxWriter, one instance main, .jurisdiction("eu"); owns inbox keys, ledger, dedupe set, fingerprint registry, alert state |
GRAFANA_BOX |
Durable Object + Container | class GrafanaBox, one instance box, .jurisdiction("eu"), container jurisdiction: "eu" |
O11Y_INBOX |
R2 | bucket handsontable-demos-o11y-inbox (EU) |
O11Y_LOKI_STATE |
R2 | bucket handsontable-demos-o11y-loki (EU); the Worker reads only state/wakes/<wakeId>/clean markers |
O11Y_MAPS |
R2 | bucket handsontable-demos-o11y-maps (EU) |
RUNNER_EVENTS |
Analytics Engine | dataset runner_events |
API |
service binding | handsontable-demos-api (o11y usage metering, o11y spend, later AdminReads) |
RATE_LIMITER |
Rate Limiting binding | gates POST /telemetry/collect and POST /telemetry/lite (ADR §B.5): 100 requests / 60 s per cf-connecting-ip, 429 with Retry-After: 60, which the browser's Faro transport waits out (runbook "Ingest rate limit") |
O11Y_ENV |
var | production | local |
LOGIN_BROKER_URL |
var | Handsontable login broker base URL (ADR-0007, K1) — same value as workers/api/wrangler.jsonc's own LOGIN_BROKER_URL |
GITHUB_OIDC_REPOSITORY |
var | handsontable/examples |
GITHUB_OIDC_WORKFLOW_REF |
var | <owner>/<repo>/<workflow file path>@<ref>, exact match — the deploy webhook's OIDC workflow_ref claim (ADR §B.5's "issuer, audience, repository, workflow") |
CLOUDFLARE_ACCOUNT_ID |
var | duplicates wrangler.jsonc's top-level account_id — a Worker has no runtime way to read its own account id, and GrafanaBox needs it to build the Loki bucket's R2 S3 endpoint and the Analytics Engine SQL API URL |
SERVICE_VERSION |
--var in the deploy script |
full GITHUB_SHA, same pattern as the API worker's own row below; falls back to "dev" when unset (wrangler dev) |
LOKI_S3_BUCKET |
var, optional | probe-only override of the Loki bucket name (COMMON.md probe rules); falls back to the production bucket name when unset, so wrangler.jsonc need not set it at all |
O11Y_EXPORT_SECRET |
secret | x-o11y-secret on the export destination and the deploy fallback |
SENTRY_HOOK_SECRET |
secret | Sentry internal-integration client secret |
AE_SQL_TOKEN |
secret | Analytics Engine SQL API (alert cron; also added by GrafanaBox's outbound handler to Grafana's queries, never passed to the box) |
LOKI_S3_ACCESS_KEY_ID, LOKI_S3_SECRET_ACCESS_KEY |
secrets | R2 S3 credentials, passed to the box as envVars |
SLACK_WEBHOOK_URL |
secret | alert channel; never passed to the box |
O11Y_SESSION_SECRET |
secret | HMAC key for the Worker's own __Host-o11y_session/__Host-o11y_login cookies (K1); rotating it logs every signed-in person out at once |
RUNNER_EVENTS_CLICKHOUSE_URL |
.dev.vars only |
local-mode stand-in for the Analytics Engine SQL API's URL (alert queries, §10); defaults to http://localhost:8123 when absent |
O11Y_LOCAL_MINIO_PORT, O11Y_LOCAL_CLICKHOUSE_PORT |
.dev.vars only |
host ports containers/o11y/compose.yml's minio/clickhouse are published on, reached from the box's Container via host.docker.internal; never set in production |
O11Y_LOCAL_PUBLIC_ORIGIN |
.dev.vars only |
the origin wrangler dev is actually reachable on, for Grafana's own GF_SERVER_ROOT_URL |
O11Y_SLEEP_AFTER |
--var only |
shortens GrafanaBox's 15-minute idle window ("20s") for containers/o11y/local/idle-stop.mjs; honored only under O11Y_ENV=local, ignored in production |
DEV_ADMIN |
.dev.vars only |
fail-closed local bypass of the session check |
The box reaches the Loki bucket over S3 at
https://<account-id>.eu.r2.cloudflarestorage.com with LOKI_S3_*, scoped to that bucket
only; it writes Loki data and the clean markers there. Lifecycle rules: browser/ chunks
30 d, worker/ chunks 90 d, index 90 d, state/ 30 d.
API worker additions (workers/api/wrangler.jsonc):
| Name | Kind | Value / purpose |
|---|---|---|
RUNNER_EVENTS |
Analytics Engine | dataset runner_events |
O11Y |
service binding, entrypoint O11yHeartbeat |
handsontable-demos-o11y, heartbeat() RPC for the watchdog (not an HTTP route — A-C1) |
SERVICE_VERSION |
--var in the deploy script |
full GITHUB_SHA |
SENTRY_SCOPE |
var | full | uncaught (§11) |
CF_ACCOUNT_ID |
var | GraphQL Analytics API account tag; also scopes the Analytics Engine SQL API read below |
AE_SQL_TOKEN |
secret | Analytics Engine SQL API (Account Analytics Read) — production read side of the nightly example_daily rollup (ADR-0042 §5, C-I1). Unset means reconcile.ts#queryExampleEventTotals throws instead of rolling up an empty day |
o11y-logs |
export destination name | referenced from observability.logs.destinations |
*/5 * * * * |
cron | pool.gauge, budget.gauge, o11y heartbeat check |
Authoring app: VITE_SENTRY_RELEASE (the GITHUB_SHA define) doubles as
service.version; VITE_TELEMETRY_LOCAL=1 enables the local path (§10) and is never set
for a production build; VITE_SENTRY_SCOPE = full | uncaught (§11).
Headers: x-hot-session (page-load id, browser → API worker), x-o11y-secret,
sentry-hook-signature, the __Host-o11y_session/__Host-o11y_login cookies
(/grafana/*'s own session, never forwarded to the container). Both cookies use Path=/
(the __Host- prefix requires it), so the browser also sends them to /api, /d and
the authoring app — none of those read them (the API worker reads only
Authorization/X-MCP-Secret; authoring is static), but it means the cookie header
itself must be stripped before proxying to the box (see /grafana/*'s gate row above).
x-o11y-grafana-user (set by the o11y worker for Grafana auth.proxy; stripped from
every client request), X-Scope-OrgID (Loki tenant: browser | worker).
All of these are OTLP resource attributes on every record, so Loki's otlp_config
can promote them to labels.
| Key | Values | Loki label | AE slot |
|---|---|---|---|
service.name |
demos-authoring, demos-api, demos-o11y, demos-embed |
service_name |
blob1 |
service.version |
full git SHA | no | blob2 |
deployment.environment.name |
production, local |
deployment_environment_name |
blob3 |
hot.surface |
authoring, share, embed, d, api, demo-runtime, o11y |
hot_surface |
blob4 |
hot.tier |
1, 2, static, none |
hot_tier |
blob5 |
hot.framework |
a key of config/frameworks.json (every docs-example framework is one), or none |
hot_framework |
blob6 |
hot.ht_major |
15…19, next, none |
hot_ht_major |
blob7 |
hot.outcome |
per metric, see §5; none on a record no metric describes |
hot_outcome |
blob8 |
Ingest bounds both open labels, since each distinct label tuple is a Loki stream
(5000 per tenant): a hot.framework outside the list above, or a hot.outcome
outside the set of the item's metric (none when there is none), becomes other.
A stored browser record (exception, log, event) always carries hot.outcome =
none; only a measurement's AE point keeps its metric outcome. That caps the
browser tenant at 4410 label tuples (collect 7 × 4 × 21 × 7, lite 2 × 1 × 21 × 7).
The box's Loki sets max_global_streams_per_user to 20000, over 4× that worst case;
pipeline/o11y-label-cardinality.test.mjs reads the limit from the config and fails
when the reachable tuples cross it.
Ingester memory grows with the streams that actually receive lines and the bytes
pushed, not with the limit, and a stream costs kilobytes (labels, index entry, head
block), so 20000 fits easily in the box's 4 GiB standard-1 container.
packages/runtime/src/telemetry/attrs.ts#KNOWN_FRAMEWORKS mirrors
config/frameworks.json.
Structured metadata only — never a Loki label, never an Analytics Engine index:
hot.demo_id, session.id (an in-memory page-load id), cf.ray, hot.kind (the Faro item
kind: exception, log, event, measurement).
Diagnostic tags — flat, non-dotted, never a Loki label, never an Analytics Engine
index, but hoisted to structured metadata alongside the dotted keys above (§6,
convert.ts#hoistAttributes's STRUCTURED_KEY_SET): handled, context,
sentry_event_id, versions_fetch_attempts, versions_fetch_outcome,
versions_fetch_elapsed_bucket, versions_fetch_online, api_base_origin,
net_effective_type. Each is a boolean flag, an enum-like/bucketed value, an
opaque platform id, or the reporting call site's own name — never user or
request content.
AE-only transport keys — never a Loki label, never structured metadata, never
hoisted by hoistAttributes at all: hot.bucket, hot.reason, hot.fingerprint,
hot.metric_kind, hot.ref, hot.area. Survive the same browser/ingest attribute
allowlist as every key above (so the browser can transmit them at all), but exist
only to carry an Analytics Engine column (§4/§5) through a Faro item's raw
context/attributes, read directly from the wire body by
normalise/browser-attrs.ts#readAeOnlyAttrs before hoistAttributes ever runs —
for a metric such as example.open that is never written to the inbox or Loki at
all (§6, ADR-0042).
Never sent to the o11y stack: the user pseudonym, an email, an IP, a user-agent
string, a query string or fragment, authored code (including Babel code frames), chat
text, console output, url.full, geo or ASN attributes.
One bounded exception, demo-runtime records (§6): a relayed preview message is sent
only as its §7 fingerprint shape (fingerprintShape: code frame stripped; quoted
strings, numbers, URLs, timestamps and keystroke-ladder identifiers replaced; ≤200
chars), never raw, never with its stack or URL. Unquoted prose a demo itself passes to
new Error(…) or console.error(…) survives that normalisation.
index1 = metric name (the sampling key); queries filter on index1 directly.
| Slot | Column | Meaning |
|---|---|---|
index1 |
metric |
metric name (§5) |
blob1 |
service_name |
§3 |
blob2 |
service_version |
§3 |
blob3 |
environment |
§3 |
blob4 |
surface |
§3 |
blob5 |
tier |
§3 |
blob6 |
framework |
§3 |
blob7 |
ht_major |
§3 |
blob8 |
outcome |
§3, §5 |
blob9 |
reason |
metric-specific qualifier (§5) |
blob10 |
route_class |
API route class, e.g. api/versions |
blob11 |
fingerprint |
§7 |
blob12 |
demo_id |
demo id where the metric concerns one demo |
blob13 |
model |
LLM model id |
blob14 |
provider |
upstream or import provider |
blob15 |
device |
desktop, mobile, tablet |
blob16 |
bucket |
docs or starter bucket (18.1, next) |
blob17 |
kind |
ADR-0042: docs, starter, saved, import, payload |
blob18 |
ref |
ADR-0042: guide path or starter id |
blob19 |
area |
ADR-0042: first breadcrumb element of a docs example |
blob20 |
— | unassigned |
double1 |
count |
1 per point unless pre-aggregated |
double2 |
duration_ms |
latency |
double3 |
value |
generic measurement (web-vital value, gauge level, seconds, percent) |
double4 |
usd |
cost |
double5 |
tokens_in |
LLM input tokens |
double6 |
tokens_out |
LLM output tokens |
double7 |
bytes |
payload or artifact size |
double8 |
cap |
the limit a gauge is measured against |
double9–double20 |
— | unassigned |
Reading rule: Analytics Engine samples at write and read time. Every count is
SUM(_sample_interval * double1), every percentile a weighted quantile, never COUNT().
Queries go through one helper that allowlists Analytics Engine's documented functions;
the local ClickHouse shim accepts more. Dashboard queries follow the same rule: no
SELECT DISTINCT, subqueries, joins, boolean arithmetic (sumIf instead), a string
literal compared with timestamp (wrap it in toDateTime()), and a literal
FROM runner_events rather than the plugin's $table, which expands to
default.runner_events. A multi-value variable in a predicate is written
('__all__' IN (${v:sqlstring}) OR col IN (${v:sqlstring})) with allValue set to
'__all__', so an empty option list never reaches Analytics Engine as IN ().
pipeline/o11y-dashboards.test.mjs enforces this on the query text the plugin sends.
Outcome values are the only strings allowed in blob8 for that metric.
| Metric | Emitted by | Blobs used | Doubles | Outcomes / reason |
|---|---|---|---|---|
preview.ready_ms |
browser | surface, tier, framework, ht_major, outcome, bucket | duration_ms | ready, error, timeout, abandoned |
sandpack.compile_ms |
browser; the settled (last) compile of each edit burst, closed by 2 s without a compile or compile error. The mount's compile is not sent (preview.ready_ms covers first load) |
tier, framework, ht_major, outcome | duration_ms | ok, error |
sandpack.compile_error |
browser | framework, ht_major, fingerprint, bucket | count | — |
sandpack.bundler_unreachable |
browser | ht_major | count, duration_ms | — |
preview.runtime_error |
browser | surface=demo-runtime, tier, framework, ht_major, fingerprint, reason |
count | reason: uncaught, console, network, stderr |
version.switch |
browser | framework, ht_major (to), reason (from), bucket | count | — |
bucket.resolve_ms |
browser | bucket, outcome | duration_ms | ok, error |
session.start_ms |
browser | framework, ht_major, outcome, reason | duration_ms | outcomes as session.start; reason cold, warm |
hmr.roundtrip_ms |
browser | framework, ht_major | duration_ms | — |
web_vital |
browser, beacon | surface, framework, ht_major, reason, device, demo_id | value | reason LCP, INP, CLS, TTFB |
error.uncaught |
browser, beacon | surface, fingerprint, demo_id | count | — |
error.handled |
browser, API worker | surface, route_class, fingerprint | count | — |
example.open |
browser (ADR-0042) | kind, ref, area, framework, ht_major, bucket, reason (entry) |
count | reason deep-link, picker, switch, version-switch, fork |
example.engaged, example.forked, example.shared, example.downloaded |
browser (ADR-0042) | kind, ref, area, framework, ht_major, bucket | count | — |
example.saved |
API worker (ADR-0042) | kind, ref, area, framework, ht_major, bucket | count | — |
api.request |
API worker | route_class, outcome, reason | count, duration_ms | 2xx, 3xx, 4xx, 5xx; reason = the exact status of a 5xx (503, 500), empty otherwise |
session.start |
API worker | framework, ht_major, outcome | count, duration_ms | ready, at_capacity, container_starting, boot_timeout, budget_denied, error |
session.end |
API worker | framework, reason | count, value (awake s) | reason pagehide, sleep_after, teardown_failed, budget_closed |
container.boot_ms |
API worker | framework, outcome, reason | duration_ms | ready, window_exceeded, error; reason cold, warm |
pool.gauge |
API worker */5 |
reason (live, builder) |
value (awake), cap | — |
budget.gauge |
API worker */5 |
reason (tier) | value (percent of ceiling), usd | — |
snapshot.build |
API worker | framework, outcome, reason, demo_id | count, duration_ms, bytes | ok, failed; reason inline, detached |
serve.share, serve.d, serve.embed |
API worker | outcome, demo_id | count, bytes | 2xx, 304, 4xx, 5xx |
chat.answer |
API worker | model, outcome | count, duration_ms, usd, tokens_in, tokens_out | answered, denied, error |
chat.edit |
API worker | outcome | count | proposed, applied, undone |
theme.ai |
API worker | model, outcome | count, duration_ms, usd | answered, denied, error |
import.url |
API worker | provider, outcome, reason | count, duration_ms | ok, refused, error |
payload.boot |
API worker | framework, outcome | count | ok, error |
reconcile.run |
API worker cron | outcome | count, duration_ms, usd (billing total WRITTEN this run — not a delta against the estimate; D-M15 fix round) | ok, skipped, error |
o11y.ingest |
o11y worker | reason, outcome | count, bytes | accepted, dropped, duplicate; reason = gate |
o11y.drain |
o11y worker | reason, outcome | count (objects), duration_ms (the not-ready give-up error point: how long the box was not ready), bytes, value (records dropped for being too old — reject_old_samples_max_age) |
ok, partial, error; reason backlog, visit, reopen (this batch replayed reopened keys) |
o11y.wake |
o11y worker | reason, outcome | count, duration_ms (to ready) | reason backlog, visit; outcome clean, unclean |
o11y.ae_degraded |
o11y worker | reason | count | reason start (a wake started with the ClickHouse route off), reload (re-applying it on an already-running box failed) |
o11y.backlog |
o11y worker cron | — | value (oldest age s), bytes | — |
o11y.alert |
o11y worker cron | reason (rule id), outcome | count | fired, resolved |
o11y.new_fingerprint |
o11y worker cron | fingerprint | count (one point per fingerprint the registry saw for the first time; at most 100 per tick, the rest stay in the registry; no Slack message) | — |
session.end's value is the awake seconds the cost ledger booked for the session: the sum
of every slice recordContainerUsage booked for it, each capped at the 300 s awake
window, so a hidden tab or an abandoned session is not credited with the quiet gaps
between ticks or the hours before a late teardown. Sessions whose meter predates the
running total fall back to the wall-clock span of their ticks plus the final slice. It is omitted (read back as 0) when the KV meter is already gone,
and the tier-2 panel excludes those rows. The admin panel's kill button emits no point.
o11y.ingest dropped reasons ingest_timeout and ingest_error come from the v1/logs,
deploy and hooks/sentry routes when InboxWriter.ingest exceeds its 10 s deadline
(INGEST_DEADLINE_MS) or rejects; the route answers 503 with Retry-After: 30. A commit whose
reply is lost is counted ingest_timeout, and its redelivery only as duplicate, so that batch is
never counted accepted. Cloudflare's exporter behaviour (timeout, retries and backoff,
back-pressure on the source Worker, what it treats as failure) is undocumented and remains a known
unknown, watched through the "Observability self" dashboard.
Two reasons in this table have no emitter and never produce data: session.end reason
sleep_after (nothing observes the Sandbox SDK's idle-timeout stop) and pool.gauge
reason builder (the builder pool is not metered; the */5 cron emits live only).
o11y.wake's duration_ms is wake-to-ready time — from the wake starting to the
box's first successful isReady(), sourced from wake:<wakeId>.readyMs (§8) — on both
the clean and unclean outcome. duration_ms = 0 means the box never became ready
during that wake, not a genuinely instant boot.
The point's own timestamp is a different moment: it is written at
resolution — the next backlog() call that runs resolveWakes() (§8) and finds this
wake over, not when the wake itself started. Locally, with no cron firing, that can lag
the actual wake by an hour or more; in production it lags by at most one */10 cron
period. Two wakes resolved in the same resolveWakes() call get timestamps identical to
the millisecond even though their wakes started at different times — expected, not a bug.
sandpack.compile_error counts Tier-1 compile failures from
both places they occur:
- a bundler diagnostic (
show-errorwith no frames — the module never evaluated); - the parcel pre-transpile's own babel parse failure (every starter except
vue-cli). That source never reaches the bundler and the last good render stays on screen, soSandpackRuntimereports it from the transpile catch itself (isTranspileFailure): at mount (a saved/shared/?payload=demo that does not parse — counted at once), and on the edit path for the newest push only.
The signal is the transpile catch, never the message shape: a runtime SyntaxError
(JSON.parse, new Function) is relayed by the preview like any other throw and stays
preview.runtime_error. Compile errors go through the same edit-burst collapse as
preview.runtime_error (below), keyed by kind, so a burst counts at most one — from its
final state. A compile failure of the burst's newest edit also replaces the run: the
preview never ran that code, so what it relays for the rest of the burst (a
keystroke-prefix rung still in flight, a re-render warning) is from code already typed
past and is dropped. One typed broken line = one sandpack.compile_error, no
preview.runtime_error. The next edit re-arms runtime reports. No error card and no
Sentry capture is added for the edit-path failure; the mount-path Sentry capture
(Tier1CompileError) is unchanged.
preview.runtime_error counts broken preview states, not relays. Reason console is a
console.error; a console warning is not counted (§6). The preview
re-runs on every keystroke, so one typed line relays a whole keystroke-prefix ladder
(s is not defined, se is not defined, …, then the line's real error). The browser
collapses it (apps/authoring/src/demoEventCollapse.ts) before the facade:
- an edit that re-runs the preview (a non-quiet workspace write, a file add, delete or rename) opens or extends a burst, and discards what the previous run reported;
- on Tier 1, an edit whose transpiled sandbox matches the running one (a closing
;, whitespace, a trailing comma) re-runs nothing, so the burst ends with the running sandbox's reports that have not been counted yet. If the bundler rejected that sandbox (a framelessshow-error), itssandpack.compile_erroris that result and replaces the run; a pre-transpile failure never ran, so it is not the running sandbox's; - on Tier 1, the bundler runs one compile at a time, so a run's reports can still arrive
after the next edit has been dispatched. A new run starts at the bundler's
startmessage for a pushed compile (onPushOutcome("rerun")), not at dispatch, and what the burst held until then came from the run it replaces and is dropped. A pre-transpile failure of the newest edit is kept, because no run of that edit will start. A run that never starts (a stalled or unreachable bundler) drops nothing, and the burst closes on its quiet window as usual; - 2 s (
DEMO_EDIT_SETTLE_MS) after the last edit the burst closes, and the last run's reports are emitted, one per §7 fingerprint; - outside a burst (first load, a user interaction, a Tier-2 rebuild that reports after the burst closed) a report is emitted at once;
- a fingerprint counts once until the next edit or preview mount, and at most 50
(
DEMO_COLLAPSE_CEILING) points per page load. An async report of the previous run (a timer, a rejected promise, a failed request) that fires after the next run started can still add one point. On Tier 2, where a rebuild outlasts the 2 s window, a superseded rebuild's report can land after the burst closed and count on its own. The Sentry side is not behind this collapse; its relay budgets are unchanged.
example.saved is written by the API worker when an editor Save (PATCH /api/demos/:id
with files) finishes its rebuild, because the rebuild can outlast the visitor's stay on
the page. It is attributed to the example the demo came from, resolved from the row's
forked_from (a bare demo id is followed up to 5 saved-demo hops, with a cycle guard): a
docs fork gives kind=docs, ref = the guide, area = its first breadcrumb element and
bucket from the lineage (the same values the example's example.open carried, from the
taxonomy workers/api/src/docs-taxonomy.generated.ts bundles); a starter fork gives
kind=starter, ref = the framework key; import and payload forks give their own kind and
source. Where nothing resolves (an MCP demo, a docs path missing from the bundled taxonomy,
a broken or looping lineage, a failed lineage read) it falls back to kind=saved, ref =
the demo id, and empty area/bucket. framework is the demo row's and ht_major is the
body's exampleHtMajor, the major the editor opened the demo at. blob1/blob2 name
demos-api.
What gates the count is a successful rebuild whose request carries a valid
exampleHtMajor (one of §3's ht_major values), whoever sends it. The editor sends the
field only while its own telemetry gate is open (§10). The rebuild and the point are
registered with ctx.waitUntil, so a client disconnect within the 30 s grace does not
cancel them. Every rebuild response carries exampleSaved (whether the point was
written). The API is the only emitter of example.saved; the editor never emits it.
A build that fails on the demo's own input is client input (isUserBuildError in
workers/api/src/share.ts), on any route that builds inline: POST /api/demos, PATCH /api/demos/:id, POST /api/mcp/demos and PATCH /api/mcp/demos/:id. That means one of:
- the build command exited with a code from 1 to 125;
- the install failed with
ERR_PNPM_NO_MATCHING_VERSION,ERR_PNPM_FETCH_404,ERR_PNPM_SPEC_NOT_SUPPORTED_BY_ANY_RESOLVERorERR_PNPM_BAD_PM_VERSION(a dependency the author named).
It answers 422 {"error":"<phase> failed: <detail>","code":"build_failed","detail":<the build error, one line>}. error carries the diagnostic because MCP clients read only
that field; the editor keys on code. api.request records it as 4xx, so api-5xx-rate
does not count it. snapshot.build still records failed, and the
snapshot-build-failed-rate alert is the backstop: it fires when one framework has more
than 50 % failed builds over 30 min with at least 10 failed across at least 3 distinct
demos (one author retrying one broken demo does not fire it). The stored demo is
unchanged, because the build runs before anything is written.
Everything else stays 5xx: exit code 126, 127 or 128 and above (not executable, not
found, killed by a signal), a result without an exit code, any other install failure
(ERR_PNPM_FETCH_5xx, a reset or timed-out connection), and any other throw. So does
any failure whose output names infrastructure, whatever the exit code: ENOTFOUND,
ECONNRESET, ETIMEDOUT, EAI_AGAIN, heap out of memory, signal SIGKILL, worker exited, and Next's `next/font` error or a fonts.googleapis.com fetch, anywhere in
the output. fetch failed and Failed to fetch count only on an error line
(TypeError: fetch failed), never on a bare line, because an author's build script can
print them. This last rule is a match on message text: a tool that rewords these lines
moves its failure into the 422 class, and an author who throws that text as an Error
reads as ours.
serve.share locally: under vite dev (what pnpm dev:full serves), React
StrictMode runs the share page's load effect twice, so one /share/<id> view gives 2
points. A production build gives 1 (measured on vite preview).
payload.boot records the Theme Builder hand-off at both ends:
okanderroratPOST /api/payload, when the link is minted;error(framework=other) at the playground boot,GET /api/payload/:id, when the link cannot boot: a miss (expired or never minted), a malformed id, or a KV failure. A link that boots adds no point, so each hand-off counts oneokat most.
The server-side error point is the record of a ?payload= boot that failed. The
browser shows the "expired" message for the 404 without reporting it. error.handled context=payload-boot is only the browser failing to reach the API at all (a network
error), which the server never sees.
The app never calls Faro directly; it calls one facade, implemented with Faro:
interface Telemetry {
metric(name: MetricName, values: { duration_ms?: number; value?: number; count?: number }, attrs: HotAttrs): void;
event(name: EventName, attrs: HotAttrs & Record<string, string>): void;
error(err: unknown, context: string, attrs?: HotAttrs): void; // handled errors
pageLoadId(): string; // minted in memory at page load
}Faro configuration: session tracking disabled; only the errors and web-vitals
instrumentations; no user meta; the facade sets session.id = page-load id on every
item; transport to same-origin /telemetry/collect; beforeSend = scrubTelemetry then
the shared noise gates.
Faro's global dedupe stays on for pushError: consecutive identical
exceptions (same type, message, stack, context) collapse to one item, with no time
window. So error.uncaught/error.handled counts — including the dashboard panels
built on them — are reports, not occurrences: a burst of identical errors counts as
one. Sentry, not this pipeline, is the occurrence counter (ADR §F.3).
What the o11y worker does with each Faro item at ingest:
| Faro item | Analytics Engine | Inbox (Loki browser tenant) |
|---|---|---|
measurement whose type is a browser metric in §5 |
one point | none |
web-vitals measurement |
one web_vital point per vital |
none |
exception with context.handled = "true" |
error.handled |
one log record, symbolicated at drain |
| other exception | error.uncaught |
one log record, symbolicated at drain |
event named example.* |
one point | none |
| other event, log | — | one log record |
Demo-runtime records. Each report that survives the collapse (§5)
is also one handled Faro exception, through Telemetry.error with context
demo-runtime. Its type is DemoError, DemoUnhandledRejection, DemoConsoleError,
DemoNetworkError or DemoStderr. Its value is the §7 fingerprint shape (§3). It has
no stack, and the §3 labels include hot.surface = demo-runtime. At ingest it becomes
an error.handled point (surface demo-runtime, never feeding the new-fingerprint
alert) and one Loki line, which the "Recent demo-runtime errors" panels read with
{hot_surface="demo-runtime"} | hot_kind="exception". A console-warn report is
neither counted nor recorded: a warning is context, not a fault (DEV-2539), and
Handsontable's own load-time notices would otherwise count on every preview load. Faro's
pushError dedupe applies, so an identical record in two consecutive bursts is sent
once. The preview.runtime_error metric, not the line count, is the counter.
A Faro measurement or web-vitals item (this table's scope — the item
kinds normalise/faro.ts handles) is AE-only; Loki holds logs, events and exceptions.
This was a contract/ADR mismatch, not an implementation bug — ADR §F.1 ("Counts and
latencies go to Analytics Engine; Loki holds the text") already said this; requiring a
stored record for every measurement too would make measurements
~99% of the browser Loki tenant's lines and drained bytes for no reader:
no dashboard panel parses a measurement's {"duration_ms":N}-shaped body, so the AE
point was always the only consumer. A measurement/web-vitals item still gets a
hash-only ingestItem (no record) so a retried/redelivered batch cannot double-write
its Analytics Engine point — the same dedupe-only shape an example.* event already
used above.
The lite-beacon path (§9, workers/o11y/src/lite.ts) follows the same rule: only a t:"err" beacon
stores a record, and a t:"vital" beacon is a hash-only item that writes its web_vital point alone.
fingerprint(context, message) = <context>:<16 hex chars of FNV-1a 64 over the normalised message>, synchronous and identical in browser and Worker; the normalisation
is normalizeMonitorMessage plus stripCodeFrame. For hot.surface = demo-runtime,
keystroke-ladder shapes ("<identifier> is not defined" and similar) collapse to one
fingerprint per shape. Demo-runtime fingerprints never feed the new-fingerprint alert.
context is caller-chosen (typically hot.surface or a metric name) and MAY itself
contain further :-separated segments, e.g. docs-example-load:fetch or
npm-registry:version-exists — a call-site path, not always a single flat token. A
client-supplied fingerprint (Faro's own payload.fingerprint wire field, or
context["hot.fingerprint"]) is trusted only when it passes the ONE shared validator
(isValidFingerprint, packages/runtime/src/telemetry/fingerprint.ts — also used by
normalise/faro.ts's resolveFingerprint and normalise/otlp.ts's
apiFingerprintFeed, never a second, independently drifting copy of the shape):
anchor on the LAST :, followed by exactly 16 lowercase hex characters, with zero or
more earlier :-separated segments in context, each drawn from [a-z][a-z0-9._-]*;
the whole context half is capped at 128 characters. A value that does not match is
discarded, never stored or forwarded to Slack verbatim.
Normalised records are OTLP JSON log records (resourceLogs shape), one tenant per
object:
inbox/<tenant>/<yyyy-mm-dd>/<hh>/<seq:012d>.ndjson.gz # inbox bucket; tenant = browser | worker
state/wakes/<wakeId>/clean # Loki bucket; written by the box on a clean stop
<seq> is a counter in InboxWriter storage, incremented in the transaction that
records the key. Each NDJSON line is one OTLP ResourceLogs object. The arrival time is
not part of the record (it would break dedupe); it lives on the storage row. The
dedupe hash is computed over the decoded, scrubbed record before timestamps are stamped.
InboxWriter storage:
| Key | Value |
|---|---|
seq |
last issued sequence |
row:<n:012d> |
pending records with their arrival time, ≤ 1 MB per row. <n> is zero-padded to 12 digits so native ascending key order equals arrival order (pack.ts#collectRowBatch) |
key:<inbox key> |
written | provisional:<wakeId> | rejected:<reason> — never committed (see done:, below) |
done:<inbox key> |
1 — a committed key, moved OUT of key: on commit (same write that deletes key:<inbox key>) |
hash:<yyyymmdd>:<sha256> |
first-seen epoch ms; 24 h window, checked across the current and previous UTC-day bucket |
fp:<fingerprint> |
first-seen epoch ms (exact registry behind the new-fingerprint detection, charted through o11y.new_fingerprint, §5) |
fpts:<firstSeenMs:015d>:<fingerprint> |
same first-seen epoch ms as its fp: twin — a time-ordered secondary index (G1 fix round, B-C1/A-I1 remainder) so the new-fingerprint alert can do a bounded start/end range read instead of listing the whole (alphabetically, not chronologically, ordered) fp: prefix every tick. Written/deleted together with its fp: twin, always |
admission:<windowStartMs:015d> |
{ fp, fpDropped, hash, hashDropped } — one row per 10-minute window: how many NEW fp: fingerprints (budget 200) and NEW hash: dedupe entries (budget 5,000) were admitted, and how many were dropped instead of stored. Browser ingest is only rate-limited per IP, so this is the global cap on a distributed flood. A dropped fingerprint's record is still stored; a dropped hash only means a redelivery inside the 24 h window may be stored twice. Rows older than 1 h are pruned; the admission-overflow alert reads the dropped counts (Slack fire/resolve plus the o11y.alert point) |
fpCount / fpLapSeen |
the fp: registry size counter and the running count of the prune lap in progress. The counter is incremented at ingest, lowered by TTL deletes, and re-measured from fpLapSeen when a full prune lap finishes. While it exceeds 50,000, each cron tick evicts the oldest entries (via fpts:), at most 5,000 per tick. A flood can therefore push real fingerprints out, at the cost of a re-alert if one reappears |
alert:<rule> |
{ state: firing | resolved, since, lastNotified } |
alertMeta:newFingerprintCursorKey / alertMeta:newFingerprintAnnouncedKeys |
the new-fingerprint detection's keyset cursor (an fpts: key) and a JSON array of the fpts: keys it already reported past that cursor. The cursor lags 2 minutes behind the tick, so the next tick reads recent entries again; the reported set makes sure each fingerprint becomes exactly one o11y.new_fingerprint point (F35). Both are saved only after the point write returns: a write that throws or rejects saves neither, and the next tick announces the fingerprint again (a binding that returns no result cannot report failure) |
wake:<wakeId> |
{ startedAt, reason, over: boolean, readyMs? } — over when a newer wake started or the container is not running; deleted once fully resolved (see below). readyMs (F8) is wake-to-ready time in ms, written once by InboxWriterApi.recordWakeReady(wakeId, readyMs) — called by GrafanaBox on the wake's first successful isReady(), first call wins, a no-op for an already-resolved (deleted) wake — and copied onto the resolved o11y.wake point as duration_ms (§5); absent while the box has not yet become ready |
rejectedEvent:<ms:015d>:<inbox key> |
rejection reason (string) — a chronological audit/alert log (G1 fix round, row 19 / B-C1/A-I1 remainder), written by both a full rejection (ledger.ts#rejectKey) and a partial one (ledger.ts#recordPartialReject, see below). The rejected-inbox-key alert fires on a RECENT (last hour) count here, not on rejectedKeyCount()'s never-pruned total, so it resolves once rejections stop instead of firing forever after the first one ever seen |
drainsPaused |
boolean (o11y spend cap). Set on every alert tick from the o11y-spend-cap result, so it follows the runtime budget override. The cron decides its backlog wake only after the alerts ran (F37). While it is set, the cron never wakes the box for the backlog, and GrafanaBox.drainStep pushes nothing, whatever woke the box. A visit wake still starts the box and serves Grafana (ADR §G) |
heartbeat |
{ lastCron, lastIngest } |
Limits: records over 256 KB are dropped; requests to Loki carry at most 1 MB
decompressed. Every free-text string (Faro/OTLP body,
message, attribute value) is truncated to this same 256 KB before any scrub/redact
regex runs over it (SCRUB_TEXT_MAX_CHARS, packages/runtime/src/telemetry/scrub.ts)
— a ReDoS defense-in-depth independent of each pattern also being made linear-time.
The fingerprint normaliser (normalizeMonitorMessage, §7) is bounded separately, to a
much smaller 4096 chars, since its own output is always sliced to 200 chars regardless.
The pre-scrub truncation above shares the same 256 KB limit as the
record-size drop, so a field that actually gets truncated still leaves the record over
the drop cap — an oversize record is genuinely dropped, never truncated down to fit.
Only the inbox/Loki record and its first-seen fp: entry are lost this way; the item's
Analytics Engine point (e.g. error.uncaught/error.handled) is still written — dropping
is a size decision at the inbox/Loki layer only, not an ingest-wide refusal.
Bounded storage (the resolve/drain/backlog paths must never scan committed history):
- A
key:entry only ever holds a live state (written,provisional:<wakeId>, or a genuinerejected:<reason>). The moment a key is confirmed clean-committed, itskey:<inbox key>entry is deleted and adone:<inbox key>marker takes its place in the same write —key:therefore never grows with committed history, only with what is currently open or in flight.done:entries are pruned once their embedded date is older than the 7-day inbox-object retention (the ten-minute cron path, viaInboxWriter.backlog()), using a boundedstart/endrange delete (done:inbox/<tenant>/through the cutoff date), never a full-prefix scan. - A
wake:<wakeId>entry is deleted as soon asresolveOverWakesfully resolves it (every provisional key under it moved towrittenordone:) — not merely flagged. Thewake:prefix therefore only ever holds the (at most one) currently-active wake plus any wake whose resolution crashed mid-way, never all-time history. hash:is bucketed by UTC calendar day (hash:<yyyymmdd>:<sha256>) instead of one flat set; a dedupe check reads exactly the current and previous day's buckets (the 24 h window can never span more than those two), and stale buckets (2+ days old) are pruned with a bounded range delete on the same cron path.fp:keeps its flat shape (nothing reads it by date range), but is swept by a bounded, cursor-paginated TTL prune (default 90 days) on the same cron path, so it does not grow forever either.POST /grafana/_o11y/reopen's window WIDTH (toMs - fromMs) is capped to the same 7-day retention (ledger.ts#reopenWindowExceedsRetention) — a wider request is refused up front. This does not require the window itself to be recent: a[from, to)pair from long ago, 7 days wide or narrower, is accepted too, it just finds nothing to reopen, sincedone:/hash:past the retention window are already pruned (Loki's ownreject_old_samples_max_ageis 7d too). The route also accepts a request whosecontent-typemerely containsapplication/jsonas a parameter (e.g. with a charset), not only an exact match — still enough to force a CORS preflight for any cross-origin caller (the CSRF hardening this exists for).- A manual reopen of a committed key reads
done:, moves it back tokey:<inbox key> = written, and deletes thedone:entry.
Additional bounded-storage invariants:
- Pack alarm, bounded. The alarm pages
row:in small chunks rather than loading every pending row into memory before packing, accumulating up to one packed object's own ~4 MB budget per round, looping until the backlog is drained orMAX_OBJECTS_PER_ALARM(25) objects have been packed this invocation. Seepack.ts#collectRowBatch. - DO storage 128-key batch limit. Cloudflare's SQLite-backed Durable Object
storage API caps
get/put/deleteat 128 keys/pairs per call (https://developers.cloudflare.com/durable-objects/api/storage-api/, fetched 2026-09-24: "Supports up to 128 keys at a time" / "up to 128 key-value pairs at a time"). Every multi-key call inInboxWriterchunks throughstorage.ts'sgetManyChunked/putChunked/deleteChunked—checkDuplicates,newFingerprintWrites,pruneLedger/pruneHashBuckets/pruneFingerprintRegistry,finalizeWakeResolution,markKeysProvisional,reopenWindow,commitPackedObject, andingest's own transactionput.finalizeWakeResolution/reopenWindowrun their whole put+delete sequence inside onestorage.transaction()— chunking alone, without that, would let a crash between chunks leave a partial write. - Prune throughput.
hash:/done:prune batch sizes are 5,000 rows/tick (still chunked to 128 per actualdelete()call) — ADR §D's own 10× headroom projects ~220,000 worker records/day, which 500/tick × 144 ten-minute ticks/day (72,000/day) falls behind at roughly 3× today's traffic.rejected:key:entries are pruned too, past the same 7-day retention (filtered by value, sincekey:mixes live and rejected states chronologically — seeledger.ts#pruneLedger). - Drain partial-400 durability. A key with at least one chunk
accepted (2xx) and at least one chunk permanently rejected (400) now stays
provisional(notrejected) — its accepted content follows the normal written→provisional→committed path, so an unclean stop before Loki's local flush still triggers an automatic replay instead of being silently unrecoverable except by manual reopen. Only a key with ZERO accepted chunks staysrejected. Seedrain.ts#drainKey's own doc comment. - Drain refusals. A
429whose body names Loki's stream limit (Maximum active stream limit exceeded) is never retried (a retry can answer 204 with the excess streams dropped) and never rejects: the table may have been filled by earlier keys, so that key is deferred and stayswritten, with norejectedEvent. Its tenant is then excluded for the rest of the wake (streamLimitedTenantsin the box's storage;nextWrittenKeyspages past it), so the batch fills with the other tenant's keys, and oneo11y.drain.stream_limitwarning line names the tenant and Loki's message. Any other429, and a5xx, stays transient and stops the batch. An inbox read that throws defers only that key; the rest of the batch still pushes and commits. So does a source-map read that still fails after two in-call retries (deferral: "map_fetch_error", logged aso11y.drain.error): the key stayswrittenand nothing is pushed for it, so a replay never sends an unsymbolicated copy first. The ledger counts no attempts, so the deferral lasts only while the key's inbox hour is under 6 hours old (MAP_RETRY_MAX_AGE_MS); after that the key pushes with its frames as they are, so a replay is byte-stable only once the maps are final. A batch of only deferred keys ends the wake's drain once no un-excluded tenant has keys left. Either deferral (inbox read or source-map read) also excludes that key for the rest of the wake (deferredKeysin the box's storage, reset by a new wake;nextWrittenKeysskips it), so a deferred key is not re-read on every step and cannot starve the batch: the keys behind it still drain. When only deferred keys remain and a Grafana visitor keeps the box up, the drain clears that set and rechecks after 60 s, like the spend-cap pause; with no visitor the box stops as for an empty backlog. - Symbolication read caps. One inbox object reads at most 32 distinct maps
(
MAX_MAP_KEYS_PER_CALL, first-seen order), one body adds at most 8 of them (MAX_NEW_MAP_KEYS_PER_BODY), and at most 128 frames are looked up per body (MAX_FRAMES_PER_BODY), so a drain step of 10 objects stays at a few hundred of the Workers limit of 10,000 subrequests. Frames past a cap stay byte-for-byte and are reported aso11y.symbolicate.skipwith reasonover_cap, plus one aggregateover_capline with the call's cappedframesandkeysthatMAX_SKIP_REPORTSnever suppresses. Before any read, each distinctservice.versionthat has a resolvable frame is listed once (sourcemaps/<version>/, paged through the R2 cursor); only a key the listing shows is admitted, so a forged frame path costs no read and is reported asno_map.service.versionis client-supplied, so a call lists at most 8 versions (MAX_LISTED_VERSIONS_PER_CALL); frames of the rest stay as they were and are reported per version prefix asover_version_cap. A list that throws (or runs past 5 pages) is reported aslist_error; while the key is under 6 hours old (MAP_RETRY_MAX_AGE_MS) the object is deferred like a failing map read, so a replay never admits by the caps what the first pass admitted by the listing, and past that age the version falls back to the caps above. The per-call version cap is a stated residual risk, not closed: at least 8 forged requests with distinctservice.versionpacked ahead of a real one can still push that real version toover_version_cap.
POST /telemetry/lite, a JSON body, ≤ 2 KB, sent with navigator.sendBeacon(url, json) —
a plain string, not a Blob, so the browser sends its own default
Content-Type: text/plain;charset=UTF-8, never application/json (the
route itself does not check or require a content type either way, so this is a fact about
what ships on the wire, not a gate). The body itself is still JSON:
{"v":1,"t":"err","s":"embed","demo":"r-react-18-0-0","ht":"18","fw":"react","n":"TypeError","m":"<normalized, ≤500 chars>","st":"<stack, ≤2000 chars>","val":null,"dev":"desktop","ts":1695463200000,"id":"a1b2c3d4"}t = err | vital; s = embed | d; for vital, n is LCP | INP | CLS |
TTFB and val carries the value. No page path: docs pages send no referrer. Vitals are
sampled at 10 % per page view, decided once per page; errors are sent up to the
monitor.ts event ceiling. The o11y worker converts beacons with beaconToRecord, a sibling of the Faro
converter faroItemToRecord (packages/runtime/src/telemetry/convert.ts) that shares its
clamp of ts to the receive time ± 5 minutes and its resource-attribute handling but
builds the record from the beacon's own fields (workers/o11y/src/lite.ts).
The dedupe hash covers the whole converted record (body with message and stack, attributes)
plus the raw ts and, when present, id: a per-beacon random value, used only in the
dedupe hash.
| Piece | Local stand-in |
|---|---|
RUNNER_EVENTS |
ClickHouse at http://localhost:8123, table runner_events with the §4 columns plus timestamp and _sample_interval (always 1), DDL in containers/o11y/local/clickhouse-init.sql |
| Loki S3 | Miniflare's local S3 endpoint for R2, or MinIO from containers/o11y/compose.yml |
| Grafana session (K1) | DEV_ADMIN in workers/o11y/.dev.vars; a real broker login also works locally (LOGIN_BROKER_URL defaults to the production broker) |
| Cloudflare OTLP export | fixtures in pipeline/fixtures/otlp/ (scrubbed sandbox-probe captures plus hand-built edge cases), replayed by scripts/o11y-replay-fixtures.mjs |
| Slack | scripts/o11y-slack-capture.mjs, a local HTTP capture server started by pnpm dev:full (NOT by pnpm o11y:dev, which doesn't start it — point SLACK_WEBHOOK_URL at your own instance if you need one from the standalone o11y-only command); prints and keeps the last 50 posts, GET /_captured to inspect |
deployment.environment.name = local; the production o11y worker drops local data.
Local telemetry gate in the browser. Faro runs on the local path only when the build
was made with VITE_TELEMETRY_LOCAL=1 and the host is localhost or 127.0.0.1. It
checks neither import.meta.env.DEV nor navigator.webdriver, because Playwright serves
a production vite preview under automation. The production gate (resolveReporting)
is unchanged and stays closed under automation. A production build never sets the flag,
and the post-build leak check fails if the local path survives into it.
SENTRY_SCOPE / VITE_SENTRY_SCOPE = full (default): handled diagnostic reports go to
Sentry and the new stack. uncaught: they go only to the new stack; Sentry keeps
errors that escape a handler (browser onerror, unhandledrejection,
Sentry.ErrorBoundary; Worker fetch catch-all, DO alarms, cron, snapshot-job failures)
and the budget-alert captureMessage. The launch plan flips to uncaught after the
pipeline is seen working in production.