Skip to content

Latest commit

 

History

History
713 lines (630 loc) · 53.7 KB

File metadata and controls

713 lines (630 loc) · 53.7 KB

Observability contract

The names, shapes and slot positions every observability component shares (ADR-0041 revision 3, ADR-0042). Permanent document: it outlives the implementation task board and is what a dashboard author or a new emitter reads.

Implemented once, in packages/runtime/src/telemetry/, exported as @handsontable/demo-runtime/telemetry — the subpath-export pattern monitor-inject.ts already uses for @handsontable/demo-runtime/monitor. The module is pure (no DOM, no Cloudflare imports), so the API worker, the o11y worker, the authoring app and pipeline/ tests import the same definitions. pipeline/telemetry-contract.test.mjs parses the tables below and fails when the module disagrees with this file.

Changing this file: in its own small PR, together with the module, before any code relies on the change. Metric rows and slots are append-only: never reuse, move or rename a slot or a metric name. Analytics Engine columns are positional, so a moved slot silently corrupts every stored row and no type checks it.

1. Deployables, routes, ports

Name Path Notes
handsontable-demos-o11y workers/o11y/ Worker + InboxWriter DO + GrafanaBox Container class; workers_dev: false, preview_urls: false
Grafana box image containers/o11y/ Loki + Grafana only

Routes on demos.handsontable.com, owned by the o11y worker, passed as --routes flags in its deploy script (ADR-0020), never in wrangler.jsonc:

Route Purpose Gate (ADR-0041 §B.5)
POST /telemetry/collect Faro payloads from the authoring app host/env, bot filter, caps, kind allowlist, rate limit, server scrub
POST /telemetry/lite lite beacon from /d and /embed same as collect
POST /telemetry/v1/logs Cloudflare OTLP log export x-o11y-secret
POST /telemetry/deploy deploy event from CI GitHub OIDC token, x-o11y-secret fallback
POST /telemetry/hooks/sentry Sentry issue-alert webhook sentry-hook-signature HMAC
GET /grafana/_o11y/login start a broker sign-in per-IP rate limit; mints the __Host-o11y_login nonce cookie
GET /grafana/_o11y/callback broker return; static page, no server-side check of its own none (fix round M6: the __Host-o11y_login cookie + the live /broker/userinfo call happen on the session row below, not here — this route only serves a hash-pinned static page)
POST /grafana/_o11y/session mint the __Host-o11y_session cookie per-IP rate limit; __Host-o11y_login cookie (nonce bound), same-origin Origin, @handsontable.com broker identity
GET /grafana/_o11y/logout a same-origin sign-out page (fix round M5) none; the page itself fires the POST below
POST /grafana/_o11y/logout clear both cookies same-origin Origin
/grafana/* Grafana UI, waking page the Worker's own session cookie (O11Y_SESSION_SECRET, K1 — not Access)
POST /grafana/_o11y/reopen manual ledger re-open session cookie, same-origin Origin
GET /grafana/_o11y/admin/<name> ADR-0043 read forwarder (after launch) session cookie, name allowlist, GET only

There is no trace route: traces are not exported (ADR-0041 §C.4).

Ports inside the Grafana box, reached only through GrafanaBox.containerFetch:

Port Service
3000 Grafana (served from sub-path /grafana/)
3100 Loki HTTP (/otlp/v1/logs, /ready, /metrics)

2. Bindings, variables, secrets

o11y worker (workers/o11y/wrangler.jsonc):

Name Kind Value / purpose
INBOX_WRITER Durable Object class InboxWriter, one instance main, .jurisdiction("eu"); owns inbox keys, ledger, dedupe set, fingerprint registry, alert state
GRAFANA_BOX Durable Object + Container class GrafanaBox, one instance box, .jurisdiction("eu"), container jurisdiction: "eu"
O11Y_INBOX R2 bucket handsontable-demos-o11y-inbox (EU)
O11Y_LOKI_STATE R2 bucket handsontable-demos-o11y-loki (EU); the Worker reads only state/wakes/<wakeId>/clean markers
O11Y_MAPS R2 bucket handsontable-demos-o11y-maps (EU)
RUNNER_EVENTS Analytics Engine dataset runner_events
API service binding handsontable-demos-api (o11y usage metering, o11y spend, later AdminReads)
RATE_LIMITER Rate Limiting binding gates POST /telemetry/collect and POST /telemetry/lite (ADR §B.5): 100 requests / 60 s per cf-connecting-ip, 429 with Retry-After: 60, which the browser's Faro transport waits out (runbook "Ingest rate limit")
O11Y_ENV var production | local
LOGIN_BROKER_URL var Handsontable login broker base URL (ADR-0007, K1) — same value as workers/api/wrangler.jsonc's own LOGIN_BROKER_URL
GITHUB_OIDC_REPOSITORY var handsontable/examples
GITHUB_OIDC_WORKFLOW_REF var <owner>/<repo>/<workflow file path>@<ref>, exact match — the deploy webhook's OIDC workflow_ref claim (ADR §B.5's "issuer, audience, repository, workflow")
CLOUDFLARE_ACCOUNT_ID var duplicates wrangler.jsonc's top-level account_id — a Worker has no runtime way to read its own account id, and GrafanaBox needs it to build the Loki bucket's R2 S3 endpoint and the Analytics Engine SQL API URL
SERVICE_VERSION --var in the deploy script full GITHUB_SHA, same pattern as the API worker's own row below; falls back to "dev" when unset (wrangler dev)
LOKI_S3_BUCKET var, optional probe-only override of the Loki bucket name (COMMON.md probe rules); falls back to the production bucket name when unset, so wrangler.jsonc need not set it at all
O11Y_EXPORT_SECRET secret x-o11y-secret on the export destination and the deploy fallback
SENTRY_HOOK_SECRET secret Sentry internal-integration client secret
AE_SQL_TOKEN secret Analytics Engine SQL API (alert cron; also added by GrafanaBox's outbound handler to Grafana's queries, never passed to the box)
LOKI_S3_ACCESS_KEY_ID, LOKI_S3_SECRET_ACCESS_KEY secrets R2 S3 credentials, passed to the box as envVars
SLACK_WEBHOOK_URL secret alert channel; never passed to the box
O11Y_SESSION_SECRET secret HMAC key for the Worker's own __Host-o11y_session/__Host-o11y_login cookies (K1); rotating it logs every signed-in person out at once
RUNNER_EVENTS_CLICKHOUSE_URL .dev.vars only local-mode stand-in for the Analytics Engine SQL API's URL (alert queries, §10); defaults to http://localhost:8123 when absent
O11Y_LOCAL_MINIO_PORT, O11Y_LOCAL_CLICKHOUSE_PORT .dev.vars only host ports containers/o11y/compose.yml's minio/clickhouse are published on, reached from the box's Container via host.docker.internal; never set in production
O11Y_LOCAL_PUBLIC_ORIGIN .dev.vars only the origin wrangler dev is actually reachable on, for Grafana's own GF_SERVER_ROOT_URL
O11Y_SLEEP_AFTER --var only shortens GrafanaBox's 15-minute idle window ("20s") for containers/o11y/local/idle-stop.mjs; honored only under O11Y_ENV=local, ignored in production
DEV_ADMIN .dev.vars only fail-closed local bypass of the session check

The box reaches the Loki bucket over S3 at https://<account-id>.eu.r2.cloudflarestorage.com with LOKI_S3_*, scoped to that bucket only; it writes Loki data and the clean markers there. Lifecycle rules: browser/ chunks 30 d, worker/ chunks 90 d, index 90 d, state/ 30 d.

API worker additions (workers/api/wrangler.jsonc):

Name Kind Value / purpose
RUNNER_EVENTS Analytics Engine dataset runner_events
O11Y service binding, entrypoint O11yHeartbeat handsontable-demos-o11y, heartbeat() RPC for the watchdog (not an HTTP route — A-C1)
SERVICE_VERSION --var in the deploy script full GITHUB_SHA
SENTRY_SCOPE var full | uncaught (§11)
CF_ACCOUNT_ID var GraphQL Analytics API account tag; also scopes the Analytics Engine SQL API read below
AE_SQL_TOKEN secret Analytics Engine SQL API (Account Analytics Read) — production read side of the nightly example_daily rollup (ADR-0042 §5, C-I1). Unset means reconcile.ts#queryExampleEventTotals throws instead of rolling up an empty day
o11y-logs export destination name referenced from observability.logs.destinations
*/5 * * * * cron pool.gauge, budget.gauge, o11y heartbeat check

Authoring app: VITE_SENTRY_RELEASE (the GITHUB_SHA define) doubles as service.version; VITE_TELEMETRY_LOCAL=1 enables the local path (§10) and is never set for a production build; VITE_SENTRY_SCOPE = full | uncaught (§11).

Headers: x-hot-session (page-load id, browser → API worker), x-o11y-secret, sentry-hook-signature, the __Host-o11y_session/__Host-o11y_login cookies (/grafana/*'s own session, never forwarded to the container). Both cookies use Path=/ (the __Host- prefix requires it), so the browser also sends them to /api, /d and the authoring app — none of those read them (the API worker reads only Authorization/X-MCP-Secret; authoring is static), but it means the cookie header itself must be stripped before proxying to the box (see /grafana/*'s gate row above). x-o11y-grafana-user (set by the o11y worker for Grafana auth.proxy; stripped from every client request), X-Scope-OrgID (Loki tenant: browser | worker).

3. Attributes

All of these are OTLP resource attributes on every record, so Loki's otlp_config can promote them to labels.

Key Values Loki label AE slot
service.name demos-authoring, demos-api, demos-o11y, demos-embed service_name blob1
service.version full git SHA no blob2
deployment.environment.name production, local deployment_environment_name blob3
hot.surface authoring, share, embed, d, api, demo-runtime, o11y hot_surface blob4
hot.tier 1, 2, static, none hot_tier blob5
hot.framework a key of config/frameworks.json (every docs-example framework is one), or none hot_framework blob6
hot.ht_major 15…19, next, none hot_ht_major blob7
hot.outcome per metric, see §5; none on a record no metric describes hot_outcome blob8

Ingest bounds both open labels, since each distinct label tuple is a Loki stream (5000 per tenant): a hot.framework outside the list above, or a hot.outcome outside the set of the item's metric (none when there is none), becomes other. A stored browser record (exception, log, event) always carries hot.outcome = none; only a measurement's AE point keeps its metric outcome. That caps the browser tenant at 4410 label tuples (collect 7 × 4 × 21 × 7, lite 2 × 1 × 21 × 7). The box's Loki sets max_global_streams_per_user to 20000, over 4× that worst case; pipeline/o11y-label-cardinality.test.mjs reads the limit from the config and fails when the reachable tuples cross it. Ingester memory grows with the streams that actually receive lines and the bytes pushed, not with the limit, and a stream costs kilobytes (labels, index entry, head block), so 20000 fits easily in the box's 4 GiB standard-1 container. packages/runtime/src/telemetry/attrs.ts#KNOWN_FRAMEWORKS mirrors config/frameworks.json.

Structured metadata only — never a Loki label, never an Analytics Engine index: hot.demo_id, session.id (an in-memory page-load id), cf.ray, hot.kind (the Faro item kind: exception, log, event, measurement).

Diagnostic tags — flat, non-dotted, never a Loki label, never an Analytics Engine index, but hoisted to structured metadata alongside the dotted keys above (§6, convert.ts#hoistAttributes's STRUCTURED_KEY_SET): handled, context, sentry_event_id, versions_fetch_attempts, versions_fetch_outcome, versions_fetch_elapsed_bucket, versions_fetch_online, api_base_origin, net_effective_type. Each is a boolean flag, an enum-like/bucketed value, an opaque platform id, or the reporting call site's own name — never user or request content.

AE-only transport keys — never a Loki label, never structured metadata, never hoisted by hoistAttributes at all: hot.bucket, hot.reason, hot.fingerprint, hot.metric_kind, hot.ref, hot.area. Survive the same browser/ingest attribute allowlist as every key above (so the browser can transmit them at all), but exist only to carry an Analytics Engine column (§4/§5) through a Faro item's raw context/attributes, read directly from the wire body by normalise/browser-attrs.ts#readAeOnlyAttrs before hoistAttributes ever runs — for a metric such as example.open that is never written to the inbox or Loki at all (§6, ADR-0042).

Never sent to the o11y stack: the user pseudonym, an email, an IP, a user-agent string, a query string or fragment, authored code (including Babel code frames), chat text, console output, url.full, geo or ASN attributes.

One bounded exception, demo-runtime records (§6): a relayed preview message is sent only as its §7 fingerprint shape (fingerprintShape: code frame stripped; quoted strings, numbers, URLs, timestamps and keystroke-ladder identifiers replaced; ≤200 chars), never raw, never with its stack or URL. Unquoted prose a demo itself passes to new Error(…) or console.error(…) survives that normalisation.

4. Analytics Engine layout (runner_events)

index1 = metric name (the sampling key); queries filter on index1 directly.

Slot Column Meaning
index1 metric metric name (§5)
blob1 service_name §3
blob2 service_version §3
blob3 environment §3
blob4 surface §3
blob5 tier §3
blob6 framework §3
blob7 ht_major §3
blob8 outcome §3, §5
blob9 reason metric-specific qualifier (§5)
blob10 route_class API route class, e.g. api/versions
blob11 fingerprint §7
blob12 demo_id demo id where the metric concerns one demo
blob13 model LLM model id
blob14 provider upstream or import provider
blob15 device desktop, mobile, tablet
blob16 bucket docs or starter bucket (18.1, next)
blob17 kind ADR-0042: docs, starter, saved, import, payload
blob18 ref ADR-0042: guide path or starter id
blob19 area ADR-0042: first breadcrumb element of a docs example
blob20 — unassigned
double1 count 1 per point unless pre-aggregated
double2 duration_ms latency
double3 value generic measurement (web-vital value, gauge level, seconds, percent)
double4 usd cost
double5 tokens_in LLM input tokens
double6 tokens_out LLM output tokens
double7 bytes payload or artifact size
double8 cap the limit a gauge is measured against
double9–double20 — unassigned

Reading rule: Analytics Engine samples at write and read time. Every count is SUM(_sample_interval * double1), every percentile a weighted quantile, never COUNT(). Queries go through one helper that allowlists Analytics Engine's documented functions; the local ClickHouse shim accepts more. Dashboard queries follow the same rule: no SELECT DISTINCT, subqueries, joins, boolean arithmetic (sumIf instead), a string literal compared with timestamp (wrap it in toDateTime()), and a literal FROM runner_events rather than the plugin's $table, which expands to default.runner_events. A multi-value variable in a predicate is written ('__all__' IN (${v:sqlstring}) OR col IN (${v:sqlstring})) with allValue set to '__all__', so an empty option list never reaches Analytics Engine as IN (). pipeline/o11y-dashboards.test.mjs enforces this on the query text the plugin sends.

5. Metric registry

Outcome values are the only strings allowed in blob8 for that metric.

Metric Emitted by Blobs used Doubles Outcomes / reason
preview.ready_ms browser surface, tier, framework, ht_major, outcome, bucket duration_ms ready, error, timeout, abandoned
sandpack.compile_ms browser; the settled (last) compile of each edit burst, closed by 2 s without a compile or compile error. The mount's compile is not sent (preview.ready_ms covers first load) tier, framework, ht_major, outcome duration_ms ok, error
sandpack.compile_error browser framework, ht_major, fingerprint, bucket count —
sandpack.bundler_unreachable browser ht_major count, duration_ms —
preview.runtime_error browser surface=demo-runtime, tier, framework, ht_major, fingerprint, reason count reason: uncaught, console, network, stderr
version.switch browser framework, ht_major (to), reason (from), bucket count —
bucket.resolve_ms browser bucket, outcome duration_ms ok, error
session.start_ms browser framework, ht_major, outcome, reason duration_ms outcomes as session.start; reason cold, warm
hmr.roundtrip_ms browser framework, ht_major duration_ms —
web_vital browser, beacon surface, framework, ht_major, reason, device, demo_id value reason LCP, INP, CLS, TTFB
error.uncaught browser, beacon surface, fingerprint, demo_id count —
error.handled browser, API worker surface, route_class, fingerprint count —
example.open browser (ADR-0042) kind, ref, area, framework, ht_major, bucket, reason (entry) count reason deep-link, picker, switch, version-switch, fork
example.engaged, example.forked, example.shared, example.downloaded browser (ADR-0042) kind, ref, area, framework, ht_major, bucket count —
example.saved API worker (ADR-0042) kind, ref, area, framework, ht_major, bucket count —
api.request API worker route_class, outcome, reason count, duration_ms 2xx, 3xx, 4xx, 5xx; reason = the exact status of a 5xx (503, 500), empty otherwise
session.start API worker framework, ht_major, outcome count, duration_ms ready, at_capacity, container_starting, boot_timeout, budget_denied, error
session.end API worker framework, reason count, value (awake s) reason pagehide, sleep_after, teardown_failed, budget_closed
container.boot_ms API worker framework, outcome, reason duration_ms ready, window_exceeded, error; reason cold, warm
pool.gauge API worker */5 reason (live, builder) value (awake), cap —
budget.gauge API worker */5 reason (tier) value (percent of ceiling), usd —
snapshot.build API worker framework, outcome, reason, demo_id count, duration_ms, bytes ok, failed; reason inline, detached
serve.share, serve.d, serve.embed API worker outcome, demo_id count, bytes 2xx, 304, 4xx, 5xx
chat.answer API worker model, outcome count, duration_ms, usd, tokens_in, tokens_out answered, denied, error
chat.edit API worker outcome count proposed, applied, undone
theme.ai API worker model, outcome count, duration_ms, usd answered, denied, error
import.url API worker provider, outcome, reason count, duration_ms ok, refused, error
payload.boot API worker framework, outcome count ok, error
reconcile.run API worker cron outcome count, duration_ms, usd (billing total WRITTEN this run — not a delta against the estimate; D-M15 fix round) ok, skipped, error
o11y.ingest o11y worker reason, outcome count, bytes accepted, dropped, duplicate; reason = gate
o11y.drain o11y worker reason, outcome count (objects), duration_ms (the not-ready give-up error point: how long the box was not ready), bytes, value (records dropped for being too old — reject_old_samples_max_age) ok, partial, error; reason backlog, visit, reopen (this batch replayed reopened keys)
o11y.wake o11y worker reason, outcome count, duration_ms (to ready) reason backlog, visit; outcome clean, unclean
o11y.ae_degraded o11y worker reason count reason start (a wake started with the ClickHouse route off), reload (re-applying it on an already-running box failed)
o11y.backlog o11y worker cron — value (oldest age s), bytes —
o11y.alert o11y worker cron reason (rule id), outcome count fired, resolved
o11y.new_fingerprint o11y worker cron fingerprint count (one point per fingerprint the registry saw for the first time; at most 100 per tick, the rest stay in the registry; no Slack message) —

session.end's value is the awake seconds the cost ledger booked for the session: the sum of every slice recordContainerUsage booked for it, each capped at the 300 s awake window, so a hidden tab or an abandoned session is not credited with the quiet gaps between ticks or the hours before a late teardown. Sessions whose meter predates the running total fall back to the wall-clock span of their ticks plus the final slice. It is omitted (read back as 0) when the KV meter is already gone, and the tier-2 panel excludes those rows. The admin panel's kill button emits no point. o11y.ingest dropped reasons ingest_timeout and ingest_error come from the v1/logs, deploy and hooks/sentry routes when InboxWriter.ingest exceeds its 10 s deadline (INGEST_DEADLINE_MS) or rejects; the route answers 503 with Retry-After: 30. A commit whose reply is lost is counted ingest_timeout, and its redelivery only as duplicate, so that batch is never counted accepted. Cloudflare's exporter behaviour (timeout, retries and backoff, back-pressure on the source Worker, what it treats as failure) is undocumented and remains a known unknown, watched through the "Observability self" dashboard.

Two reasons in this table have no emitter and never produce data: session.end reason sleep_after (nothing observes the Sandbox SDK's idle-timeout stop) and pool.gauge reason builder (the builder pool is not metered; the */5 cron emits live only).

o11y.wake's duration_ms is wake-to-ready time — from the wake starting to the box's first successful isReady(), sourced from wake:<wakeId>.readyMs (§8) — on both the clean and unclean outcome. duration_ms = 0 means the box never became ready during that wake, not a genuinely instant boot.

The point's own timestamp is a different moment: it is written at resolution — the next backlog() call that runs resolveWakes() (§8) and finds this wake over, not when the wake itself started. Locally, with no cron firing, that can lag the actual wake by an hour or more; in production it lags by at most one */10 cron period. Two wakes resolved in the same resolveWakes() call get timestamps identical to the millisecond even though their wakes started at different times — expected, not a bug.

sandpack.compile_error counts Tier-1 compile failures from both places they occur:

  • a bundler diagnostic (show-error with no frames — the module never evaluated);
  • the parcel pre-transpile's own babel parse failure (every starter except vue-cli). That source never reaches the bundler and the last good render stays on screen, so SandpackRuntime reports it from the transpile catch itself (isTranspileFailure): at mount (a saved/shared/?payload= demo that does not parse — counted at once), and on the edit path for the newest push only.

The signal is the transpile catch, never the message shape: a runtime SyntaxError (JSON.parse, new Function) is relayed by the preview like any other throw and stays preview.runtime_error. Compile errors go through the same edit-burst collapse as preview.runtime_error (below), keyed by kind, so a burst counts at most one — from its final state. A compile failure of the burst's newest edit also replaces the run: the preview never ran that code, so what it relays for the rest of the burst (a keystroke-prefix rung still in flight, a re-render warning) is from code already typed past and is dropped. One typed broken line = one sandpack.compile_error, no preview.runtime_error. The next edit re-arms runtime reports. No error card and no Sentry capture is added for the edit-path failure; the mount-path Sentry capture (Tier1CompileError) is unchanged.

preview.runtime_error counts broken preview states, not relays. Reason console is a console.error; a console warning is not counted (§6). The preview re-runs on every keystroke, so one typed line relays a whole keystroke-prefix ladder (s is not defined, se is not defined, …, then the line's real error). The browser collapses it (apps/authoring/src/demoEventCollapse.ts) before the facade:

  • an edit that re-runs the preview (a non-quiet workspace write, a file add, delete or rename) opens or extends a burst, and discards what the previous run reported;
  • on Tier 1, an edit whose transpiled sandbox matches the running one (a closing ;, whitespace, a trailing comma) re-runs nothing, so the burst ends with the running sandbox's reports that have not been counted yet. If the bundler rejected that sandbox (a frameless show-error), its sandpack.compile_error is that result and replaces the run; a pre-transpile failure never ran, so it is not the running sandbox's;
  • on Tier 1, the bundler runs one compile at a time, so a run's reports can still arrive after the next edit has been dispatched. A new run starts at the bundler's start message for a pushed compile (onPushOutcome("rerun")), not at dispatch, and what the burst held until then came from the run it replaces and is dropped. A pre-transpile failure of the newest edit is kept, because no run of that edit will start. A run that never starts (a stalled or unreachable bundler) drops nothing, and the burst closes on its quiet window as usual;
  • 2 s (DEMO_EDIT_SETTLE_MS) after the last edit the burst closes, and the last run's reports are emitted, one per §7 fingerprint;
  • outside a burst (first load, a user interaction, a Tier-2 rebuild that reports after the burst closed) a report is emitted at once;
  • a fingerprint counts once until the next edit or preview mount, and at most 50 (DEMO_COLLAPSE_CEILING) points per page load. An async report of the previous run (a timer, a rejected promise, a failed request) that fires after the next run started can still add one point. On Tier 2, where a rebuild outlasts the 2 s window, a superseded rebuild's report can land after the burst closed and count on its own. The Sentry side is not behind this collapse; its relay budgets are unchanged.

example.saved is written by the API worker when an editor Save (PATCH /api/demos/:id with files) finishes its rebuild, because the rebuild can outlast the visitor's stay on the page. It is attributed to the example the demo came from, resolved from the row's forked_from (a bare demo id is followed up to 5 saved-demo hops, with a cycle guard): a docs fork gives kind=docs, ref = the guide, area = its first breadcrumb element and bucket from the lineage (the same values the example's example.open carried, from the taxonomy workers/api/src/docs-taxonomy.generated.ts bundles); a starter fork gives kind=starter, ref = the framework key; import and payload forks give their own kind and source. Where nothing resolves (an MCP demo, a docs path missing from the bundled taxonomy, a broken or looping lineage, a failed lineage read) it falls back to kind=saved, ref = the demo id, and empty area/bucket. framework is the demo row's and ht_major is the body's exampleHtMajor, the major the editor opened the demo at. blob1/blob2 name demos-api.

What gates the count is a successful rebuild whose request carries a valid exampleHtMajor (one of §3's ht_major values), whoever sends it. The editor sends the field only while its own telemetry gate is open (§10). The rebuild and the point are registered with ctx.waitUntil, so a client disconnect within the 30 s grace does not cancel them. Every rebuild response carries exampleSaved (whether the point was written). The API is the only emitter of example.saved; the editor never emits it.

A build that fails on the demo's own input is client input (isUserBuildError in workers/api/src/share.ts), on any route that builds inline: POST /api/demos, PATCH /api/demos/:id, POST /api/mcp/demos and PATCH /api/mcp/demos/:id. That means one of:

  • the build command exited with a code from 1 to 125;
  • the install failed with ERR_PNPM_NO_MATCHING_VERSION, ERR_PNPM_FETCH_404, ERR_PNPM_SPEC_NOT_SUPPORTED_BY_ANY_RESOLVER or ERR_PNPM_BAD_PM_VERSION (a dependency the author named).

It answers 422 {"error":"<phase> failed: <detail>","code":"build_failed","detail":<the build error, one line>}. error carries the diagnostic because MCP clients read only that field; the editor keys on code. api.request records it as 4xx, so api-5xx-rate does not count it. snapshot.build still records failed, and the snapshot-build-failed-rate alert is the backstop: it fires when one framework has more than 50 % failed builds over 30 min with at least 10 failed across at least 3 distinct demos (one author retrying one broken demo does not fire it). The stored demo is unchanged, because the build runs before anything is written.

Everything else stays 5xx: exit code 126, 127 or 128 and above (not executable, not found, killed by a signal), a result without an exit code, any other install failure (ERR_PNPM_FETCH_5xx, a reset or timed-out connection), and any other throw. So does any failure whose output names infrastructure, whatever the exit code: ENOTFOUND, ECONNRESET, ETIMEDOUT, EAI_AGAIN, heap out of memory, signal SIGKILL, worker exited, and Next's `next/font` error or a fonts.googleapis.com fetch, anywhere in the output. fetch failed and Failed to fetch count only on an error line (TypeError: fetch failed), never on a bare line, because an author's build script can print them. This last rule is a match on message text: a tool that rewords these lines moves its failure into the 422 class, and an author who throws that text as an Error reads as ours.

serve.share locally: under vite dev (what pnpm dev:full serves), React StrictMode runs the share page's load effect twice, so one /share/<id> view gives 2 points. A production build gives 1 (measured on vite preview).

payload.boot records the Theme Builder hand-off at both ends:

  • ok and error at POST /api/payload, when the link is minted;
  • error (framework=other) at the playground boot, GET /api/payload/:id, when the link cannot boot: a miss (expired or never minted), a malformed id, or a KV failure. A link that boots adds no point, so each hand-off counts one ok at most.

The server-side error point is the record of a ?payload= boot that failed. The browser shows the "expired" message for the 404 without reporting it. error.handled context=payload-boot is only the browser failing to reach the API at all (a network error), which the server never sees.

6. Browser facade and Faro

The app never calls Faro directly; it calls one facade, implemented with Faro:

interface Telemetry {
  metric(name: MetricName, values: { duration_ms?: number; value?: number; count?: number }, attrs: HotAttrs): void;
  event(name: EventName, attrs: HotAttrs & Record<string, string>): void;
  error(err: unknown, context: string, attrs?: HotAttrs): void; // handled errors
  pageLoadId(): string; // minted in memory at page load
}

Faro configuration: session tracking disabled; only the errors and web-vitals instrumentations; no user meta; the facade sets session.id = page-load id on every item; transport to same-origin /telemetry/collect; beforeSend = scrubTelemetry then the shared noise gates.

Faro's global dedupe stays on for pushError: consecutive identical exceptions (same type, message, stack, context) collapse to one item, with no time window. So error.uncaught/error.handled counts — including the dashboard panels built on them — are reports, not occurrences: a burst of identical errors counts as one. Sentry, not this pipeline, is the occurrence counter (ADR §F.3).

What the o11y worker does with each Faro item at ingest:

Faro item Analytics Engine Inbox (Loki browser tenant)
measurement whose type is a browser metric in §5 one point none
web-vitals measurement one web_vital point per vital none
exception with context.handled = "true" error.handled one log record, symbolicated at drain
other exception error.uncaught one log record, symbolicated at drain
event named example.* one point none
other event, log — one log record

Demo-runtime records. Each report that survives the collapse (§5) is also one handled Faro exception, through Telemetry.error with context demo-runtime. Its type is DemoError, DemoUnhandledRejection, DemoConsoleError, DemoNetworkError or DemoStderr. Its value is the §7 fingerprint shape (§3). It has no stack, and the §3 labels include hot.surface = demo-runtime. At ingest it becomes an error.handled point (surface demo-runtime, never feeding the new-fingerprint alert) and one Loki line, which the "Recent demo-runtime errors" panels read with {hot_surface="demo-runtime"} | hot_kind="exception". A console-warn report is neither counted nor recorded: a warning is context, not a fault (DEV-2539), and Handsontable's own load-time notices would otherwise count on every preview load. Faro's pushError dedupe applies, so an identical record in two consecutive bursts is sent once. The preview.runtime_error metric, not the line count, is the counter.

A Faro measurement or web-vitals item (this table's scope — the item kinds normalise/faro.ts handles) is AE-only; Loki holds logs, events and exceptions. This was a contract/ADR mismatch, not an implementation bug — ADR §F.1 ("Counts and latencies go to Analytics Engine; Loki holds the text") already said this; requiring a stored record for every measurement too would make measurements ~99% of the browser Loki tenant's lines and drained bytes for no reader: no dashboard panel parses a measurement's {"duration_ms":N}-shaped body, so the AE point was always the only consumer. A measurement/web-vitals item still gets a hash-only ingestItem (no record) so a retried/redelivered batch cannot double-write its Analytics Engine point — the same dedupe-only shape an example.* event already used above.

The lite-beacon path (§9, workers/o11y/src/lite.ts) follows the same rule: only a t:"err" beacon stores a record, and a t:"vital" beacon is a hash-only item that writes its web_vital point alone.

7. Fingerprint

fingerprint(context, message) = <context>:<16 hex chars of FNV-1a 64 over the normalised message>, synchronous and identical in browser and Worker; the normalisation is normalizeMonitorMessage plus stripCodeFrame. For hot.surface = demo-runtime, keystroke-ladder shapes ("<identifier> is not defined" and similar) collapse to one fingerprint per shape. Demo-runtime fingerprints never feed the new-fingerprint alert.

context is caller-chosen (typically hot.surface or a metric name) and MAY itself contain further :-separated segments, e.g. docs-example-load:fetch or npm-registry:version-exists — a call-site path, not always a single flat token. A client-supplied fingerprint (Faro's own payload.fingerprint wire field, or context["hot.fingerprint"]) is trusted only when it passes the ONE shared validator (isValidFingerprint, packages/runtime/src/telemetry/fingerprint.ts — also used by normalise/faro.ts's resolveFingerprint and normalise/otlp.ts's apiFingerprintFeed, never a second, independently drifting copy of the shape): anchor on the LAST :, followed by exactly 16 lowercase hex characters, with zero or more earlier :-separated segments in context, each drawn from [a-z][a-z0-9._-]*; the whole context half is capped at 128 characters. A value that does not match is discarded, never stored or forwarded to Slack verbatim.

8. Inbox

Normalised records are OTLP JSON log records (resourceLogs shape), one tenant per object:

inbox/<tenant>/<yyyy-mm-dd>/<hh>/<seq:012d>.ndjson.gz     # inbox bucket; tenant = browser | worker
state/wakes/<wakeId>/clean                                # Loki bucket; written by the box on a clean stop

<seq> is a counter in InboxWriter storage, incremented in the transaction that records the key. Each NDJSON line is one OTLP ResourceLogs object. The arrival time is not part of the record (it would break dedupe); it lives on the storage row. The dedupe hash is computed over the decoded, scrubbed record before timestamps are stamped.

InboxWriter storage:

Key Value
seq last issued sequence
row:<n:012d> pending records with their arrival time, ≤ 1 MB per row. <n> is zero-padded to 12 digits so native ascending key order equals arrival order (pack.ts#collectRowBatch)
key:<inbox key> written | provisional:<wakeId> | rejected:<reason> — never committed (see done:, below)
done:<inbox key> 1 — a committed key, moved OUT of key: on commit (same write that deletes key:<inbox key>)
hash:<yyyymmdd>:<sha256> first-seen epoch ms; 24 h window, checked across the current and previous UTC-day bucket
fp:<fingerprint> first-seen epoch ms (exact registry behind the new-fingerprint detection, charted through o11y.new_fingerprint, §5)
fpts:<firstSeenMs:015d>:<fingerprint> same first-seen epoch ms as its fp: twin — a time-ordered secondary index (G1 fix round, B-C1/A-I1 remainder) so the new-fingerprint alert can do a bounded start/end range read instead of listing the whole (alphabetically, not chronologically, ordered) fp: prefix every tick. Written/deleted together with its fp: twin, always
admission:<windowStartMs:015d> { fp, fpDropped, hash, hashDropped } — one row per 10-minute window: how many NEW fp: fingerprints (budget 200) and NEW hash: dedupe entries (budget 5,000) were admitted, and how many were dropped instead of stored. Browser ingest is only rate-limited per IP, so this is the global cap on a distributed flood. A dropped fingerprint's record is still stored; a dropped hash only means a redelivery inside the 24 h window may be stored twice. Rows older than 1 h are pruned; the admission-overflow alert reads the dropped counts (Slack fire/resolve plus the o11y.alert point)
fpCount / fpLapSeen the fp: registry size counter and the running count of the prune lap in progress. The counter is incremented at ingest, lowered by TTL deletes, and re-measured from fpLapSeen when a full prune lap finishes. While it exceeds 50,000, each cron tick evicts the oldest entries (via fpts:), at most 5,000 per tick. A flood can therefore push real fingerprints out, at the cost of a re-alert if one reappears
alert:<rule> { state: firing | resolved, since, lastNotified }
alertMeta:newFingerprintCursorKey / alertMeta:newFingerprintAnnouncedKeys the new-fingerprint detection's keyset cursor (an fpts: key) and a JSON array of the fpts: keys it already reported past that cursor. The cursor lags 2 minutes behind the tick, so the next tick reads recent entries again; the reported set makes sure each fingerprint becomes exactly one o11y.new_fingerprint point (F35). Both are saved only after the point write returns: a write that throws or rejects saves neither, and the next tick announces the fingerprint again (a binding that returns no result cannot report failure)
wake:<wakeId> { startedAt, reason, over: boolean, readyMs? } — over when a newer wake started or the container is not running; deleted once fully resolved (see below). readyMs (F8) is wake-to-ready time in ms, written once by InboxWriterApi.recordWakeReady(wakeId, readyMs) — called by GrafanaBox on the wake's first successful isReady(), first call wins, a no-op for an already-resolved (deleted) wake — and copied onto the resolved o11y.wake point as duration_ms (§5); absent while the box has not yet become ready
rejectedEvent:<ms:015d>:<inbox key> rejection reason (string) — a chronological audit/alert log (G1 fix round, row 19 / B-C1/A-I1 remainder), written by both a full rejection (ledger.ts#rejectKey) and a partial one (ledger.ts#recordPartialReject, see below). The rejected-inbox-key alert fires on a RECENT (last hour) count here, not on rejectedKeyCount()'s never-pruned total, so it resolves once rejections stop instead of firing forever after the first one ever seen
drainsPaused boolean (o11y spend cap). Set on every alert tick from the o11y-spend-cap result, so it follows the runtime budget override. The cron decides its backlog wake only after the alerts ran (F37). While it is set, the cron never wakes the box for the backlog, and GrafanaBox.drainStep pushes nothing, whatever woke the box. A visit wake still starts the box and serves Grafana (ADR §G)
heartbeat { lastCron, lastIngest }

Limits: records over 256 KB are dropped; requests to Loki carry at most 1 MB decompressed. Every free-text string (Faro/OTLP body, message, attribute value) is truncated to this same 256 KB before any scrub/redact regex runs over it (SCRUB_TEXT_MAX_CHARS, packages/runtime/src/telemetry/scrub.ts) — a ReDoS defense-in-depth independent of each pattern also being made linear-time. The fingerprint normaliser (normalizeMonitorMessage, §7) is bounded separately, to a much smaller 4096 chars, since its own output is always sliced to 200 chars regardless.

The pre-scrub truncation above shares the same 256 KB limit as the record-size drop, so a field that actually gets truncated still leaves the record over the drop cap — an oversize record is genuinely dropped, never truncated down to fit. Only the inbox/Loki record and its first-seen fp: entry are lost this way; the item's Analytics Engine point (e.g. error.uncaught/error.handled) is still written — dropping is a size decision at the inbox/Loki layer only, not an ingest-wide refusal.

Bounded storage (the resolve/drain/backlog paths must never scan committed history):

  • A key: entry only ever holds a live state (written, provisional:<wakeId>, or a genuine rejected:<reason>). The moment a key is confirmed clean-committed, its key:<inbox key> entry is deleted and a done:<inbox key> marker takes its place in the same write — key: therefore never grows with committed history, only with what is currently open or in flight. done: entries are pruned once their embedded date is older than the 7-day inbox-object retention (the ten-minute cron path, via InboxWriter.backlog()), using a bounded start/end range delete (done:inbox/<tenant>/ through the cutoff date), never a full-prefix scan.
  • A wake:<wakeId> entry is deleted as soon as resolveOverWakes fully resolves it (every provisional key under it moved to written or done:) — not merely flagged. The wake: prefix therefore only ever holds the (at most one) currently-active wake plus any wake whose resolution crashed mid-way, never all-time history.
  • hash: is bucketed by UTC calendar day (hash:<yyyymmdd>:<sha256>) instead of one flat set; a dedupe check reads exactly the current and previous day's buckets (the 24 h window can never span more than those two), and stale buckets (2+ days old) are pruned with a bounded range delete on the same cron path.
  • fp: keeps its flat shape (nothing reads it by date range), but is swept by a bounded, cursor-paginated TTL prune (default 90 days) on the same cron path, so it does not grow forever either.
  • POST /grafana/_o11y/reopen's window WIDTH (toMs - fromMs) is capped to the same 7-day retention (ledger.ts#reopenWindowExceedsRetention) — a wider request is refused up front. This does not require the window itself to be recent: a [from, to) pair from long ago, 7 days wide or narrower, is accepted too, it just finds nothing to reopen, since done:/hash: past the retention window are already pruned (Loki's own reject_old_samples_max_age is 7d too). The route also accepts a request whose content-type merely contains application/json as a parameter (e.g. with a charset), not only an exact match — still enough to force a CORS preflight for any cross-origin caller (the CSRF hardening this exists for).
  • A manual reopen of a committed key reads done:, moves it back to key:<inbox key> = written, and deletes the done: entry.

Additional bounded-storage invariants:

  • Pack alarm, bounded. The alarm pages row: in small chunks rather than loading every pending row into memory before packing, accumulating up to one packed object's own ~4 MB budget per round, looping until the backlog is drained or MAX_OBJECTS_PER_ALARM (25) objects have been packed this invocation. See pack.ts#collectRowBatch.
  • DO storage 128-key batch limit. Cloudflare's SQLite-backed Durable Object storage API caps get/put/delete at 128 keys/pairs per call (https://developers.cloudflare.com/durable-objects/api/storage-api/, fetched 2026-09-24: "Supports up to 128 keys at a time" / "up to 128 key-value pairs at a time"). Every multi-key call in InboxWriter chunks through storage.ts's getManyChunked/putChunked/deleteChunked — checkDuplicates, newFingerprintWrites, pruneLedger/pruneHashBuckets/pruneFingerprintRegistry, finalizeWakeResolution, markKeysProvisional, reopenWindow, commitPackedObject, and ingest's own transaction put. finalizeWakeResolution/reopenWindow run their whole put+delete sequence inside one storage.transaction() — chunking alone, without that, would let a crash between chunks leave a partial write.
  • Prune throughput. hash:/done: prune batch sizes are 5,000 rows/tick (still chunked to 128 per actual delete() call) — ADR §D's own 10× headroom projects ~220,000 worker records/day, which 500/tick × 144 ten-minute ticks/day (72,000/day) falls behind at roughly 3× today's traffic. rejected: key: entries are pruned too, past the same 7-day retention (filtered by value, since key: mixes live and rejected states chronologically — see ledger.ts#pruneLedger).
  • Drain partial-400 durability. A key with at least one chunk accepted (2xx) and at least one chunk permanently rejected (400) now stays provisional (not rejected) — its accepted content follows the normal written→provisional→committed path, so an unclean stop before Loki's local flush still triggers an automatic replay instead of being silently unrecoverable except by manual reopen. Only a key with ZERO accepted chunks stays rejected. See drain.ts#drainKey's own doc comment.
  • Drain refusals. A 429 whose body names Loki's stream limit (Maximum active stream limit exceeded) is never retried (a retry can answer 204 with the excess streams dropped) and never rejects: the table may have been filled by earlier keys, so that key is deferred and stays written, with no rejectedEvent. Its tenant is then excluded for the rest of the wake (streamLimitedTenants in the box's storage; nextWrittenKeys pages past it), so the batch fills with the other tenant's keys, and one o11y.drain.stream_limit warning line names the tenant and Loki's message. Any other 429, and a 5xx, stays transient and stops the batch. An inbox read that throws defers only that key; the rest of the batch still pushes and commits. So does a source-map read that still fails after two in-call retries (deferral: "map_fetch_error", logged as o11y.drain.error): the key stays written and nothing is pushed for it, so a replay never sends an unsymbolicated copy first. The ledger counts no attempts, so the deferral lasts only while the key's inbox hour is under 6 hours old (MAP_RETRY_MAX_AGE_MS); after that the key pushes with its frames as they are, so a replay is byte-stable only once the maps are final. A batch of only deferred keys ends the wake's drain once no un-excluded tenant has keys left. Either deferral (inbox read or source-map read) also excludes that key for the rest of the wake (deferredKeys in the box's storage, reset by a new wake; nextWrittenKeys skips it), so a deferred key is not re-read on every step and cannot starve the batch: the keys behind it still drain. When only deferred keys remain and a Grafana visitor keeps the box up, the drain clears that set and rechecks after 60 s, like the spend-cap pause; with no visitor the box stops as for an empty backlog.
  • Symbolication read caps. One inbox object reads at most 32 distinct maps (MAX_MAP_KEYS_PER_CALL, first-seen order), one body adds at most 8 of them (MAX_NEW_MAP_KEYS_PER_BODY), and at most 128 frames are looked up per body (MAX_FRAMES_PER_BODY), so a drain step of 10 objects stays at a few hundred of the Workers limit of 10,000 subrequests. Frames past a cap stay byte-for-byte and are reported as o11y.symbolicate.skip with reason over_cap, plus one aggregate over_cap line with the call's capped frames and keys that MAX_SKIP_REPORTS never suppresses. Before any read, each distinct service.version that has a resolvable frame is listed once (sourcemaps/<version>/, paged through the R2 cursor); only a key the listing shows is admitted, so a forged frame path costs no read and is reported as no_map. service.version is client-supplied, so a call lists at most 8 versions (MAX_LISTED_VERSIONS_PER_CALL); frames of the rest stay as they were and are reported per version prefix as over_version_cap. A list that throws (or runs past 5 pages) is reported as list_error; while the key is under 6 hours old (MAP_RETRY_MAX_AGE_MS) the object is deferred like a failing map read, so a replay never admits by the caps what the first pass admitted by the listing, and past that age the version falls back to the caps above. The per-call version cap is a stated residual risk, not closed: at least 8 forged requests with distinct service.version packed ahead of a real one can still push that real version to over_version_cap.

9. Lite beacon payload

POST /telemetry/lite, a JSON body, ≤ 2 KB, sent with navigator.sendBeacon(url, json) — a plain string, not a Blob, so the browser sends its own default Content-Type: text/plain;charset=UTF-8, never application/json (the route itself does not check or require a content type either way, so this is a fact about what ships on the wire, not a gate). The body itself is still JSON:

{"v":1,"t":"err","s":"embed","demo":"r-react-18-0-0","ht":"18","fw":"react","n":"TypeError","m":"<normalized, ≤500 chars>","st":"<stack, ≤2000 chars>","val":null,"dev":"desktop","ts":1695463200000,"id":"a1b2c3d4"}

t = err | vital; s = embed | d; for vital, n is LCP | INP | CLS | TTFB and val carries the value. No page path: docs pages send no referrer. Vitals are sampled at 10 % per page view, decided once per page; errors are sent up to the monitor.ts event ceiling. The o11y worker converts beacons with beaconToRecord, a sibling of the Faro converter faroItemToRecord (packages/runtime/src/telemetry/convert.ts) that shares its clamp of ts to the receive time ± 5 minutes and its resource-attribute handling but builds the record from the beacon's own fields (workers/o11y/src/lite.ts).

The dedupe hash covers the whole converted record (body with message and stack, attributes) plus the raw ts and, when present, id: a per-beacon random value, used only in the dedupe hash.

10. Local mode

Piece Local stand-in
RUNNER_EVENTS ClickHouse at http://localhost:8123, table runner_events with the §4 columns plus timestamp and _sample_interval (always 1), DDL in containers/o11y/local/clickhouse-init.sql
Loki S3 Miniflare's local S3 endpoint for R2, or MinIO from containers/o11y/compose.yml
Grafana session (K1) DEV_ADMIN in workers/o11y/.dev.vars; a real broker login also works locally (LOGIN_BROKER_URL defaults to the production broker)
Cloudflare OTLP export fixtures in pipeline/fixtures/otlp/ (scrubbed sandbox-probe captures plus hand-built edge cases), replayed by scripts/o11y-replay-fixtures.mjs
Slack scripts/o11y-slack-capture.mjs, a local HTTP capture server started by pnpm dev:full (NOT by pnpm o11y:dev, which doesn't start it — point SLACK_WEBHOOK_URL at your own instance if you need one from the standalone o11y-only command); prints and keeps the last 50 posts, GET /_captured to inspect

deployment.environment.name = local; the production o11y worker drops local data.

Local telemetry gate in the browser. Faro runs on the local path only when the build was made with VITE_TELEMETRY_LOCAL=1 and the host is localhost or 127.0.0.1. It checks neither import.meta.env.DEV nor navigator.webdriver, because Playwright serves a production vite preview under automation. The production gate (resolveReporting) is unchanged and stays closed under automation. A production build never sets the flag, and the post-build leak check fails if the local path survives into it.

11. Sentry scope

SENTRY_SCOPE / VITE_SENTRY_SCOPE = full (default): handled diagnostic reports go to Sentry and the new stack. uncaught: they go only to the new stack; Sentry keeps errors that escape a handler (browser onerror, unhandledrejection, Sentry.ErrorBoundary; Worker fetch catch-all, DO alarms, cron, snapshot-job failures) and the budget-alert captureMessage. The launch plan flips to uncaught after the pipeline is seen working in production.