Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 2 additions & 4 deletions runner/docs/observability-contract.md
Original file line number Diff line number Diff line change
Expand Up @@ -444,10 +444,8 @@ hash-only `ingestItem` (no `record`) so a retried/redelivered batch cannot doubl
its Analytics Engine point — the same dedupe-only shape an `example.*` event already
used above.

**Scope note, not yet fixed**: §9's lite-beacon path (`workers/o11y/src/lite.ts`,
`POST /telemetry/lite`) is a separate converter and still stores a vital beacon's record
today (the triage's own `LCP=172` inbox-record example) — this ruling was not extended
there. A future consistency pass may want to.
The lite-beacon path (§9, `workers/o11y/src/lite.ts`) follows the same rule: only a `t:"err"` beacon
stores a record, and a `t:"vital"` beacon is a hash-only item that writes its `web_vital` point alone.

## 7. Fingerprint

Expand Down
28 changes: 23 additions & 5 deletions runner/docs/run-and-deploy.md
Original file line number Diff line number Diff line change
Expand Up @@ -273,10 +273,11 @@ panel carries no message text. `reportDemoEvent` files the preview relay as
a `preview.runtime_error` count metric (a Faro measurement, `toFacade` →
`telemetry.metric` → `pushMeasurement`) — per the ruling above (§6),
every measurement is Analytics-Engine-only, so this produces no Loki line
at all any more, not even a count-only one. The one demo-runtime Loki line
that still exists is the Tier-1 compile-failure branch, whose message is
always the constant `"Tier-1 compile failed"` — the real diagnostic detail
goes to `extra` there, and for everything else lives in Sentry, when
at all any more, not even a count-only one. Each collapsed demo-runtime
report (one per fingerprint per edit burst) is also filed as a handled Faro exception
(`emitCollapsedDemoEvent` in `apps/authoring/src/sentry.ts`), which is the Loki line, and its message
is the fingerprint shape, never the relayed text; the Tier-1 compile-failure branch's message is
the constant `"Tier-1 compile failed"`, with the diagnostic detail in `extra`. The full text lives in Sentry, when
`MONITOR_DEMOS`/`VITE_MONITOR_DEMOS` is on.) `/admin`'s
header also has an **Open Grafana** link
(otherwise nothing in the app points at it): `href={GRAFANA_URL}` in
Expand All @@ -299,14 +300,31 @@ Viewer save a change back to a provisioned dashboard or datasource — those
stay read-only, and Grafana's state is disposable anyway (a fresh DB on
every wake).

**What the `browser` tenant holds.** Exceptions, handled or uncaught, including the
demo-runtime relays above; Faro logs; Faro events not named `example.*` (in the authoring app,
`sentry.event` and `versions_fetch_unreachable`); and lite-beacon `t:"err"` records from `/d` and
`/embed`. Everything else goes to Analytics Engine only and never reaches Loki: every measurement
(web-vitals, `preview.ready_ms`, `sandpack.compile_ms`, `preview.runtime_error`), every `example.*`
event, and lite `t:"vital"` beacons. So an open, healthy demo produces no Loki lines until
something throws. To force a probe line, add `throw new Error("o11y probe")` to the demo code in the
editor (a throw from the DevTools console skips the demo-runtime relay), wait 60 s for the pack,
let the box stop after its 15 min idle timeout or wait for the next wake, then query
`{service_name="demos-authoring"}` over the last 24 h in Explore.

**Explore on the Analytics Engine datasource.** The ClickHouse plugin's default Format As is
Time series, which drops string columns and shows the time as NaN for a query that is not a
time series. Set Format As to Table for ad-hoc queries in Explore.

**Logs are only as fresh as the last wake.** The box drains its packed
objects into Loki once, right after it wakes, and nothing re-arms that
drain while it stays awake (ADR-0041 §B.3's out-of-order window assumes the
drain replays into an empty ingester, which only holds true at wake start).
So any ingest that arrives *while* the box is already up sits undrained and
invisible in Grafana/Explore until the *next* wake. There is no staleness
indicator on the dashboards for this; treat Logs/Explore as "as of the last
wake started", not live — a design change may follow.
wake started", not live — a design change may follow. A new browser error can therefore take
until the next wake to appear; the cron forces one when the oldest inbox key is over 1 h old or the
backlog exceeds 64 MB (`workers/o11y/src/index.ts`).

**Bot traffic is filtered locally too.** The o11y worker's bot gate drops
any request whose user agent matches `HeadlessChrome` — including local
Expand Down
Loading