You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit d09c2c6
Browse filesBrowse the repository at this point in the historyBrowse files
Document latency limits for voice simulations (#1288)
## Summary
Voice simulations can now fail when the assistant responds too slowly. The feature shipped in VapiAI/vapi#21959 (API and worker), VapiAI/vapi#21962 (run results) and VapiAI/vapi#21963 (editor), tracked in TEST-140. The API reference already shows the new `latencyExpectations` and `latencyEvaluations` fields through the nightly spec update; this PR covers the guides.
- **Simulations advanced**: new **Set latency limits** section and card:
- Dashboard and cURL setup;
- the `turn`, `model` and `voice` metrics and the four aggregations, including that `p95` equals `max` under 20 turns;
- required versus optional limits;
- when limits are skipped (chat mode, GPT-Live assistants under test) and that a voice run with no measured turns fails;
- reading results in the run details and through `results.latencyEvaluations`;
- two troubleshooting rows.
- **Simulations overview**: the scenario definition, the "how it works" paragraph and the Evals comparison table now mention latency limits next to structured outputs.
- **Manage simulations**: run results mention latency results and `latencyEvaluations`, "Design evaluations" points to latency limits, and the voice-or-chat table gains a "Gate on response latency" row.
- **GPT-Live testing**: notes that latency limits are skipped when the assistant under test uses GPT-Live.
Every behavior described was checked against the merged code: field names, the default of a required median turn latency of 1,200 ms, the 1 to 60,000 ms range, the 20-limit maximum and the skip rules.
## Not included
A "What's new" changelog entry. The weekly changelog files look curated by one author (e.g. #1280), so I left the entry to the next weekly post. Happy to add it here instead.
## Testing
`fern check` with the pinned CLI 5.112.0 (`node scripts/fern/run.cjs check`): 0 errors. The 14 warnings are pre-existing discriminator warnings in the API spec. The preview link from CI is the place to check rendering.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Copy file name to clipboardExpand all lines: fern/gpt-live/testing.mdx
+2Lines changed: 2 additions & 0 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -62,6 +62,8 @@ To run one:
62
62
63
63
Unmocked tools run for real. Use [tool mocks](/observability/simulations-advanced#mock-tool-responses), or point tools at a test service, before running scenarios that change data. A passing evaluation doesn't cover everything about the conversation, so listen to some recordings too.
64
64
65
+
[Latency limits](/observability/simulations-advanced#set-latency-limits) are skipped when the assistant under test uses GPT-Live, because its per-turn latency isn't measured yet. Measure GPT-Live latency with the call's timestamps, as described in [Latency](#latency).
66
+
65
67
[Evals](/observability/evals-quickstart) check supported text-model decisions, such as which tool is called, using mock conversations. They don't test GPT-Live's listening, speech, or turn-taking. Use Voice Simulations to test the spoken conversation.
Copy file name to clipboardExpand all lines: fern/observability/simulations-advanced.mdx
+78-3Lines changed: 78 additions & 3 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -1,11 +1,11 @@
1
1
---
2
2
title: Simulations advanced
3
-
subtitle: Mock tools, send lifecycle webhooks, and reuse structured outputs in simulations
4
-
description: "Configure advanced simulations with variables, tool mocks, lifecycle webhooks, and reusable structured outputs for consistent testing."
3
+
subtitle: Mock tools, send lifecycle webhooks, set latency limits, and reuse structured outputs in simulations
4
+
description: "Configure advanced simulations with variables, tool mocks, lifecycle webhooks, latency limits, and reusable structured outputs for consistent testing."
5
5
slug: observability/simulations-advanced
6
6
---
7
7
8
-
Advanced simulation options let you set variable values, mock tool responses, trigger lifecycle webhooks, and reuse structured outputs. Use them after you complete the [**Simulations quickstart**](/observability/simulations-quickstart) and need more control over behavior or test conditions.
8
+
Advanced simulation options let you set variable values, mock tool responses, trigger lifecycle webhooks, set latency limits, and reuse structured outputs. Use them after you complete the [**Simulations quickstart**](/observability/simulations-quickstart) and need more control over behavior or test conditions.
9
9
10
10
## How it works
11
11
@@ -29,6 +29,9 @@ Open a suite and select **Edit**, then **Next**. The review step contains the **
A latency limit fails a voice simulation when the [**assistant**](/assistants) or [**squad**](/squads) under test responds too slowly. Vapi measures latency on each turn of the call, combines the turns into one number, and passes the limit when that number is at or below the limit you set.
248
+
249
+
Latency limits work alongside structured-output evaluations. A scenario still needs at least one evaluation, and a required limit that fails marks the simulation as failed even when every evaluation passes.
250
+
251
+
<Tabs>
252
+
<Tabtitle="Dashboard">
253
+
254
+
On the **Success criteria** tab, under **Latency expectations**, select **Add latency expectation**. Configure each limit with four fields:
255
+
256
+
-**Aggregation**: How the call's turns are combined: **Median**, **Mean**, **P95**, or **Max**.
257
+
-**Metric**: **Turn latency**, **Model latency**, or **Voice latency**.
258
+
-**Limit**: The maximum in milliseconds, as a whole number from 1 to 60,000.
259
+
-**Required**: Turn off to make the limit informational. An optional limit is reported but never fails the simulation.
260
+
261
+
A new limit starts as a required median turn latency of 1,200 ms. Each simulation can have up to 20 limits.
262
+
263
+
</Tab>
264
+
<Tabtitle="cURL">
265
+
266
+
Latency limits are the `latencyExpectations` property of the scenario. Add them when you create a scenario, or patch an existing one:
`required` defaults to `true`. Sending `latencyExpectations` replaces the scenario's existing limits, and an empty array removes them.
281
+
282
+
</Tab>
283
+
</Tabs>
284
+
285
+
### Metrics and aggregations
286
+
287
+
| Metric | What it measures |
288
+
| --- | --- |
289
+
|`turn`| Total time from the end of the AI tester's speech to the start of the assistant's reply |
290
+
|`model`| Time for the model to return its first token |
291
+
|`voice`| Time for the voice provider to return the first audio |
292
+
293
+
The aggregation reduces the call's turns to one value: `mean`, `median`, `p95`, or `max`. `p95` uses the nearest-rank method, so on a call with fewer than 20 turns it equals `max`. Measured values are rounded to the nearest millisecond before they are compared with the limit.
294
+
295
+
<Tip>
296
+
Gate on median turn latency. A short simulation has only a few turns, so `p95` and `max` can fail on a single slow turn. To watch the slowest turns without failing runs, add an optional `p95` or `max` limit next to a required `median` limit on the same metric.
297
+
</Tip>
298
+
299
+
### When a limit is skipped
300
+
301
+
| Situation | Result |
302
+
| --- | --- |
303
+
| The run uses chat mode | Skipped. Chat runs have no audio, so there is no latency to measure. |
304
+
| The assistant under test uses [**GPT-Live**](/gpt-live/overview)| Skipped. Per-turn latency isn't measured for GPT-Live assistants yet. A GPT-Live AI tester talking to a classic assistant is still measured. |
305
+
| A voice run where no turn measured the metric | The limit fails. A latency that couldn't be measured never counts as within the limit. |
306
+
307
+
Skipped limits don't affect the result.
308
+
309
+
### Review latency results
310
+
311
+
Open a run and select a simulation. The **Latency** section shows average turn latency when at least one turn was measured. Each limit shows its threshold and result. When at least one turn measured that limit's metric, the result also shows the measured value and number of turns. Optional and skipped limits are labeled. When a simulation fails, the count in the results list, such as `1 of 3 evaluations failed`, includes its latency limits.
312
+
313
+
Through the API, each run item's `results.latencyEvaluations` holds one entry per limit with `actualMs`, `thresholdMs`, `sampleCount`, `passed`, `required`, and, for skipped limits, `isSkipped` and `skipReason`. `results.passed` accounts for required latency limits.
314
+
242
315
## Set variable values for the assistant or squad
243
316
244
317
Variables provide values for the [**dynamic variables**](/assistants/dynamic-variables) used by the [**assistant**](/assistants) or [**squad**](/squads) during the simulation. On the **Variables** tab, add **Name** and **Value** pairs. Each name must match a `{{variable}}` placeholder in the assistant's prompt, and the value is substituted during the run.
@@ -255,6 +328,8 @@ Use variables to test specific inputs, such as a customer name or account tier,
255
328
| There is no audio or recording | Check whether the run used chat mode. Use voice mode to test or record audio. |
256
329
| Start or end webhooks do not trigger | Confirm the corresponding hook and its URL are configured on the scenario. |
257
330
| A tool mock does not apply | Confirm that `toolName` matches the configured tool name exactly and that `enabled` is `true`. |
331
+
| A latency limit fails without a measured value | No turn in the call measured that metric. Confirm the run used voice mode and that the conversation completed at least one exchange. |
332
+
| Latency limits show as skipped | Check whether the run used chat mode or the assistant under test uses GPT-Live. Both skip latency limits. |
Open **Simulations**, then select **Runs** to see every run with its suite, iterations, [**assistant**](/assistants) or [**squad**](/squads), run date, and overall result (`Passed` or a failed count such as `1/1 failed`). Filter by **time range**, **status**, or **assistant or squad** to find a run.
75
75
76
-
Open a run to see whether each simulation passed or failed, its evaluations, and its transcript. Open an active run to watch the transcript and evaluations update. If a run or iteration can't complete, it shows a **Failure reason** describing what went wrong, such as no payment method on file. See the [**Simulations quickstart**](/observability/simulations-quickstart) for the full walkthrough.
76
+
Open a run to see whether each simulation passed or failed, its evaluations, its latency results for voice runs, and its transcript. Open an active run to watch the transcript and evaluations update. If a run or iteration can't complete, it shows a **Failure reason** describing what went wrong, such as no payment method on file. See the [**Simulations quickstart**](/observability/simulations-quickstart) for the full walkthrough.
77
77
78
78
</Tab>
79
79
<Tabtitle="cURL">
@@ -90,7 +90,7 @@ curl -X GET "https://api.vapi.ai/eval/simulation/run/<run-id>" \
90
90
-H "Authorization: Bearer $VAPI_API_KEY"
91
91
```
92
92
93
-
The run includes `status` and `itemCounts` (`total`, `passed`, `failed`, and so on). Fetch the result for each simulation iteration with `GET /eval/simulation/run/<run-id>/item`.
93
+
The run includes `status` and `itemCounts` (`total`, `passed`, `failed`, and so on). Fetch the result for each simulation iteration with `GET /eval/simulation/run/<run-id>/item`. Each item's `results` holds its `evaluations` and, when the scenario sets latency limits, its `latencyEvaluations`. See [**Review latency results**](/observability/simulations-advanced#review-latency-results).
94
94
95
95
</Tab>
96
96
</Tabs>
@@ -144,12 +144,15 @@ Test realistic variations such as an ambiguous request, an impatient customer, a
144
144
145
145
Keep each evaluation focused on one observable outcome. Use a descriptive name, choose a Boolean or numeric value that can be measured consistently, and avoid combining several independent requirements into one evaluation.
146
146
147
+
To catch slow responses, add a [**latency limit**](/observability/simulations-advanced#set-latency-limits) rather than an evaluation, and run the suite in voice mode. Median turn latency is the most stable limit to gate a release on.
148
+
147
149
### Choose voice or chat mode
148
150
149
151
| Testing goal | Mode | Why |
150
152
| --- | --- | --- |
151
153
| Iterate on prompts, tools, and conversation logic | Chat (`vapi.webchat`) | Runs without audio processing, so it is faster and costs less. |
152
154
| Investigate speech recognition, voice output, or interruptions | Voice (`vapi.websocket`) | Exercises synthetic audio. Review the recording, not just the pass/fail label. |
155
+
| Gate on response latency | Voice (`vapi.websocket`) | Latency limits are skipped in chat mode. |
153
156
| Receive call-specific webhook data | Voice (`vapi.websocket`) | Start and end webhooks fire in both modes, but chat payloads omit call-specific fields. |
154
157
| Check representative voice journeys before launch | Voice (`vapi.websocket`) | Adds voice coverage. Also make controlled calls through the real phone path and sandbox integrations. |
Copy file name to clipboardExpand all lines: fern/observability/simulations-overview.mdx
+3-3Lines changed: 3 additions & 3 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -13,7 +13,7 @@ Simulations are automated tests that run an AI tester through a real conversatio
13
13
| -- | -- |
14
14
|**Simulation suite**| A simulation suite groups one or more simulations that you can run against one or more [**assistants**](/assistants) or [**squads**](/squads). |
15
15
|**Simulation**| A simulation pairs a scenario with a personality to test one situation. |
16
-
|**Scenario**| A scenario defines the AI tester's intent and the success criteria that determine the result. |
16
+
|**Scenario**| A scenario defines the AI tester's intent and the success criteria that determine the result: structured-output evaluations and, for voice runs, any latency limits. |
17
17
|**Personality**| A personality defines how the AI tester behaves, including its model, transcriber, and voice. |
18
18
|**AI tester**| An AI tester drives the simulated conversation according to a scenario and personality. |
19
19
@@ -54,7 +54,7 @@ directly instead of synthesized speech and transcription.
54
54
Use **chat mode** for rapid iteration during development, then switch to **voice mode** for final validation.
55
55
</Tip>
56
56
57
-
When you run the suite, Vapi runs each simulation for the number of iterations you select. During each iteration, the AI tester holds a live conversation with the assistant or squad. Vapi then evaluates each conversation with [**structured outputs**](/assistants/structured-outputs-quickstart) and reports which criteria passed or failed.
57
+
When you run the suite, Vapi runs each simulation for the number of iterations you select. During each iteration, the AI tester holds a live conversation with the assistant or squad. Vapi then evaluates each conversation with [**structured outputs**](/assistants/structured-outputs-quickstart), checks any [**latency limits**](/observability/simulations-advanced#set-latency-limits) on voice runs, and reports which criteria passed or failed.
58
58
59
59
## When to use Simulations
60
60
@@ -82,7 +82,7 @@ Simulations test live interactions. An AI tester with a defined personality and
82
82
|**What you provide**| A scripted conversation with expected responses | An AI tester personality and an intent |
83
83
|**How it runs**| Fixed-context, turn-by-turn checks | Dynamic; the AI tester improvises the conversation |
84
84
|**Transport**| Chat (mock conversations) | Voice, or text-only |
85
-
|**Evaluation**| Exact match, regex, AI judge, tool-call checks | Structured outputs on the conversation outcome |
85
+
|**Evaluation**| Exact match, regex, AI judge, tool-call checks | Structured outputs on the conversation outcome, plus latency limits on voice runs|
86
86
|**Reach for it when**| You need focused checks at a known point | You need realistic end-to-end or voice behavior |
87
87
88
88
Choose Evals when you want to lock down a specific response or verify a tool call's arguments with fast, rerunnable checks.
0 commit comments