Skip to content

Give evals real MCP tools, and fix the clustering threshold expectation - #86

Open
mattpodwysocki wants to merge 1 commit into
eval-runner-improvementsfrom
wire-evals-to-mcp
Open

mattpodwysocki wants to merge 1 commit into
eval-runner-improvementsfrom
wire-evals-to-mcp

Conversation

@mattpodwysocki

Copy link
Copy Markdown
Contributor

Stacked on #80 — base is fix-caching-tos-snippets, since this builds on the --surface and --repeats runner work there. Retarget to main once #80 merges.

Two evals were measuring the wrong thing.

mapbox-location-grounding was grading willingness to fabricate

The skill is about composing live tool calls into a cited answer. The runner only ever sent a system prompt and a question — no tools. The task was impossible, so the scores recorded which model was most willing to answer anyway.

Asked to describe the neighbourhood at a coordinate, Opus refused to invent POIs and walk times, said why, and showed the exact ground_location_tool call it would make. That scored 1/12. Sonnet produced a confident answer from training data and scored 84%. The eval rewarded fabrication — the one behaviour this skill exists to prevent.

An eval set can now opt in by name:

{
  "skill_name": "mapbox-location-grounding",
  "mcp_servers": ["mapbox"]
}

The runner resolves those against a known-server table, attaches them with the mcp-client beta header, and fails loudly when the token is missing rather than quietly grading a tool-less response against tool-use expectations. Only mapbox (https://mcp.mapbox.com/mcp, MAPBOX_ACCESS_TOKEN) is wired up.

Judging had to change too: MCP tool use happens server-side and often leaves no trace in the prose, so "calls ground_location_tool with the coordinates" was unscoreable from text alone. Responses are now flattened into a transcript rendering tool calls and truncated tool results alongside the text. Skills without MCP emit text blocks only, so their transcripts are unchanged.

model before after
haiku-4.5 60.2% 81.5%
sonnet-5 84.3% 91.7%
opus-5 50.9% 91.7%

That one eval set was inverting the suite's scaling signal

A usable eval suite should score higher with stronger models. It did not, and this was why. Full suite, skill surface, all 20 skills:

haiku-4.5 sonnet-5 opus-5
before 90.8% 95.9% 93.9% ✗ non-monotonic
after 92.7% 96.6% 97.5% ✓ monotonic

The suite now passes the model-scaling check with every skill included, instead of only when location-grounding was excluded.

web-performance #2: the 10,000 threshold

The expectation asserted that 75,000 points do not require clustering because "clustering is for 100,000+", contradicting SKILL.md:175, which recommends clustering from 10,000. It scored 80% on both surfaces, so it was never AGENTS.md drift — the expectation was wrong.

Settled on 10,000. The expectation now looks for clustering enabled on the GeoJSON source, and the first expectation no longer frames clustering as an alternative to layers, since clustering is a property of the source that still renders through one.

That eval moves 80% → 95%; the skill 91.7% → 94.2%.

Follow-up

evals/baseline.json predates this change, so its location-grounding rows reflect tool-less runs. Worth regenerating once a repeat count is settled.

Test plan

  • npm run check passes
  • mapbox-location-grounding run at all three model tiers with MCP attached
  • mapbox-web-performance-patterns re-run at --repeats=5
  • Verified a missing MAPBOX_ACCESS_TOKEN fails the run rather than silently degrading it

🤖 Generated with Claude Code

@mattpodwysocki
mattpodwysocki requested review from a team as code owners October 1, 2026 14:15
@mattpodwysocki
mattpodwysocki requested review from underoot and removed request for a team October 1, 2026 14:15
@mattpodwysocki
mattpodwysocki force-pushed the fix-caching-tos-snippets branch from 35c5de4 to 021e9f0 Compare October 1, 2026 15:09
mapbox-location-grounding is about composing live tool calls into a cited
answer, but the runner only ever sent a system prompt and a question — no
tools. The task was impossible, so the scores recorded which model was
most willing to answer anyway. Asked to describe a neighbourhood from
coordinates, Opus declined to invent POIs and walk times, said why, and
showed the exact ground_location_tool call it would make: 1/12. Sonnet
answered confidently from training data: 84%. The eval rewarded
fabrication, the one behaviour this skill exists to prevent.

An eval set can now opt in by name:

  "mcp_servers": ["mapbox"]

The runner resolves those against a known-server table, attaches them with
the mcp-client beta header, and fails loudly when the required token is
missing rather than quietly grading a tool-less response against tool-use
expectations.

Judging had to change too. MCP tool use happens server-side and often
leaves no trace in the prose, so "calls ground_location_tool with the
coordinates" was unscoreable from text alone. Responses are flattened into
a transcript that renders tool calls and truncated tool results alongside
the text. Skills without MCP emit text blocks only, so their transcripts
are unchanged.

mapbox-location-grounding, skill surface:

  haiku-4.5   60.2% -> 81.5%
  sonnet-5    84.3% -> 91.7%
  opus-5      50.9% -> 91.7%

That one eval set was inverting the suite's model-scaling signal. Full
suite, skill surface, all 20 skills:

  before   haiku 90.8%  ->  sonnet 95.9%  ->  opus 93.9%   (non-monotonic)
  after    haiku 92.7%  ->  sonnet 96.6%  ->  opus 97.5%   (monotonic)

Separately, mapbox-web-performance-patterns #2 asserted that 75,000 points
do not require clustering because "clustering is for 100,000+",
contradicting SKILL.md, which recommends clustering from 10,000. It scored
80% on both surfaces, so it was never AGENTS.md drift — the expectation was
wrong. Settled on 10,000: the expectation now looks for clustering enabled
on the GeoJSON source, and the first expectation no longer frames
clustering as an alternative to layers, since it is a property of the
source that still renders through one.
@mattpodwysocki
mattpodwysocki changed the base branch from fix-caching-tos-snippets to eval-runner-improvements October 1, 2026 15:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant