Give evals real MCP tools, and fix the clustering threshold expectation - #86
Open
mattpodwysocki wants to merge 1 commit into
Open
mattpodwysocki wants to merge 1 commit into
mattpodwysocki wants to merge 1 commit into
Conversation
mattpodwysocki
force-pushed
the
fix-caching-tos-snippets
branch
from
October 1, 2026 15:09
35c5de4 to
021e9f0
Compare
mapbox-location-grounding is about composing live tool calls into a cited answer, but the runner only ever sent a system prompt and a question — no tools. The task was impossible, so the scores recorded which model was most willing to answer anyway. Asked to describe a neighbourhood from coordinates, Opus declined to invent POIs and walk times, said why, and showed the exact ground_location_tool call it would make: 1/12. Sonnet answered confidently from training data: 84%. The eval rewarded fabrication, the one behaviour this skill exists to prevent. An eval set can now opt in by name: "mcp_servers": ["mapbox"] The runner resolves those against a known-server table, attaches them with the mcp-client beta header, and fails loudly when the required token is missing rather than quietly grading a tool-less response against tool-use expectations. Judging had to change too. MCP tool use happens server-side and often leaves no trace in the prose, so "calls ground_location_tool with the coordinates" was unscoreable from text alone. Responses are flattened into a transcript that renders tool calls and truncated tool results alongside the text. Skills without MCP emit text blocks only, so their transcripts are unchanged. mapbox-location-grounding, skill surface: haiku-4.5 60.2% -> 81.5% sonnet-5 84.3% -> 91.7% opus-5 50.9% -> 91.7% That one eval set was inverting the suite's model-scaling signal. Full suite, skill surface, all 20 skills: before haiku 90.8% -> sonnet 95.9% -> opus 93.9% (non-monotonic) after haiku 92.7% -> sonnet 96.6% -> opus 97.5% (monotonic) Separately, mapbox-web-performance-patterns #2 asserted that 75,000 points do not require clustering because "clustering is for 100,000+", contradicting SKILL.md, which recommends clustering from 10,000. It scored 80% on both surfaces, so it was never AGENTS.md drift — the expectation was wrong. Settled on 10,000: the expectation now looks for clustering enabled on the GeoJSON source, and the first expectation no longer frames clustering as an alternative to layers, since it is a property of the source that still renders through one.
mattpodwysocki
force-pushed
the
wire-evals-to-mcp
branch
from
October 1, 2026 15:23
0c68438 to
da20907
Compare
mattpodwysocki
changed the base branch from
fix-caching-tos-snippets
to
eval-runner-improvements
October 1, 2026 15:23
This was referenced Oct 1, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Two evals were measuring the wrong thing.
mapbox-location-groundingwas grading willingness to fabricateThe skill is about composing live tool calls into a cited answer. The runner only ever sent a system prompt and a question — no tools. The task was impossible, so the scores recorded which model was most willing to answer anyway.
Asked to describe the neighbourhood at a coordinate, Opus refused to invent POIs and walk times, said why, and showed the exact
ground_location_toolcall it would make. That scored 1/12. Sonnet produced a confident answer from training data and scored 84%. The eval rewarded fabrication — the one behaviour this skill exists to prevent.An eval set can now opt in by name:
{ "skill_name": "mapbox-location-grounding", "mcp_servers": ["mapbox"] }The runner resolves those against a known-server table, attaches them with the
mcp-clientbeta header, and fails loudly when the token is missing rather than quietly grading a tool-less response against tool-use expectations. Onlymapbox(https://mcp.mapbox.com/mcp,MAPBOX_ACCESS_TOKEN) is wired up.Judging had to change too: MCP tool use happens server-side and often leaves no trace in the prose, so "calls
ground_location_toolwith the coordinates" was unscoreable from text alone. Responses are now flattened into a transcript rendering tool calls and truncated tool results alongside the text. Skills without MCP emit text blocks only, so their transcripts are unchanged.That one eval set was inverting the suite's scaling signal
A usable eval suite should score higher with stronger models. It did not, and this was why. Full suite, skill surface, all 20 skills:
The suite now passes the model-scaling check with every skill included, instead of only when
location-groundingwas excluded.web-performance#2: the 10,000 thresholdThe expectation asserted that 75,000 points do not require clustering because "clustering is for 100,000+", contradicting
SKILL.md:175, which recommends clustering from 10,000. It scored 80% on both surfaces, so it was neverAGENTS.mddrift — the expectation was wrong.Settled on 10,000. The expectation now looks for clustering enabled on the GeoJSON source, and the first expectation no longer frames clustering as an alternative to layers, since clustering is a property of the source that still renders through one.
That eval moves 80% → 95%; the skill 91.7% → 94.2%.
Follow-up
evals/baseline.jsonpredates this change, so itslocation-groundingrows reflect tool-less runs. Worth regenerating once a repeat count is settled.Test plan
npm run checkpassesmapbox-location-groundingrun at all three model tiers with MCP attachedmapbox-web-performance-patternsre-run at--repeats=5MAPBOX_ACCESS_TOKENfails the run rather than silently degrading it🤖 Generated with Claude Code