Skip to content

Remove API response caching from skills per standard terms - #80

Open
mattpodwysocki wants to merge 2 commits into
mainfrom
fix-caching-tos-snippets
Open

mattpodwysocki wants to merge 2 commits into
mainfrom
fix-caching-tos-snippets

Conversation

@mattpodwysocki

@mattpodwysocki mattpodwysocki commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Why

Several skills shipped code samples that cache Mapbox API responses — Directions routes held for five minutes, all MCP tool results held for an hour, geocoding results held in an LRU cache. Caching or storing API responses is not permitted under Mapbox's standard self-serve (pay-as-you-go) terms, and caching exceptions are not granted on pay-as-you-go accounts.

So the pattern itself was wrong, not just missing a caveat. Adding a warning on top of code that tells people to cache Directions responses would not have fixed it. These sections are replaced with request-reduction patterns that never retain a response.

The two most visible cases

File Was Now
mapbox-navigation-patterns/references/best-practices.md "Route Caching" — 5-minute TTL Map of Directions responses "Reducing Directions API Calls" — in-flight request deduplication (entry dropped as soon as the request settles), debouncing, overview=simplified, Matrix API for many-to-many travel times
mapbox-mcp-runtime-patterns/references/production.md "Caching Strategy" — 1-hour TTL on all tool results, Infinity for offline tools "Request Deduplication and Offline Memoization" — memoize only the 9 offline Turf.js tools (local compute, no Mapbox content); API-backed tools get deduplication only

Six more with the same problem

  • mapbox-search-integration/references/best-practices.md — the LRU SearchCache of geocoding results, replaced with session tokens + debouncing + deduplication, and pointing at permanent geocoding (permanent=true) as the supported way to store coordinates across sessions.
  • mapbox-navigation-patterns/AGENTS.md and mapbox-mcp-runtime-patterns/AGENTS.md — the Codex-side copies of the same two snippets. Same code, same problem.
  • mapbox-navigation-patterns/references/web-directions-api.md — listed "caching" as a reason to prefer polyline6.
  • mapbox-search-patterns/references/optimization-combining.md — "Cache coordinates for session", now explicit that temporary results are in-session only.
  • mapbox-token-security/AGENTS.md — recommended "Cache tiles in CDN / Implement client-side caching" as a rate-limit workaround.

Each rewritten section carries a short note linking the Terms of Service and Product Terms. Pointer text in three SKILL.md files and README.md updated to match.

Please check the licensing wording before merging

The note added to each section reads: "may not be cached or stored under Mapbox's standard self-serve (pay-as-you-go) terms, and Mapbox does not grant caching exceptions for pay-as-you-go accounts."

That phrasing is not quoted from published text. The public ToS, the Product Terms page, and the Directions API reference do not state a caching rule for Directions specifically. The only cleanly published line covers geocoding: temporary results are "for use during the current user session only," and permanent=true grants storage rights. Since this wording asserts the rule in public-facing docs, it is worth a sign-off from whoever owns licensing language.

Also in this PR

  • Regression evals. No eval covered this behavior, so nothing guarded the rewritten sections. One per affected skill, with negative expectations on the old pattern. Scores: nav Add MCP integration patterns for development and runtime #19 15/15, MCP Add Mapbox search integration workflow skill #5 18/18, search Add Mapbox ↔ MapLibre GL JS migration skill #4 13/15.
  • AGENTS.md as a second eval surface. AGENTS.md is hand-maintained and no eval read it, which is why these snippets needed fixing twice. scripts/eval.js now runs each eval against both surfaces (--surface=skill|agents|both).
  • --repeats=N for variance. Every eval ran once, so a delta could not be told from judge jitter — one eval scored 15/15 and 11/15 on identical content. Repeats report a mean, a 95% interval, and flag unstable evals.
  • Empty responses no longer score 0. With max_tokens at 4096 a thinking-heavy response could spend the whole budget before emitting text, scoring 0 for reasons unrelated to the skill. Default raised to 8192; empty responses now throw.
  • Eval answers no longer ship. build-plugin.js copied skills/ wholesale, so every published package included skills/*/evals/evals.json — the answer key for the suite that grades these skills.
  • depart_at guidance fixed. The text said only that "both profiles support depart_at", so the model picked driving for future departures. It now states that driving-traffic uses historical data with live traffic mixed in as the time approaches.

Follow-up, not in this PR

evals/baseline.json was regenerated here but is still single-sample, so it carries no interval. Worth regenerating with repeats once a value for N is settled.

Test plan

  • npm run check passes (format, spellcheck, markdownlint, Codex plugin build + validation, skill validation)
  • Affected skills re-run at --repeats=5 on both surfaces

🤖 Generated with Claude Code

@mattpodwysocki
mattpodwysocki requested a review from a team as a code owner September 29, 2026 21:30
@mattpodwysocki

mattpodwysocki commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor Author

Evals added and run

No eval covered this behavior, so nothing guarded these sections against regressing. Added one per affected skill, phrased as the question a pay-as-you-go user would realistically ask, with negative expectations on the old pattern.

Run against the rewritten skills (claude-sonnet-5, judge claude-sonnet-5):

Skill New eval Score Skill total
mapbox-navigation-patterns #19 — cache Directions responses in Redis for an hour 15/15 (100%) 218/228 (95.6%)
mapbox-mcp-runtime-patterns #5 — 1-hour TTL cache over callTool() for every tool 18/18 (100%) 77/78 (98.7%)
mapbox-search-integration #4 — persist geocoding results to Postgres 13/15 (87%) 44/51 (86.3%)

The search one lost 2 points on expectations written too demandingly (it didn't spell out the localStorage/IndexedDB contrast, and didn't mention debouncing). It got the substance right: refused the LRU/TTL cache and pointed at permanent=true.

Also fixed: the runner scored empty responses as legitimate zeros

mapbox-mcp-runtime-patterns #4 came back 0/15 with an empty model response. Cause was not the skill: max_tokens was 4096, and a thinking-heavy response spent 3357 tokens on thinking before emitting any text, so response.content had no text blocks and the judge dutifully scored nothing against every expectation.

scripts/eval.js now defaults to 8192 (EVAL_MAX_TOKENS to override) and throws on an empty response so the cause is visible rather than being laundered into a skill failure. That eval scores 15/15 now. Any past run with a suspicious zero is worth re-reading in that light.

Pre-existing failures, not touched here

evals/baseline.json was stale at the time of this comment: it predated this change and its stored mapbox-mcp-runtime-patterns response contained the CachedGeospatialAgent the old skill taught. Regenerated in a later commit.

Note the runner loads SKILL.md + references/*.md only — it never read AGENTS.md, so the Codex-side copies of these snippets were not covered by any eval. Addressed in a later commit on this branch.

@mattpodwysocki

Copy link
Copy Markdown
Contributor Author

AGENTS.md is now eval-covered, and the drift is bigger than expected

AGENTS.md is maintained by hand — nothing in scripts/ generates it — and no eval had ever read it. That is the root cause of the caching snippets needing the same fix twice, and of mapbox-navigation-patterns/AGENTS.md being 454 lines standing in for a 148-line SKILL.md plus 1,755 lines of references.

scripts/eval.js now runs every eval against both surfaces (--surface=skill|agents|both, default both). Skills with no AGENTS.md skip that surface. Results carry a surface field, the summary has a per-surface breakdown, and diff keys became skill#id@surface with surface-less baseline entries read as skill.

The measured gap is consistent across both skills that have an AGENTS.md here:

Skill SKILL.md + refs AGENTS.md
mapbox-navigation-patterns 226/228 (99%) 187/228 (82%)
mapbox-mcp-runtime-patterns 77/78 (99%) 64/78 (82%)

The caching evals hold up on both: nav #19 is 15/15 and 14/15, MCP #5 is 18/18 and 18/18.

The AGENTS.md losses are concentrated in Android Navigation SDK specifics that live only in references/:

So agents reading AGENTS.md get materially worse Android guidance, and some of it is wrong rather than merely absent. Not fixed in this PR — this only makes the drift visible. Closing it is a separate call: either generate AGENTS.md from SKILL.md + references the way plugins/mapbox/skills/ already is, or accept the condensed copy and shrink what it claims to cover. Worth its own issue.

Navigation #5 fixed (was 33%)

Asked which profile to use for a future depart_at with traffic-aware ETAs, the model recommended driving, reasoning live traffic cannot help a future departure. The text only said "both profiles support depart_at" without saying which to pick.

Per the Directions API reference, driving-traffic with a future depart_at uses "historical travel data and live traffic," with live traffic "gently mixed with historical data when the value for depart_at is close to current time." The guidance now says that outright, so a future departure is not a reason to drop to driving, and keeps arrive_by as the only reason to switch. Fixed in both copies; #5 now scores 9/9 on both surfaces, and the skill surface for this skill went 95.6% → 99%.

CI

The red Check links job was never this PR's doing — three 404s on one dead Flutter example URL that is present on main and in AGENTS.md + SKILL.md of mapbox-flutter-patterns. The fix already existed as 2f18086, stranded inside #79 behind a whole new skill. Lifted into #82, which unblocks this PR, #81, and anything else branched off main. My added ToS links all passed (342 links checked, 3 errors, all Flutter).

@zmofei zmofei left a comment •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The dedup-instead-of-cache rewrite looks right to me. A few things before merge:

Blocking

  1. Terms note goes beyond the published terms. It now says every API-backed MCP tool (Search, Matrix, Isochrone, Static Images, …) may not be cached, and that no exceptions are granted. The published terms don't say that. Please get a sign-off on the wording, or narrow it to what's published. The new evals (mapbox-navigation-patterns eval 19, mapbox-mcp-runtime-patterns eval 5) require the model to repeat this wording, so they should change with it.
  2. Evals still ship in the Claude Code plugin. build-plugin.js drops evals/ from the Codex package only. .claude-plugin/plugin.json points at ./skills/, so evals.json is still installed there.
  3. Wrong tool name. point_in_polygon_tool was removed from mcp-server; use points_within_polygon_tool. The list also isn't "the nine offline tools": the server has ~17 (convex_tool, length_tool, union_tool, …).

Should fix

  1. Errors are still scored as 0. The new empty-response throw is caught in runTask, which returns score: 0. So the summary and baseline still count it as a skill failure, and with --repeats it's averaged into the mean. Error runs should be left out of the scores.
  2. Please split the PR. The eval runner changes, Android AGENTS.md additions, depart_at change, Pydantic/LangChain rewrites, and the Flutter link (duplicates #82) each need their own review. Keeping this one to the caching fix makes it much easier to approve.
  3. The description is out of date. It says the baseline isn't regenerated, but the last commit regenerates it.
  4. --surface defaults to both. That doubles eval cost, and the headline total can't be compared with older baselines. I'd default to skill.

Minor

  • best-practices.md says Directions responses may not be stored, then suggests keeping the active route in memory for the session. Say where the line is, like the search section does.
  • The offline memo map is never evicted.

Several skills shipped code samples that cache Mapbox API responses —
Directions routes held for five minutes, all MCP tool results held for an
hour, geocoding results in an LRU cache. Storing or caching API responses
is restricted, so these patterns put readers at risk rather than helping
them, and a caveat on top of code that says to cache would not have fixed
that.

Each is replaced with a way to issue fewer requests that never retains a
response:

- mapbox-navigation-patterns: "Route Caching" becomes "Reducing Directions
  API Calls" — in-flight request deduplication, debouncing,
  overview=simplified, and the Matrix API for many-to-many travel times.
- mapbox-mcp-runtime-patterns: "Caching Strategy" (1-hour TTL on every
  tool result) becomes "Request Deduplication and Offline Memoization" —
  memoize only the local tools, which compute on the client and return no
  Mapbox content, and deduplicate the rest.
- mapbox-search-integration: the LRU SearchCache of geocoding results
  becomes session tokens, debouncing and deduplication, with permanent
  geocoding (permanent=true) as the documented way to store coordinates
  across sessions.

Also removed: "caching" as a listed reason to prefer polyline6, "Cache
coordinates for session" in mapbox-search-patterns, and the rate-limit
advice in mapbox-token-security that suggested caching tiles in a customer
CDN. The duplicated copies in the two AGENTS.md files are fixed alongside
their SKILL.md counterparts.

The notes added to each section describe only what Mapbox publishes. For
geocoding that is specific — temporary results are documented as being for
the current session, and permanent=true grants storage rights. Elsewhere
the notes say storing or caching API responses is restricted and point to
the Terms of Service and Product Terms for what a given plan permits,
rather than asserting a blanket prohibition the published terms do not
state.

best-practices.md now also draws the line it previously left implicit:
holding a response in memory to serve the request it was made for is
in-session use; writing it to disk, localStorage or a database is storage.

The local-tool list in the MCP samples was wrong. point_in_polygon_tool no
longer exists on the server — it is points_within_polygon_tool — and the
local set is 17 tools, not 9. Corrected against a tools/list call, with a
note to re-check per server version. The memo map is now bounded, since an
unbounded one leaks in a long-lived agent.
No eval covered this behavior, so nothing stopped the caching samples
coming back. One per affected skill, phrased as the question a
pay-as-you-go user would realistically ask, with negative expectations on
the old pattern:

- mapbox-navigation-patterns #19: caching Directions responses in Redis
- mapbox-mcp-runtime-patterns #5: a 1-hour TTL cache over callTool()
- mapbox-search-integration #4: persisting geocoding results to Postgres

Expectations check that the model treats caching as a licensing question
and points at the Mapbox terms, rather than requiring it to recite a
specific policy sentence — the skills no longer assert one, so an eval
demanding it would pin the docs to wording that is not published.

Scores at --repeats=3 on the skill surface: nav #19 97%, MCP #5 100%,
search #4 82%.
@mattpodwysocki

Copy link
Copy Markdown
Contributor Author

Thanks — this was a good catch list. All seven points are addressed. Three of them were things I'd got wrong, including two I'd previously reported as done.

Blocking

1. Terms note goes beyond the published terms. You're right, and the phrasing came from a paraphrase rather than from published text. Every instance of "may not be cached or stored under Mapbox's standard self-serve terms" and "does not grant caching exceptions" is gone (grep is clean). What's left says only what Mapbox publishes:

  • For geocoding, still specific, because it is published: temporary results are documented as being for the current session, and permanent=true grants storage rights.
  • Elsewhere: "Storing or caching Mapbox API responses is restricted — check the Terms of Service and Product Terms for what your plan permits," plus Sales for storage rights.

You were also right that the evals pinned the docs to that wording. Eval 19 and eval 5 now check that the model treats caching as a licensing question and points at the terms, rather than requiring it to recite a policy sentence. The "does not grant caching exceptions" expectation is deleted. Re-verified at --repeats=3: nav #19 97%, MCP #5 100%.

2. Evals still ship in the Claude Code plugin. Confirmed — .claude-plugin/plugin.json declares "skills": "./skills/", so my build-plugin.js filter only ever covered the Codex package. I'd reported this as fixed; it wasn't.

Rather than filter the second packager, eval definitions moved out of skills/ entirely — now evals/<skill-name>/evals.json. There are two manifests and nothing stops a third, so keeping answers outside skills/ holds for any packager. validate-skills.js and CONTRIBUTING.md follow the move. In #87.

3. Wrong tool name. Confirmed against a live tools/list: it's points_within_polygon_tool, and the local set is 17, not 9. I'd propagated the stale name into new samples without checking a server I had access to.

Fixed, and it was wider than this PR — 37 occurrences across mcp-runtime-patterns, geospatial-operations and search-patterns. I also inventoried every *_tool in the repo against all three servers in .mcp.json (31 + 22 + 3 = 56 tools). Three more are documented but exist on none of them: version_tool, create_token_tool, get_latest_mapbox_docs_tool. Those have no confirmed replacement so I've flagged rather than guessed. Ten runtime tools are also documented nowhere. In #90.

Should fix

4. Errors are still scored as 0. Correct, and my commit message claimed the opposite. The throw was caught in runTask, which returned score: 0, so it landed in the summary, the baseline and the --repeats mean. Failed runs are now excluded from the mean, the per-skill and per-surface totals, and the saved results, and get their own ERRORS (not scored) section. Verified with EVAL_MAX_TOKENS=1: 3 failed, 0 scored, totals print n/a instead of 0%. In #87.

5. Please split the PR. Done. #80 is now the caching fix only — 16 files, no scripts/, no baseline:

PR Scope
#80 caching → request-reduction, and the regression evals for it
#87 eval runner: surfaces, repeats, error handling, eval relocation
#88 depart_at profile guidance
#89 navigation AGENTS.md Android custom-UI APIs
#90 MCP tool names + Pydantic/LangChain APIs
#86 MCP wiring + clustering threshold — rebased onto #87

The Flutter link commit is dropped; #82 merged separately and main has it.

6. Description is out of date. Rewritten, and it no longer claims the baseline is unregenerated — the baseline isn't in #80 at all now.

7. --surface defaults to both. Agreed, changed to skill in #87. Routine runs cost half as much and totals stay comparable with older baselines; --surface=agents or both checks the drop-in copy deliberately. That's how the AGENTS.md drift PRs (#83, #84, #85, #89) were measured.

Minor

  • The in-memory contradiction. Fair — it said responses may not be stored and then suggested keeping the active route in memory. best-practices.md now draws the line explicitly: holding a response in memory to serve the request it was made for, and discarding it at session end, is in-session use; writing it to disk, localStorage or a database is storage.
  • Unbounded memo. Bounded now, oldest-first at 500 entries, in both the production.md and AGENTS.md samples.

One note on ordering: #86 is based on #87, and #90 touches the same eval files #87 relocates, so #87 wants to land before those two. #80, #88 and #89 are independent of all of it.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants