Skip to content

fix: reasoning opt-in for OpenAI-format gateways + learn-once observability - #13

Merged
jkyberneees merged 3 commits into
mainfrom
fix/reasoning-optin-learn-observer
Oct 5, 2026
Merged

jkyberneees merged 3 commits into
mainfrom
fix/reasoning-optin-learn-observer

Conversation

@jkyberneees

Copy link
Copy Markdown
Contributor

Fixes reasoning visibility through OpenAI-format gateways (OpenRouter, LiteLLM, vLLM).

What was wrong

  • The SDK never sent the gateway reasoning opt-in, so gateways like OpenRouter returned no reasoning at all.
  • Response reasoning arriving as message/delta.reasoning (OpenRouter string alias) or the typed message/delta.reasoning_details array was silently dropped; only reasoning_content was parsed.
  • The learn-once stream→buffered downgrade engaged silently and permanently; clients had no way to tell why streaming stopped.

What changed

  • Quirks.IncludeReasoning — sends include_reasoning: true (OpenRouter-documented legacy form of reasoning: {}) on OpenAI-format requests. Off by default.
  • Buffered + streaming parsers accept reasoning_content, reasoning, and reasoning_details (text/summary folded in order, encrypted entries skipped; precedence in that order).
  • SetLearnObserver(func(LearnEvent)) — one callback per engaged learn-once fallback (buffered, responses, none_effort, drop_stream_options) with provider, HTTP status, and provider message. Nil by default; no behavior change.

Verification

  • 12 new tests (wire + parse + observer), RED-first
  • go vet ./..., go test -race -count=1 ., golangci-lint run ./... all clean
  • 3-reviewer adversarial panel: 0 blockers, 0 majors; review fixes applied

…bility

OpenAI-format gateways (OpenRouter, LiteLLM, vLLM) can return reasoning in
shapes the SDK never read: the documented request opt-in (OpenRouter legacy
include_reasoning, equivalent to reasoning: {}) was never sent, and response
reasoning arriving as message/delta.reasoning or the typed
message/delta.reasoning_details array was silently dropped. Only the
LiteLLM-standardized reasoning_content was understood.

- Quirks.IncludeReasoning sends include_reasoning: true on OpenAI-format
  requests; off by default (strict endpoints reject unknown parameters).
- Buffered and streaming parsers now accept reasoning_content, reasoning,
  and reasoning_details (text/summary folded in order, encrypted skipped),
  precedence in that order.
- SetLearnObserver exposes the previously silent learn-once fallbacks
  (buffered downgrade, /responses retry, effort pinning, stream_options
  drop): one callback per engagement with kind, provider, status, and the
  provider message. The permanent stream-to-buffered downgrade is now
  observable instead of invisible.
Review findings from the adversarial PR panel:

- Gateways in the wild send reasoning/reasoning_content as JSON objects or
  arrays (e.g. OpenRouter reasoning objects). String-typed fields made the
  whole response or stream chunk fail to parse and abort the request. Both
  response-side fields now decode through a tolerant string type: JSON
  strings parse as before, every other shape (null, object, array, number,
  bool) is skipped instead of failing the request. This is a strict
  robustness improvement for shapes the SDK previously rejected outright.
- The streaming quirk test now asserts stream:true on the captured body, so
  a regression that routes CallStream through the buffered builder fails.
- Simplified a redundant fold condition (i > 0 && b.Len() > 0).

Known follow-up (documented, not in this PR): a gateway that rejects the
include_reasoning parameter outright has no learn-down path yet, unlike
stream_options. The flag is opt-in per provider, and the 400 surfaces
visibly, so the failure mode is loud rather than silent.
Tag-gated (//go:build e2e), OPENROUTER_API_KEY from env or .env, never
logged. Covers the PR #13 contract end to end against the live API:

- control-buffered: no-quirk call must succeed (reasoning presence is
  provider-side, logged not asserted)
- optin-buffered: quirk on, request succeeds, and ReasoningContent must
  mirror the documented fold of the raw wire response (wire-mirror
  assertion via a capturing transport — plaintext, encrypted-only, or
  object-shaped reasoning are all provider-side shapes)
- optin-stream: three attempts; every attempt must complete without a
  parse abort (the non-string reasoning regression class) with content
  flowing and ReasoningContent == concatenated reasoning deltas;
  plaintext reasoning deltas are probed across attempts because OpenRouter
  toggles them per request (observed: 74-delta run and zero-delta runs of
  the identical request within minutes)
@jkyberneees
jkyberneees merged commit 1a70f1e into main Oct 5, 2026
7 checks passed
@jkyberneees
jkyberneees deleted the fix/reasoning-optin-learn-observer branch October 5, 2026 15:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant