Repository navigation
Top-10 improvements: thinking/signature replay, Gemini fixes, retry policy, request controls, embeddings - #15
Merged
Merged
Conversation
…stream completion - Anthropic: max_tokens always exceeds budget_tokens (unset → budget+8192, preset clamped below explicit cap, impossible explicit budget → ConfigError); temperature/top_p omitted under extended thinking - Anthropic: every thinking and redacted_thinking block is captured in order (ChatResult.ThinkingBlocks) and replayed verbatim; streamed signatures are per block, never concatenated across blocks - Gemini: thoughtSignature captured per function call (ToolCall.Signature) and on text parts (ThinkingSignature) and replayed on the same parts - Gemini: a function-call turn finishes as tool_calls, like every format - Streams that end at EOF without any completion signal are a premature close on every format, Gemini included; unmapped finish reasons still count as completion - Delta.ToolIndex is the call's position in ChatResult.ToolCalls on every format (was the Anthropic content-block / Responses output index) - Gemini model ids are path-escaped in the request URL - ChatResult.AssistantMessage() builds the replay turn with every field Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N76SRxxjJ4FgZGWmv4sGPS
…rver, ListModels hardening
- RetryPolicy{MaxAttempts, MaxBackoff} via WithRetryPolicy, honored by
Call, CallStream, Speak and Transcribe; one shared withRetry ladder
replaces the three copied loops (chat, tts, stt)
- Buffered Call, Speak and Transcribe share the request timeout as a
whole-call budget, retries included (as streaming already did) — a
failing provider can no longer hold a caller for timeout x attempts
- Backoff shift is clamped so long ladders never overflow to zero delay
- WithStreamIdleTimeout and WithLearnObserver: per-SDK settings that take
precedence over the process-wide defaults; SetStreamIdleTimeout is now
atomic (race-free while streams run)
- ListModels: malformed bodies are terminal (no retry), the page cap
returns ErrModelListTruncated instead of a silently truncated listing
(cap raised to 100 pages), and concurrent cache misses share one
upstream fetch; a follower never inherits its leader's cancellation
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N76SRxxjJ4FgZGWmv4sGPS
- Gemini: cachedContentTokenCount moves to CacheReadTokens (CacheReported set); thoughts are added to CompletionTokens so it includes reasoning on every format - DeepSeek: a cache miss is ordinary uncached input, not a cache write; hit/miss fields are authoritative and never double-counted with an echoed cached_tokens - Anthropic: cumulative message_delta usage updates input and cache volumes when present - Usage.InputTokens() / TotalTokens(); CachedTokens documented as a diagnostic subset of CacheReadTokens (never summed) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N76SRxxjJ4FgZGWmv4sGPS
…ls, Seed
- ToolChoice{auto,none,required,tool}: OpenAI/Responses tool_choice,
Anthropic tool_choice (auto/none/any/tool), Gemini
toolConfig.functionCallingConfig (AUTO/NONE/ANY + allowedFunctionNames)
- ResponseFormat{text,json_object,json_schema}: OpenAI response_format,
Responses text.format, Gemini responseMimeType + responseJsonSchema;
Anthropic emulates it with a forced synthetic tool whose input is
folded back into Content on buffered and streaming paths
- ParallelToolCalls: OpenAI/Responses parallel_tool_calls (only with
tools), Anthropic disable_parallel_tool_use
- Seed: OpenAI chat completions and Gemini generationConfig.seed
- Controls are validated at the SDK boundary (ConfigError), including
Anthropic forced tool use under extended thinking
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N76SRxxjJ4FgZGWmv4sGPS
- ProviderConfig.Headers / WithHeaders: applied last on every request (chat, streaming, models, speech, transcription); an empty value removes an SDK header (e.g. Authorization for api-key gateways); values never appear in String() or errors; the map is copied - ChatRequest.Extra: top-level body fields merged over the SDK's own on every format (Responses included); "stream" is reserved and non-JSON values fail fast with a ConfigError Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N76SRxxjJ4FgZGWmv4sGPS
…items
- SDK.Embed (EmbedRequest/EmbedResult): OpenAI-compatible /embeddings and
Gemini batchEmbedContents; vectors returned in input order with count
and index checks; shared retry ladder, budget, headers and error types;
Anthropic (no endpoint) is a ConfigError
- Message.IsError for tool results: Anthropic is_error, Gemini
{"error": …}; ConfigError outside the tool role
- Tool results accept text+image Parts: Anthropic tool_result blocks,
Responses input_text/input_image output, Gemini inlineData in the same
turn, chat completions images in a follow-up user message
- Responses: every reasoning item becomes a ThinkingBlock and is replayed
in order (was: last encrypted_content only); top-level stream "error"
events fail the stream with the provider message
- Unauthenticated errors name the env vars actually consulted (custom
providers without EnvKeys no longer suggest a variable nobody reads)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N76SRxxjJ4FgZGWmv4sGPS
Found by four independent reviewers (each finding reproduced RED first): - ListModels single-flight: a panicking fetch no longer wedges the provider (flight always released, followers get an error); an upstream timeout is the shared result instead of re-leading per follower - Anthropic: a signed thinking block with empty text keeps the required "thinking" key; preset budgets clamp to half an explicit MaxTokens (never a one-token answer); disable_parallel_tool_use never rides a "none" choice; JSON mode rejects Tools and non-object schemas; sparse message_delta usage never zeroes output tokens - Gemini: promptFeedback.blockReason completes as content_filter (was retried 8x as a premature close); seeds must fit in 32 bits - Responses: incomplete+content_filter maps to content_filter - Usage: gateway Anthropic+OpenAI cache fields never double-count; Moonshot top-level cached_tokens is understood - DeepSeek: Quirks.NoJSONSchema fails json_schema fast (json_object only) - Embed: Gemini models/ prefix accepted; batches of 100 (Gemini) / 2048 (OpenAI); gateways omitting index are trusted in order - Security: cross-host redirects are never followed (x-api-key and custom headers would leak); Provider/ProviderConfig redact under %v/%+v/%#v; Extra cannot replace validated keys (messages, tools, …); STT Filename/MIMEType control characters are rejected (header injection) - withRetry reports the context error on an interrupted last attempt; a mid-stream buffered fallback stays within the call's budget - README.md and AGENTS.md describe every new option, control and invariant Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N76SRxxjJ4FgZGWmv4sGPS
- Responses: an incomplete turn with incomplete_details content_filter now finishes content_filter (was length / tool_calls) - A request that cannot be built (*ConfigError from the transport layer) is terminal instead of being retried 8 times; base URLs that do not parse are rejected at wiring time - Dead branches removed (Responses failed-with-message, streaming learn retry without an APIError, an unreachable exhaustion return, assistant-role Parts in the OpenAI builder) - Edge tests for body read failures, learn triggers on the final attempt, 429-then-definitive ladders, long backoff ladders, same-host redirects, follower cancellation, Responses finish/usage/stream edges, content-part validation, model CreatedAt, speech transport errors Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N76SRxxjJ4FgZGWmv4sGPS
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N76SRxxjJ4FgZGWmv4sGPS
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Replaces #14. The commits are identical; only the branch is renamed to the repo's
fix/convention.Fixes the bugs and fills the feature gaps from the top-10 audit, plus everything found by four independent adversarial reviewers. Every behavior change was test-first: the test failed, then the fix landed. One commit per area so it can be reviewed in order.
Numbers: unit-suite coverage went from 92.2% to 99.0%. The
-racesuite runs in about 35s, down from about 108s, because ListModels no longer retries parse errors with real backoff.golangci-lintreports 0 issues, andgo vet -tags e2eplus the Windows and darwin builds are clean. I have not run the live e2e suite.Bugs fixed (top-10 audit)
max_tokensis now always larger thanbudget_tokens. An unsetMaxTokensbecomes budget + 8192. A preset larger than half an explicitMaxTokensis clamped to half of it, never below 1024. An explicit budget that cannot fit returns aConfigError.temperatureandtop_pare omitted while thinking is on.thoughtSignaturewas dropped. It is now kept per function call (ToolCall.Signature) and on text parts, and sent back on the same parts. Gemini 3 rejects tool turns that lack it.ChatResult.ThinkingBlocks/Message.ThinkingBlockskeep everythinkingandredacted_thinkingblock in order, each with its own signature."thinking"key.tool_calls.promptFeedbackfinishes withcontent_filterinstead of being retried 8 times.Delta.ToolIndexwas inconsistent. It is now the call's position inChatResult.ToolCallson every format. Before, Anthropic used the content-block index and Responses usedoutput_index.WithRetryPolicy(RetryPolicy{MaxAttempts, MaxBackoff})configures the ladder.withRetry.WithStreamIdleTimeoutandWithLearnObserverset these per SDK.SetStreamIdleTimeoutis now atomic, so calling it while streams run is race-free.CompletionTokens.PromptTokens; it is no longer counted as a cache write.cached_tokensis understood.Usage.InputTokens()andTotalTokens().incompleteturn with reasoncontent_filtermaps tocontent_filter.errorevent is handled.Features added
ToolChoice(auto / none / required / a named tool),ParallelToolCalls,Seed.ResponseFormatfor json_object / json_schema. OpenAI usesresponse_formatortext.format; Gemini usesresponseMimeType+responseJsonSchema; Anthropic uses a forced tool whose input is folded back intoContent, both buffered and streamed.WithHeaders, applied last on every request (an empty value removes a header), andChatRequest.Extrafor top-level body fields. The keys the SDK validates are reserved.SDK.Embed: OpenAI-compatible/embeddingsand GeminibatchEmbedContents. Results come back in input order, with batching and count/index checks.Message.IsError, and imagePartson tool results, mapped per format.ChatResult.AssistantMessage()builds the assistant turn with every replay field.ErrModelListTruncatedinstead of being silently truncated.Security hardening (from review)
Authorizationacross hosts, sox-api-keyand custom headers would otherwise leak.ProviderandProviderConfignow hide the API key and header values under%v,%+vand%#v.Filename/MIMETypeis rejected, to stop header injection.Behavior changes to review
Call,Speak,TranscribeandEmbednow stop at the request timeout in total (default 120s). Before, each of up to 8 attempts got the full timeout.PromptTokens, notCacheCreationTokens. This is intentional but differs from the old odek-parity test, which I updated.finishReasonis now an error, as on every other format.Quirks.NoJSONSchemais set for DeepSeek, so ajson_schemaresponse format fails fast. That DeepSeek supports onlyjson_objectcomes from provider docs and has not been checked against the live API.Thinkingon the internal Anthropic block is now a pointer. No exported field was removed.Review process
Four adversarial reviewers each worked in an isolated worktree: three Sonnet (streaming and thinking, retry and concurrency, request controls and wire formats) and one Haiku (invariants and security). Every finding they proved was fixed test-first in
4b60ce2orbd73bba.Not acted on, and documented instead:
Extramerges only top-level keys (shallow).audio/mpegMIME fallback predates this PR; left as is.The new surface and invariants 8–10 are documented in README.md and AGENTS.md.
🤖 Generated with Claude Code
https://claude.ai/code/session_01N76SRxxjJ4FgZGWmv4sGPS
Generated by Claude Code