Guidance for AI coding agents (and humans) working in this repository.
go-llm-sdk — multi-provider Go SDK for LLM inference endpoints: OpenAI, Google Gemini, DeepSeek, Z.ai, Kimi (Moonshot), Anthropic, plus any OpenAI-compatible gateway. Module github.com/BackendStack21/go-llm-sdk. Go 1.25+, stdlib only — zero external dependencies. Flat single package (package llm) at the repo root.
make quality # fmt + vet + tests
make test-race # go test -race -count=1
make lint # golangci-lint (v2 config)
go test -count=1 -timeout 120s -race ./... # full race suite
go test -tags e2e -run 'TestE2E' -timeout 15m -v . # LIVE provider e2e (see below)- Always run tests with
-count=1and an explicit-timeout(house rule: no unbounded runs). - Coverage sits at ~99.0% of statements (unit suite). The residual is documented unreachable defensive code — do not pad with fake tests to move the number.
| File | Concern |
|---|---|
llm.go |
SDK entrypoint: New, options (WithProvider, FromEnv, …), Chat, Provider (model cache, shared learn-once state) |
message.go |
Canonical types: ChatRequest, Message, ChatResult, Delta, Usage, ToolDef |
chat.go |
providerClient: retry orchestration (buffered + streaming), error classification, learn-once consumption, SSE pump wiring, httpError parsing |
openai.go / gemini.go / anthropic.go |
Per-format request builders, response/stream mappers, model listing |
responses.go |
OpenAI Responses API (/v1/responses) for GPT-5.6+ tools+reasoning |
tts.go |
Text-to-speech: Speak/SpeakRequest/SpeakResult — OpenAI-compat /audio/speech + Gemini AUDIO modality, no transcoding |
controls.go |
Request controls: ToolChoice, ResponseFormat (incl. Anthropic forced-tool JSON emulation + foldAnthropicJSON), ParallelToolCalls, per-format mapping and boundary validation |
embed.go |
Embeddings: Embed/EmbedRequest/EmbedResult — OpenAI-compat /embeddings + Gemini batchEmbedContents |
stt.go |
Speech-to-text: Transcribe/TranscribeRequest/TranscribeResult — OpenAI-compat /audio/transcriptions (multipart), 25MB input cap |
sse.go |
SSE parser (abort-safe via done channel) + idle-watchdog pump |
retry.go |
RetryPolicy + the single buffered ladder withRetry (chat, tts, stt, embed); backoff/jitter/Retry-After/retrySleep (default 8 attempts, cap 30s) |
provider.go |
Built-in registry, quirks flags, config validation |
models.go |
ListModels orchestration (retries, caps, cache backing) |
auth.go, transport.go, errors.go |
Env resolution, pooled HTTP clients, typed errors |
The canonical type system is OpenAI-shaped; gemini.go/anthropic.go translate both directions. Provider quirks are explicit registry flags — never URL sniffing.
- Canonical finish reasons are
stop | length | tool_calls | content_filter | "". Unmapped provider stop reasons map to""on every format — provider-specific strings never leak. A turn with tool calls finishestool_callson every format (Gemini'sSTOPis mapped). - API keys never appear in error text,
String(), or any typed error. - Streaming: retries only before the first emitted delta; a failure after partial output returns the partial
*ChatResult+ wrapped error and is never retried; a premature close (no completion signal — on every format, Gemini included; an unmapped finish reason still counts as one) is an error, never a silent empty or partial success; the parser goroutine is always released (abort-safedoneprotocol). A successful non-SSE response is parsed directly with the usual body cap and format mapper. It learns buffered mode for future requests, but never discards the current generation or makes a replacement request; parse/read errors remain terminal for that response. - Unknown message roles are rejected at the SDK boundary (
ConfigError) — never dropped or reinterpreted per format. - Unknown provider data stays unknown (zero values) — no static guesses, no fallback model tables.
- Learn-once fallback state is per-
Provider, monotonic, atomic, shared by everyChatClient— never move it back to per-client. - Anthropic extended-thinking rounds trip via
ThinkingBlocks(everythinking/redacted_thinkingblock, each with its own signature, in order) or the legacy singleReasoningContent+ThinkingSignaturepair — replay is signature-gated and the thinking blocks go first. Signatures are per block, never concatenated. GeminithoughtSignatures round-trip per part (ToolCall.Signature,ThinkingSignature). Delta.ToolIndexis the call's position inChatResult.ToolCallson every format — never a provider content-block or output-item index.Usagehas one meaning everywhere:PromptTokensuncached-only, cache reads/writes in their fields (a DeepSeek miss is uncached input, not a write),CompletionTokensincludes reasoning,CachedTokensis a diagnostic subset never summed.- The request timeout is a whole-call budget, retries included, on every entry point (buffered, streaming, speech, transcription, embeddings).
- RED-first TDD: failing test first, then the fix. Table-driven tests;
httptestservers for hermetic coverage;newTestClienthelper pinsbackoffUnitto 1ms — restore package vars int.Cleanup. - Timing knobs for tests: the package var
backoffUnit(usefastBackoff(t)), and the process-wide idle default (atomic; usesetIdleForTest(t, d)). Operators override the idle watchdog per SDK withWithStreamIdleTimeout, or process-wide with the race-safeSetStreamIdleTimeout(positive values only). Retry shape is per SDK viaWithRetryPolicy. - The e2e suite (
e2e_test.go) is behind ae2ebuild tag and hits live APIs. It must never lose that tag. Keys come from env or a gitignored.env; contents are never logged; tests skip when a key is absent. Adding a provider = onee2eTargetentry; models overridable via<ID>_E2E_MODEL. - Live-provider behavior (e.g. DeepSeek eliding
reasoning_content) is not an SDK contract — probe softly, assert only what the SDK guarantees (call success, parsing, canonical finish). Model answer correctness is never an assertion. Elision is absorbed by an explicit provider quirk rather than hope: on a chat-completions request that carries tools, providers flaggedQuirks.EchoReasoningWithToolsalways get thereasoning_contentkey, empty included — including on assistant turns that made no tool call themselves (a request diverted to/responsesnever reaches that builder). That flag is off for every provider exceptdeepseek, and whether DeepSeek accepts a present-but-empty value is measured live by the tag-gatedTestE2EDeepSeekEmptyReasoningEchoprobe — not asserted by the unit suite. - Timer hygiene: since Go 1.23 no drain-before-
Resetis needed fortime.Timer. Note: on some dev machinestime.After+ select-default spin loops have hung — prefer deadline loops in tests.
- Branches:
feat/(features) andfix/(fixes). PR → CI green → squash-merge — never push straight tomain. - Conventional commits (
feat:,fix:,test:,docs:). - Docs live in
README.md(public contract) and this file (repo guidance). Keep both in sync with behavior in the same commit. - CI: build+test matrix (ubuntu/macOS with race+coverage, windows build/vet) + golangci-lint v2 — keep it green; lint failures fail the PR.
- odek adds this module as a dependency; its
internal/llmbecomes a thin shim. - odek's static
KnownProfilesare deleted; context windows come fromListModels(operator override for unknown models). - odek's internal LLM tests move here; the shim is removed.