From 87b9a037bb681655b823ec4e736e49873c4bc185 Mon Sep 17 00:00:00 2001 From: Danny Avila Date: Mon, 14 Sep 2026 19:57:48 -0400 Subject: [PATCH] =?UTF-8?q?=F0=9F=A7=AD=20refactor:=20Settle=20Turn=20Deli?= =?UTF-8?q?very=20Routing=20Once=20in=20Agent=20Initialization?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit `initializeAgent` resolved the provider and its client options only after the turn's files were loaded and admitted, and the two delivery readers each rebuilt the attachment routing from the agent object at their own moment: `BaseClient` took the custom-endpoint dialect from the already-swapped `agent.provider`, the run-file encoder from whatever child config it was handed. Move `getProviderConfig`/`getOptions` ahead of file discovery, where nothing in between fed them, and settle one `deliveryRouting` value with every input final: the file policy under the endpoint's own name, the dialect its config declares, the Responses API decision the model call uses, and the transcription setting. `InitializedAgent`, the child encoder and `BaseClient` consume that value; `resolveTurnLLMDeliveryPath` is the one place a stored route is resolved again. --- CONTEXT.md | 1 + api/app/clients/BaseClient.js | 52 ++---- api/app/clients/specs/BaseClient.test.js | 50 +++--- api/server/controllers/agents/client.test.js | 3 + .../services/Endpoints/agents/skillDeps.js | 2 +- .../src/agents/__tests__/initialize.test.ts | 84 +++++++++- .../api/src/agents/files/delivery.spec.ts | 69 ++++++++ packages/api/src/agents/files/delivery.ts | 50 ++++++ packages/api/src/agents/files/encode.spec.ts | 25 ++- packages/api/src/agents/files/encode.ts | 37 ++--- packages/api/src/agents/files/index.ts | 1 + packages/api/src/agents/files/session.spec.ts | 5 + packages/api/src/agents/initialize.ts | 153 ++++++++++-------- .../src/resolve-llm-delivery-path.spec.ts | 71 ++++++++ .../src/resolve-llm-delivery-path.ts | 51 ++++++ 15 files changed, 493 insertions(+), 161 deletions(-) create mode 100644 packages/api/src/agents/files/delivery.spec.ts create mode 100644 packages/api/src/agents/files/delivery.ts diff --git a/CONTEXT.md b/CONTEXT.md index b390f49505c..62ad3558459 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -9,6 +9,7 @@ - **Agent execution host**: the protocol-neutral module that owns run admission, disconnect cancellation, provider-start fencing, and terminal settlement. Protocol implementations execute behind its callback interface; HTTP adapters retain validation and final stream rendering. - **Agent execution enrollment**: the durable, protocol-neutral lifecycle authority for an admitted Agent run. It is created under the authenticated user and tenant before user-owned initialization, rechecks the shared owner-deletion admission fence after registration, exposes the only provider abort signal, fences exact provider start, terminalizes the run, waits for every trailing usage, artifact, and stored-response write, and acknowledges provider drain last. A transient terminalization failure is reconciled after trailing writes; provider drain is never acknowledged while the exact job remains nonterminal. Delete-all holds the owner fence, drains every owner run before selecting its first persistence snapshot, and repeats both the drain and an idempotent owner-persistence sweep after any recovered fence lapse before releasing admission. Exact-conversation deletion additionally performs an unconditional idempotent cleanup over its immutable deleted-ID set because a fully drained run may leave the active index after racing the first delete; only the explicit empty result is benign, while storage failures remain fatal. Chat Completions, Responses, Channels, and future ingress adapters share this authority without moving LibreChat persistence policy into the Agents SDK. - **Agent turn execution plan**: the immutable, request-local decision compiled once after authentication, agent resolution, and tool initialization. It records the trusted turn origin, conversation lineage, pause capability, binding/action context, and the preferred checkpoint, history, or fresh state-loading strategy without executing the model or owning persistence. Checkpoint failure falls back to durable history within the same Agents lifecycle. +- **Turn delivery routing**: the per-agent, request-local value that decides how each attachment reaches the model on one turn (`provider`, `text`, or `none`). Initialization settles it once, after the provider swap and the Responses API decision, under the endpoint's own name and the media dialect its config declares. Every reader of a turn route consumes that one value rather than deriving it from the agent. A stored route is an upload-time inference that this value resolves again for the turn; a destination the user chose stands. - **Effective agent selection**: the resolved endpoint and agent identity after an enforced model spec is applied. Authorization and agent loading must consume this same identity before the Agent run envelope is initialized. - **MCP runtime request body**: trusted chat identifiers supplied only while an MCP server handles an agent request. It enables request-scoped header placeholders without retaining user-specific request data on a shared server definition. - **MCP direct OpenID bearer**: an operator-trusted remote MCP credential mode that resolves the logged-in user's live OpenID access token into an Authorization header. It may replace one rejected connection after a forced session refresh, but it never replays the rejected tool invocation automatically. diff --git a/api/app/clients/BaseClient.js b/api/app/clients/BaseClient.js index 81f89987604..21ce3ec45db 100644 --- a/api/app/clients/BaseClient.js +++ b/api/app/clients/BaseClient.js @@ -33,19 +33,15 @@ const { isCompactedLeaf, excludedKeys, EModelEndpoint, - mergeFileConfig, isParamEndpoint, isAgentsEndpoint, isEphemeralAgentId, supportsBalanceCheck, isBedrockDocumentType, HITL_MESSAGE_FILTER_FIELDS, - getEndpointFileConfig, stripReasoningLabelMetadata, - resolveUploadLLMDeliveryPath, - isSpeechProviderConfigured, + resolveTurnLLMDeliveryPath, resolveUseResponsesApi, - getCustomEndpointProvider, } = require('librechat-data-provider'); const { getStrategyFunctions } = require('~/server/services/Files/strategies'); const { logViolation } = require('~/cache'); @@ -242,10 +238,6 @@ class BaseClient { this.currentMessages = []; /** @type {import('librechat-data-provider').VisionModes | undefined} */ this.visionMode; - /** @type {import('librechat-data-provider').FileConfig | undefined} */ - this._mergedFileConfig; - /** @type {import('librechat-data-provider').EndpointFileConfig | undefined} */ - this._endpointFileConfig; } setOptions() { @@ -1804,38 +1796,9 @@ class BaseClient { }); } - /** Re-resolves an inferred upload route against the provider handling this turn. */ + /** The route a stored attachment takes on this turn, from the routing settled for the agent. */ getAttachmentDeliveryPath(file) { - if (!this._mergedFileConfig) { - this._mergedFileConfig = mergeFileConfig(this.options.req?.config?.fileConfig); - /* Agent file policy is configured under the endpoint it names, not the client - * family initialization may rewrite it to. */ - const agentEndpoint = this.options.agent?.endpoint ?? this.options.agent?.provider; - this._deliveryEndpoint = agentEndpoint ?? this.options.endpoint; - this._endpointFileConfig = getEndpointFileConfig({ - fileConfig: this._mergedFileConfig, - endpoint: this._deliveryEndpoint, - endpointType: agentEndpoint != null ? undefined : this.options.endpointType, - }); - } - - return file.llmDeliveryPath == null || file.metadata?.destinationChosen === true - ? file.llmDeliveryPath - : resolveUploadLLMDeliveryPath({ - /* Conversion changes the stored type, so use the type routing originally saw. */ - mimeType: file.metadata?.routingMimeType ?? file.type, - endpointConfig: this._endpointFileConfig, - fileConfig: this._mergedFileConfig, - endpoint: this._deliveryEndpoint, - endpointProvider: - this.options.agent?.provider ?? - getCustomEndpointProvider( - this.options.req?.config?.endpoints?.custom, - this._deliveryEndpoint, - ), - useResponsesApi: this.usesResponsesApi(), - sttConfigured: isSpeechProviderConfigured(this.options.req?.config?.speech?.stt), - }); + return resolveTurnLLMDeliveryPath(this.options.agent?.deliveryRouting, file); } async processAttachments(message, attachments) { @@ -1849,6 +1812,7 @@ class BaseClient { const allFiles = []; const provider = this.options.agent?.provider ?? this.options.endpoint; const isBedrock = provider === EModelEndpoint.bedrock; + const deliveryRouting = this.options.agent?.deliveryRouting; /* The stored path records what upload time inferred from the endpoint it saw, and this * turn may be running somewhere else: audio stored as `provider` under Google reaches @@ -1897,9 +1861,11 @@ class BaseClient { allFiles.push(file); } else if ( file.type && - this._mergedFileConfig && - this._endpointFileConfig?.supportedMimeTypes && - this._mergedFileConfig.checkType(file.type, this._endpointFileConfig.supportedMimeTypes) + deliveryRouting?.endpointConfig.supportedMimeTypes && + deliveryRouting.fileConfig.checkType( + file.type, + deliveryRouting.endpointConfig.supportedMimeTypes, + ) ) { categorizedAttachments.documents.push(file); allFiles.push(file); diff --git a/api/app/clients/specs/BaseClient.test.js b/api/app/clients/specs/BaseClient.test.js index 26d8f3ccfb6..ca86332af0b 100644 --- a/api/app/clients/specs/BaseClient.test.js +++ b/api/app/clients/specs/BaseClient.test.js @@ -1,6 +1,6 @@ const { Constants, ContentTypes, EModelEndpoint } = require('librechat-data-provider'); const BaseClientClass = require('../BaseClient'); -const { ContentFilterError } = require('@librechat/api'); +const { ContentFilterError, resolveTurnDeliveryRouting } = require('@librechat/api'); const { FakeClient, initializeFakeClient } = require('./FakeClient'); function deferred() { @@ -3192,8 +3192,6 @@ describe('BaseClient', () => { TestClient.options = { endpoint: EModelEndpoint.openAI, }; - TestClient._mergedFileConfig = undefined; - TestClient._endpointFileConfig = undefined; TestClient.addImageURLs = jest.fn(async (message, files) => { message.image_urls = ['encoded-image']; return files; @@ -3207,9 +3205,18 @@ describe('BaseClient', () => { TestClient.addAudios = jest.fn(async (_message, files) => files); }); - /* The stored path is an upload-time inference, so delivery re-resolves it for the - * endpoint running the turn. A test asserting a route has to configure that route - * rather than rely on the stored value alone. */ + /** The routing initialization settles for an agent, from the request config it reads. */ + const routedAgent = (agent) => ({ + ...agent, + deliveryRouting: resolveTurnDeliveryRouting({ + agent, + config: TestClient.options.req?.config, + }), + }); + + /* The stored path is an upload-time inference, so delivery resolves it again by the + * routing settled for the agent running the turn. A test asserting a route has to + * configure that route rather than rely on the stored value alone. */ const routeTo = (path, ...mimeTypes) => { TestClient.options.req = { config: { @@ -3224,8 +3231,10 @@ describe('BaseClient', () => { }, }, }; - TestClient._mergedFileConfig = undefined; - TestClient._endpointFileConfig = undefined; + TestClient.options.agent = routedAgent({ + provider: EModelEndpoint.openAI, + endpoint: EModelEndpoint.openAI, + }); }; test('keeps a none image in returned files without adding image URLs', async () => { @@ -3327,17 +3336,19 @@ describe('BaseClient', () => { expect(TestClient.addImageURLs).not.toHaveBeenCalled(); }); - test('reads the Responses setting from a plain conversation too', async () => { - /* A non-agent Azure chat carries it in model options, and reading only the agent - * parameters re-resolves a natively supported PDF to text, which the record has - * none of, so the model receives nothing. */ + test('reads the Responses setting the turn runs on from the settled routing', async () => { + /* Azure sends a PDF natively only under the Responses API. The routing carries the + * decision initialization made, so a record stored as `provider` is not resolved + * again to text it has none of, which would leave the model with nothing. */ TestClient.options = { - endpoint: EModelEndpoint.azureOpenAI, + endpoint: EModelEndpoint.agents, req: { config: { fileConfig: undefined } }, }; - TestClient.modelOptions = { useResponsesApi: true }; - TestClient._mergedFileConfig = undefined; - TestClient._endpointFileConfig = undefined; + TestClient.options.agent = routedAgent({ + provider: EModelEndpoint.azureOpenAI, + endpoint: EModelEndpoint.azureOpenAI, + model_parameters: { useResponsesApi: true }, + }); const message = {}; const file = { user: 'user1', @@ -3353,7 +3364,6 @@ describe('BaseClient', () => { await TestClient.processAttachments(message, [file]); expect(TestClient.addDocuments).toHaveBeenCalled(); - TestClient.modelOptions = undefined; }); test('resolves a custom endpoint policy by the name the admin configured', async () => { @@ -3377,8 +3387,7 @@ describe('BaseClient', () => { }, }, }; - TestClient._mergedFileConfig = undefined; - TestClient._endpointFileConfig = undefined; + TestClient.options.agent = routedAgent(TestClient.options.agent); const message = {}; const file = { user: 'user1', @@ -3418,8 +3427,7 @@ describe('BaseClient', () => { }, }, }; - TestClient._mergedFileConfig = undefined; - TestClient._endpointFileConfig = undefined; + TestClient.options.agent = routedAgent(TestClient.options.agent); const message = {}; const file = { user: 'user1', diff --git a/api/server/controllers/agents/client.test.js b/api/server/controllers/agents/client.test.js index 7a986a19cf0..365b762da5f 100644 --- a/api/server/controllers/agents/client.test.js +++ b/api/server/controllers/agents/client.test.js @@ -5471,6 +5471,9 @@ describe('AgentClient - titleConvo', () => { }, }; mockRes = {}; + mockAgent.deliveryRouting = jest + .requireActual('@librechat/api') + .resolveTurnDeliveryRouting({ agent: mockAgent, config: mockReq.config }); client = new AgentClient({ req: mockReq, diff --git a/api/server/services/Endpoints/agents/skillDeps.js b/api/server/services/Endpoints/agents/skillDeps.js index c8aa0f2b303..aaa51f7fc84 100644 --- a/api/server/services/Endpoints/agents/skillDeps.js +++ b/api/server/services/Endpoints/agents/skillDeps.js @@ -291,10 +291,10 @@ function buildAgentToolContext({ agent, config }) { agent, fileEncodingAgent: { provider: config.provider, - endpoint: config.endpoint, model_parameters: config.model_parameters, imageDetail: config.imageDetail, agentContextAttachments: config.agentContextAttachments, + deliveryRouting: config.deliveryRouting, }, /** Per-agent resolved endpoint token/pricing config. Retained here because * `agentToolContexts` is the one map that holds every agent — including diff --git a/packages/api/src/agents/__tests__/initialize.test.ts b/packages/api/src/agents/__tests__/initialize.test.ts index c67e45f03d3..0f20e80daf5 100644 --- a/packages/api/src/agents/__tests__/initialize.test.ts +++ b/packages/api/src/agents/__tests__/initialize.test.ts @@ -1655,7 +1655,6 @@ describe('initializeAgent — attachment scoping', () => { expect(db.updateFilesUsage).not.toHaveBeenCalled(); expect(primeResources).not.toHaveBeenCalled(); expect(loadTools).not.toHaveBeenCalled(); - expect(mockGetProviderConfig).not.toHaveBeenCalled(); }); }); @@ -4499,3 +4498,86 @@ describe('initializeAgent — provider-native web search role gate', () => { expect(result.tools).toContainEqual(OPENAI_SEARCH); }); }); + +describe('initializeAgent turn delivery routing', () => { + beforeEach(() => { + jest.clearAllMocks(); + }); + + it('settles the routing once, after the provider swap and the Responses API decision', async () => { + const { agent, req, res, loadTools, db } = createMocks({ provider: 'MyClaude' }); + req.config = { + fileConfig: { endpoints: { MyClaude: { supportedMimeTypes: ['video/mp4'] } } }, + endpoints: { custom: [{ name: 'MyClaude', provider: 'anthropic' }] }, + } as unknown as ServerRequest['config']; + mockGetProviderConfig.mockReturnValue({ + overrideProvider: Providers.ANTHROPIC, + getOptions: jest + .fn() + .mockResolvedValue({ llmConfig: { model: 'test-model', useResponsesApi: true } }), + }); + + const result = await initializeAgent( + { + req, + res, + agent, + loadTools, + endpointOption: { endpoint: EModelEndpoint.agents }, + allowedProviders: new Set(['MyClaude']), + isInitialAgent: true, + }, + db, + ); + + expect(result.provider).toBe(Providers.ANTHROPIC); + expect(result.deliveryRouting).toMatchObject({ + endpoint: 'MyClaude', + endpointProvider: 'anthropic', + useResponsesApi: true, + sttConfigured: false, + }); + expect(result.deliveryRouting.endpointConfig.supportedMimeTypes).toEqual([/video\/mp4/]); + }); + + it('resolves the provider and its options before any attachment is loaded', async () => { + const { filterFilesByEndpointRuntimeConfig } = jest.requireMock('~/files') as { + filterFilesByEndpointRuntimeConfig: jest.Mock; + }; + const { agent, req, res, loadTools, db } = createMocks(); + const getOptions = jest.fn().mockResolvedValue({ llmConfig: { model: 'test-model' } }); + mockGetProviderConfig.mockReturnValue({ overrideProvider: Providers.OPENAI, getOptions }); + const file = { + file_id: 'file-1', + filename: 'notes.txt', + type: 'text/plain', + bytes: 10, + text: 'notes', + llmDeliveryPath: 'text', + } as IMongoFile; + (db.getFiles as jest.Mock).mockResolvedValueOnce([file]); + filterFilesByEndpointRuntimeConfig.mockImplementationOnce( + (_config: ServerRequest['config'], { files }: { files: IMongoFile[] }) => files, + ); + + await initializeAgent( + { + req, + res, + agent, + loadTools, + requestFiles: [file], + endpointOption: { endpoint: EModelEndpoint.agents }, + allowedProviders: new Set([Providers.OPENAI]), + isInitialAgent: true, + }, + db, + ); + + const [optionsOrder] = getOptions.mock.invocationCallOrder; + const [filesOrder] = (db.getFiles as jest.Mock).mock.invocationCallOrder; + const [toolsOrder] = loadTools.mock.invocationCallOrder; + expect(optionsOrder).toBeLessThan(filesOrder); + expect(filesOrder).toBeLessThan(toolsOrder); + }); +}); diff --git a/packages/api/src/agents/files/delivery.spec.ts b/packages/api/src/agents/files/delivery.spec.ts new file mode 100644 index 00000000000..ddcabb800b8 --- /dev/null +++ b/packages/api/src/agents/files/delivery.spec.ts @@ -0,0 +1,69 @@ +import { EModelEndpoint } from 'librechat-data-provider'; +import type { TurnDeliveryConfig } from './delivery'; +import { resolveTurnDeliveryRouting } from './delivery'; + +const config = { + fileConfig: { + endpoints: { + MyClaude: { supportedMimeTypes: ['video/mp4'] }, + [EModelEndpoint.openAI]: { fileLimit: 3 }, + }, + }, + endpoints: { custom: [{ name: 'MyClaude', provider: 'anthropic' }] }, + speech: { stt: { openai: { apiKey: 'key', model: 'whisper-1' } } }, +} as unknown as TurnDeliveryConfig; + +describe('resolveTurnDeliveryRouting', () => { + it('reads the file policy under the endpoint name, before and after the provider swap', () => { + /* Initialization first stores the endpoint name in both fields and only later replaces + * the provider with the backing client, so neither the policy nor the dialect may move. */ + const before = resolveTurnDeliveryRouting({ + agent: { provider: 'MyClaude', endpoint: 'MyClaude' }, + config, + }); + const after = resolveTurnDeliveryRouting({ + agent: { provider: 'anthropic', endpoint: 'MyClaude' }, + config, + }); + + expect(after).toEqual(before); + expect(before.endpoint).toBe('MyClaude'); + expect(before.endpointConfig.supportedMimeTypes).toEqual([/video\/mp4/]); + expect(before.endpointProvider).toBe('anthropic'); + }); + + it('leaves the dialect undefined for a built-in or OpenAI-compatible endpoint', () => { + expect( + resolveTurnDeliveryRouting({ + agent: { provider: EModelEndpoint.openAI, endpoint: EModelEndpoint.openAI }, + config, + }).endpointProvider, + ).toBeUndefined(); + expect( + resolveTurnDeliveryRouting({ agent: { provider: 'MyGateway' }, config }).endpointProvider, + ).toBeUndefined(); + }); + + it('routes an agent loaded without an endpoint name by its provider', () => { + const routing = resolveTurnDeliveryRouting({ + agent: { provider: EModelEndpoint.openAI }, + config, + }); + + expect(routing.endpoint).toBe(EModelEndpoint.openAI); + expect(routing.endpointConfig.fileLimit).toBe(3); + }); + + it('carries the Responses API decision and the transcription setting', () => { + const routing = resolveTurnDeliveryRouting({ + agent: { provider: EModelEndpoint.azureOpenAI, model_parameters: { useResponsesApi: true } }, + config, + }); + + expect(routing.useResponsesApi).toBe(true); + expect(routing.sttConfigured).toBe(true); + expect( + resolveTurnDeliveryRouting({ agent: { provider: EModelEndpoint.openAI }, config: {} }), + ).toMatchObject({ useResponsesApi: undefined, sttConfigured: false }); + }); +}); diff --git a/packages/api/src/agents/files/delivery.ts b/packages/api/src/agents/files/delivery.ts new file mode 100644 index 00000000000..ee5b04450b1 --- /dev/null +++ b/packages/api/src/agents/files/delivery.ts @@ -0,0 +1,50 @@ +import { + mergeFileConfig, + getEndpointFileConfig, + resolveUseResponsesApi, + getCustomEndpointProvider, + isSpeechProviderConfigured, +} from 'librechat-data-provider'; +import type { TurnDeliveryRouting } from 'librechat-data-provider'; +import type { AppConfig } from '@librechat/data-schemas'; + +/** The app config a turn's attachment routing reads. */ +export type TurnDeliveryConfig = Pick; + +/** The fields of an agent that route its attachments. */ +export interface TurnDeliveryAgent { + provider: string; + /** The endpoint name initialization records before it swaps `provider` for the backing + * client; an agent loaded without one is routed by its provider. */ + endpoint?: string | null; + /** The Responses API setting the turn runs on, once initialization has decided it. */ + model_parameters?: { useResponsesApi?: boolean } | null; +} + +/** + * Settles how the agent running a turn receives its attachments. + * + * Read once, after initialization has resolved the backing provider and the Responses API + * decision: the file policy is the one configured under the endpoint's own name, the media + * dialect is the one its config declares rather than the client family it runs as, and the + * Responses setting is the one the model call uses. Every reader of a turn route consumes the + * returned value, so delivery, steering and child run-file encoding cannot answer differently. + */ +export function resolveTurnDeliveryRouting({ + agent, + config, +}: { + agent: TurnDeliveryAgent; + config?: TurnDeliveryConfig; +}): TurnDeliveryRouting { + const endpoint = agent.endpoint ?? agent.provider; + const fileConfig = mergeFileConfig(config?.fileConfig); + return { + fileConfig, + endpointConfig: getEndpointFileConfig({ fileConfig, endpoint }), + endpoint, + endpointProvider: getCustomEndpointProvider(config?.endpoints?.custom, endpoint), + useResponsesApi: resolveUseResponsesApi(agent.model_parameters?.useResponsesApi), + sttConfigured: isSpeechProviderConfigured(config?.speech?.stt), + }; +} diff --git a/packages/api/src/agents/files/encode.spec.ts b/packages/api/src/agents/files/encode.spec.ts index 753db6d0982..fe81e1f9a49 100644 --- a/packages/api/src/agents/files/encode.spec.ts +++ b/packages/api/src/agents/files/encode.spec.ts @@ -3,6 +3,7 @@ import type { TFile } from 'librechat-data-provider'; import type { RunFileEncodingAgent, RunFileMessageEncoderDeps } from './encode'; import type { ServerRequest } from '~/types'; import { AgentAttachmentLimitError, AgentAttachmentPolicyError } from '../attachments'; +import { resolveTurnDeliveryRouting } from './delivery'; import { createRunFileMessageEncoder } from './encode'; jest.mock('~/utils/tokenizer', () => ({ countTokens: (text: string) => text.length })); @@ -28,14 +29,32 @@ const nativeDocument = { }; const nativeImage = { type: 'image_url', image_url: { url: 'data:image/png;base64,aW1n' } }; +/** A child as the host loads it, before initialization settles its delivery routing. */ +type LoadedAgent = Omit & { + endpoint?: string; + useResponsesApi?: boolean; +}; + function setup({ fileConfig = {}, agents = { child: { provider: 'openAI' } }, }: { fileConfig?: NonNullable['fileConfig']; - agents?: Record; + agents?: Record; } = {}) { const req = { body: {}, config: { fileConfig } } as ServerRequest; + const initialized = Object.fromEntries( + Object.entries(agents).map(([id, { endpoint, useResponsesApi, ...agent }]) => [ + id, + { + ...agent, + deliveryRouting: resolveTurnDeliveryRouting({ + agent: { provider: agent.provider, endpoint, model_parameters: { useResponsesApi } }, + config: req.config, + }), + }, + ]), + ); const encodeImages = jest.fn(async () => ({ image_urls: [nativeImage] })); const encodeDocuments = jest.fn(async () => ({ documents: [nativeDocument] })); const encodeAudios = jest.fn(async () => ({ audios: [{ type: 'media', data: 'audio' }] })); @@ -46,7 +65,7 @@ function setup({ >(extractFileContext); const deps: RunFileMessageEncoderDeps = { req, - getAgent: (id) => agents[id], + getAgent: (id) => initialized[id], encodeImages, encodeDocuments, encodeAudios, @@ -65,7 +84,7 @@ describe('createRunFileMessageEncoder', () => { provider: 'openAI', endpoint: 'child-provider', model: 'child-model', - model_parameters: { useResponsesApi: true }, + useResponsesApi: true, imageDetail: ImageDetail.high, }, }, diff --git a/packages/api/src/agents/files/encode.ts b/packages/api/src/agents/files/encode.ts index d4de921577a..4b953ec173e 100644 --- a/packages/api/src/agents/files/encode.ts +++ b/packages/api/src/agents/files/encode.ts @@ -3,14 +3,10 @@ import { HumanMessage } from '@librechat/agents/langchain'; import { FileSources, EModelEndpoint, - mergeFileConfig, - getEndpointFileConfig, isBedrockDocumentType, - resolveUseResponsesApi, - isSpeechProviderConfigured, - resolveUploadLLMDeliveryPath, + resolveTurnLLMDeliveryPath, } from 'librechat-data-provider'; -import type { TFile, ImageDetail } from 'librechat-data-provider'; +import type { TFile, ImageDetail, TurnDeliveryRouting } from 'librechat-data-provider'; import type { BaseMessage } from '@librechat/agents/langchain'; import type { ServerRequest, StrategyFunctions } from '~/types'; import type { TokenCountFn } from '~/utils/text'; @@ -28,11 +24,12 @@ type ContentBlock = Exclude[number]; /** The already loaded child configuration; storage documents never cross this boundary. */ export interface RunFileEncodingAgent { provider: string; - endpoint?: string | null; model?: string | null; - model_parameters?: { model?: string; useResponsesApi?: boolean }; + model_parameters?: { model?: string }; imageDetail?: ImageDetail; agentContextAttachments?: readonly TFile[]; + /** How the child receives attachments, settled when its configuration was initialized. */ + deliveryRouting: TurnDeliveryRouting; } export interface RunFileEncodingParams { @@ -80,34 +77,22 @@ export function createRunFileMessageEncoder( if (!agent) { throw new Error('The target agent is not available for shared file delivery.'); } - const endpoint = agent.endpoint ?? agent.provider; - const fileConfig = mergeFileConfig(deps.req.config?.fileConfig); - const endpointConfig = getEndpointFileConfig({ fileConfig, endpoint }); - const useResponsesApi = resolveUseResponsesApi(agent.model_parameters?.useResponsesApi); + const { deliveryRouting } = agent; + const { endpoint, fileConfig, endpointConfig } = deliveryRouting; const params: RunFileEncodingParams = { provider: agent.provider, endpoint, model: agent.model_parameters?.model ?? agent.model ?? undefined, - useResponsesApi, + useResponsesApi: deliveryRouting.useResponsesApi, imageDetail: agent.imageDetail, }; const resolveDelivery = (file: TFile): TFile => { - if (file.llmDeliveryPath == null || file.metadata?.destinationChosen === true) { + const llmDeliveryPath = resolveTurnLLMDeliveryPath(deliveryRouting, file); + if (llmDeliveryPath == null || llmDeliveryPath === file.llmDeliveryPath) { return file; } - return { - ...file, - llmDeliveryPath: resolveUploadLLMDeliveryPath({ - mimeType: file.metadata?.routingMimeType ?? file.type, - endpointConfig, - fileConfig, - endpoint, - endpointProvider: agent.provider, - useResponsesApi, - sttConfigured: isSpeechProviderConfigured(deps.req.config?.speech?.stt), - }), - }; + return { ...file, llmDeliveryPath }; }; const sharedFiles = files.map(resolveDelivery); const compatibleFiles = filterFilesByEndpointRuntimeConfig(deps.req.config, { diff --git a/packages/api/src/agents/files/index.ts b/packages/api/src/agents/files/index.ts index 9293cd7f4ea..0bdab2ea815 100644 --- a/packages/api/src/agents/files/index.ts +++ b/packages/api/src/agents/files/index.ts @@ -3,3 +3,4 @@ export * from './session'; export * from './host'; export * from './binding'; export * from './encode'; +export * from './delivery'; diff --git a/packages/api/src/agents/files/session.spec.ts b/packages/api/src/agents/files/session.spec.ts index ad786df7ccd..9a01735a096 100644 --- a/packages/api/src/agents/files/session.spec.ts +++ b/packages/api/src/agents/files/session.spec.ts @@ -7,6 +7,7 @@ import type { RunFileSessionDeps } from './session'; import type { ServerRequest } from '~/types'; import { createRunFileSession, getAuthorizedRunFileSnapshot } from './session'; import { AgentAttachmentLimitError } from '../attachments'; +import { resolveTurnDeliveryRouting } from './delivery'; import { createRunFileMessageEncoder } from './encode'; function setup( @@ -219,6 +220,10 @@ it.each([{ totalSizeLimit: 1 }, { fileLimit: 1 }])( getAgent: (id) => ({ provider: 'openAI', agentContextAttachments: id === 'writer' ? writerAttachments : [], + deliveryRouting: resolveTurnDeliveryRouting({ + agent: { provider: 'openAI' }, + config: { fileConfig }, + }), }), encodeDocuments, encodeImages: async () => ({ image_urls: [] }), diff --git a/packages/api/src/agents/initialize.ts b/packages/api/src/agents/initialize.ts index 52ab19d2790..62fb2154fdc 100644 --- a/packages/api/src/agents/initialize.ts +++ b/packages/api/src/agents/initialize.ts @@ -21,6 +21,7 @@ import type { TEndpointOption, ReasoningResponseKey, StatefulCodeEnvironment, + TurnDeliveryRouting, ImageDetail, TFile, Agent, @@ -113,6 +114,7 @@ import { applyIntentLabels, sanitizeIntentLabels } from './intent'; import { ContentFilterError } from '../middleware/contentFilter'; import { resolveToolRoleGrants } from '~/tools/rolePermissions'; import { createRequestAgentExecutionContext } from './runtime'; +import { resolveTurnDeliveryRouting } from './files/delivery'; import { filterFilesByEndpointRuntimeConfig } from '~/files'; import { hasActiveFileFieldPolicy } from '~/protection'; import { PARTIAL_RESOLVED_CONVERSATION } from './guard'; @@ -626,6 +628,8 @@ export type InitializedAgent = Agent & { /** Detail level LibreChat encodes image content blocks with, from the agent's * model parameters. Absent when the agent does not configure one. */ imageDetail?: ImageDetail; + /** How this agent receives its attachments this turn, settled once every routing input is final. */ + deliveryRouting: TurnDeliveryRouting; tool_resources?: AgentToolResources; userMCPAuthMap?: Record>; /** Tool map for ToolNode to use when executing tools (required for PTC) */ @@ -1282,6 +1286,88 @@ export async function initializeAgent( const provider = agent.provider; agent.endpoint = provider; + /** Settle the provider and its client options before any attachment is judged. The file + * policy reads the endpoint's own name, but the route each attachment takes also depends on + * the backing client and on the Responses API decision `getOptions` makes, and nothing + * between here and tool loading feeds either. */ + const { getOptions, overrideProvider, customEndpointConfig } = getProviderConfig({ + provider, + appConfig, + }); + if (overrideProvider !== agent.provider) { + agent.provider = overrideProvider; + } + + const finalModelOptions = { + ...modelOptions, + model: agent.model, + }; + + const options: InitializeResultBase = await getOptions({ + runtime: { + appConfig, + user, + requestBody: runtime.requestBody, + }, + endpoint: provider, + model_parameters: finalModelOptions, + db, + }); + + const llmConfig = options.llmConfig as Record; + const webSearchDenied = + hasProviderWebSearch(options.tools, llmConfig) && + !(await resolveWebSearchGrant({ + req: params.req as Request | undefined, + user, + resolve: params.resolveWebSearchGrant, + getRoleByName: db.getRoleByName, + })); + if (webSearchDenied && stripWebSearchPlugin(llmConfig) > 0) { + logger.debug( + `[initializeAgent] Removed the OpenRouter web search plugin; role denies WEB_SEARCH.`, + ); + } + const tokensModel = + agent.provider === EModelEndpoint.azureOpenAI ? agent.model : (llmConfig?.model as string); + const maxOutputTokens = optionalChainWithEmptyCheck( + llmConfig?.maxOutputTokens as number | undefined, + llmConfig?.maxTokens as number | undefined, + 0, + ); + const agentMaxContextTokens = optionalChainWithEmptyCheck( + maxContextTokens, + getModelMaxTokens( + tokensModel ?? '', + providerEndpointMap[overrideProvider as keyof typeof providerEndpointMap], + options.endpointTokenConfig, + ), + DEFAULT_MAX_CONTEXT_TOKENS, + ); + + if ( + agent.endpoint === EModelEndpoint.azureOpenAI && + (llmConfig?.azureOpenAIApiInstanceName as string | undefined) == null + ) { + agent.provider = Providers.OPENAI; + } + + if (options.provider != null) { + agent.provider = options.provider; + } + + const deliveryRouting = resolveTurnDeliveryRouting({ + agent: { + provider: agent.provider, + endpoint: agent.endpoint, + model_parameters: { + useResponsesApi: + typeof llmConfig.useResponsesApi === 'boolean' ? llmConfig.useResponsesApi : undefined, + }, + }, + config: appConfig, + }); + /** Resolve the per-agent Code API route before resource/tool priming. A * stateful agent must perform freshness checks and recovery uploads against * the same isolated deployment its eventual `/exec` request will use. */ @@ -1917,72 +2003,6 @@ export async function initializeAgent( } } - const { getOptions, overrideProvider, customEndpointConfig } = getProviderConfig({ - provider, - appConfig, - }); - if (overrideProvider !== agent.provider) { - agent.provider = overrideProvider; - } - - const finalModelOptions = { - ...modelOptions, - model: agent.model, - }; - - const options: InitializeResultBase = await getOptions({ - runtime: { - appConfig, - user, - requestBody: runtime.requestBody, - }, - endpoint: provider, - model_parameters: finalModelOptions, - db, - }); - - const llmConfig = options.llmConfig as Record; - const webSearchDenied = - hasProviderWebSearch(options.tools, llmConfig) && - !(await resolveWebSearchGrant({ - req: params.req as Request | undefined, - user, - resolve: params.resolveWebSearchGrant, - getRoleByName: db.getRoleByName, - })); - if (webSearchDenied && stripWebSearchPlugin(llmConfig) > 0) { - logger.debug( - `[initializeAgent] Removed the OpenRouter web search plugin; role denies WEB_SEARCH.`, - ); - } - const tokensModel = - agent.provider === EModelEndpoint.azureOpenAI ? agent.model : (llmConfig?.model as string); - const maxOutputTokens = optionalChainWithEmptyCheck( - llmConfig?.maxOutputTokens as number | undefined, - llmConfig?.maxTokens as number | undefined, - 0, - ); - const agentMaxContextTokens = optionalChainWithEmptyCheck( - maxContextTokens, - getModelMaxTokens( - tokensModel ?? '', - providerEndpointMap[overrideProvider as keyof typeof providerEndpointMap], - options.endpointTokenConfig, - ), - DEFAULT_MAX_CONTEXT_TOKENS, - ); - - if ( - agent.endpoint === EModelEndpoint.azureOpenAI && - (llmConfig?.azureOpenAIApiInstanceName as string | undefined) == null - ) { - agent.provider = Providers.OPENAI; - } - - if (options.provider != null) { - agent.provider = options.provider; - } - /** * Unify code-execution tools around `bash_tool` + `read_file` when the * agent explicitly lists `execute_code` in its tools and the admin @@ -2371,6 +2391,7 @@ export async function initializeAgent( azureOptions: options.azureOptions, resendFiles, imageDetail, + deliveryRouting, toolRegistry, mcpAvailableTools, requestScopedConnections, diff --git a/packages/data-provider/src/resolve-llm-delivery-path.spec.ts b/packages/data-provider/src/resolve-llm-delivery-path.spec.ts index b62c3a26f90..2a9b254bf59 100644 --- a/packages/data-provider/src/resolve-llm-delivery-path.spec.ts +++ b/packages/data-provider/src/resolve-llm-delivery-path.spec.ts @@ -1,3 +1,4 @@ +import type { TurnDeliveryRouting } from './resolve-llm-delivery-path'; import type { TDefaultLLMDeliveryPathConfig } from './file-config'; import type { TEndpoint } from './config'; import { @@ -5,6 +6,7 @@ import { canToolResourceConsume, resolveUploadDestination, getCustomEndpointProvider, + resolveTurnLLMDeliveryPath, resolveDefaultLLMDeliveryPath, resolveUploadLLMDeliveryPath, SYSTEM_LLM_DELIVERY_DEFAULTS, @@ -640,6 +642,75 @@ describe('resolveUploadDestination', () => { }); }); +describe('resolveTurnLLMDeliveryPath', () => { + const fileConfig = mergeFileConfig({ + endpoints: { + openAI: { defaultLLMDeliveryPath: { overrides: { 'application/pdf': 'text' } } }, + MyClaude: { supportedMimeTypes: ['video/mp4'] }, + }, + }); + const routing = ( + endpoint: string, + extra?: Partial, + ): TurnDeliveryRouting => ({ + fileConfig, + endpointConfig: getEndpointFileConfig({ fileConfig, endpoint }), + endpoint, + sttConfigured: false, + ...extra, + }); + const pdf = { + type: 'application/pdf', + llmDeliveryPath: 'provider', + metadata: { destinationChosen: false }, + }; + + it('resolves an inferred route again for the endpoint running the turn', () => { + expect(resolveTurnLLMDeliveryPath(routing('openAI'), pdf)).toBe('text'); + }); + + it('keeps a destination the user chose', () => { + expect( + resolveTurnLLMDeliveryPath(routing('openAI'), { + ...pdf, + metadata: { destinationChosen: true }, + }), + ).toBe('provider'); + }); + + it('leaves a record predating routing to its legacy handling', () => { + expect(resolveTurnLLMDeliveryPath(routing('openAI'), { type: 'application/pdf' })).toBe( + undefined, + ); + expect( + resolveTurnLLMDeliveryPath(routing('openAI'), { ...pdf, llmDeliveryPath: 'legacy' }), + ).toBe(undefined); + }); + + it('takes the stored route as the route when no turn routing exists', () => { + expect(resolveTurnLLMDeliveryPath(undefined, pdf)).toBe('provider'); + }); + + it('routes by the type the upload was resolved against after conversion', () => { + const converted = { + type: 'image/webp', + llmDeliveryPath: 'provider', + metadata: { routingMimeType: 'application/pdf', destinationChosen: false }, + }; + + expect(resolveTurnLLMDeliveryPath(routing('openAI'), converted)).toBe('text'); + }); + + it('honors a custom endpoint media opt-in only under an OpenAI-format dialect', () => { + const video = { type: 'video/mp4', llmDeliveryPath: 'none', metadata: {} }; + + expect(resolveTurnLLMDeliveryPath(routing('MyClaude'), video)).toBe('provider'); + expect( + resolveTurnLLMDeliveryPath(routing('MyClaude', { endpointProvider: 'anthropic' }), video), + ).toBe('none'); + }); +}); + describe('getCustomEndpointProvider', () => { const custom = [ { name: 'My Claude', provider: 'anthropic' }, diff --git a/packages/data-provider/src/resolve-llm-delivery-path.ts b/packages/data-provider/src/resolve-llm-delivery-path.ts index bccf7f566cf..eb06d918868 100644 --- a/packages/data-provider/src/resolve-llm-delivery-path.ts +++ b/packages/data-provider/src/resolve-llm-delivery-path.ts @@ -325,6 +325,57 @@ export function resolveUploadLLMDeliveryPath({ }); } +/** + * The inputs that route every attachment for the agent running a turn. Initialization + * settles them once, after the provider swap and the Responses API decision, and every + * reader of a turn route consumes this value rather than deriving one from the agent. + */ +export interface TurnDeliveryRouting { + fileConfig: FileConfig; + endpointConfig: EndpointFileConfig; + /** The endpoint the file policy is configured under: a custom endpoint's own name, not + * the client family initialization runs it as. */ + endpoint: string; + /** The dialect a custom endpoint declares, which decides whether it receives OpenAI-format + * media; undefined for a built-in or OpenAI-compatible endpoint. */ + endpointProvider?: string; + useResponsesApi?: boolean; + sttConfigured: boolean; +} + +/** The fields of a stored attachment record that decide its route on a turn. */ +export interface TurnDeliveryFile { + type?: string; + /** Stored as an upload-time inference, so any string may be read back. */ + llmDeliveryPath?: string | null; + metadata?: { routingMimeType?: string; destinationChosen?: boolean } | null; +} + +const isLLMDeliveryPath = (value: unknown): value is TDefaultLLMDeliveryPath => + value === 'provider' || value === 'text' || value === 'none'; + +/** + * The route a stored attachment takes on the turn `routing` describes. + * + * The stored route records what upload inferred from the endpoint it saw, and this turn may + * run somewhere else, so an inferred route is resolved again for the endpoint handling the + * turn. A destination the user chose is theirs and stands, a record predating routing keeps + * its legacy handling, and without a turn routing the stored route is the route. + */ +export function resolveTurnLLMDeliveryPath( + routing: TurnDeliveryRouting | undefined, + file: TurnDeliveryFile, +): TDefaultLLMDeliveryPath | undefined { + const stored = isLLMDeliveryPath(file.llmDeliveryPath) ? file.llmDeliveryPath : undefined; + if (routing == null || stored == null || file.metadata?.destinationChosen === true) { + return stored; + } + return resolveUploadLLMDeliveryPath({ + mimeType: file.metadata?.routingMimeType ?? file.type ?? '', + ...routing, + }); +} + /** * Whether a file tool can do anything with this type. `file_search` indexes extracted * text, so it needs a type some step can turn into text and cannot use media, whose