You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Two related changes that together let a deployment treat extracted text as the source of truth and stop retaining original binaries.
1. Embed from stored extracted text when it is available
provisionToVectorDB in packages/api/src/files/provision/service.ts currently fetches a download stream from storage, pipes it to a temp file, and hands that temp file to uploadVectors:
stream=(awaitgetStorageStream(file,req,signal))??undefined;if(!stream){thrownewError(`Cannot provision file "${file.filename}" to vector DB: storage source "${file.source}" does not support download streams`,);}awaitpipeline(stream,fs.createWriteStream(tmpPath),{ signal });
There is no branch that uses the already-extracted file.text, even when it is populated.
Consequences:
The same document is parsed twice./text runs at upload time to populate file.text; /embed re-parses the identical file at tool-execution time. On a custom or resource-constrained RAG API this doubles extraction cost for every searched document, and the two endpoints can be configured differently, so the text the model reads and the text that gets embedded may not match.
Embedding is permanently coupled to binary retention. Because provisioning is deferred until file_search is first called (🚚 feat: Route Attachments and Provision Tools Lazily #15763), the original must remain in storage indefinitely — potentially forever, if the file is never searched.
Request: when file.text is present and non-empty, allow embedding from that text rather than re-fetching and re-parsing the original. A config flag would be fine if chunking quality on structured documents (PDF page boundaries, tables) is a concern for the default path.
2. Opt-in discarding of originals for text/OCR-derived uploads
Raised in #15941 and not covered by the preview fix in #15946, which deliberately keeps Download pointed at the original file.
For records with llmDeliveryPath: 'text' and populated text, the original binary is never used to build the prompt. It is retained only for Download. On deployments that process large volumes of documents this is significant storage, and for some deployments retaining the originals is undesirable on data-handling grounds rather than disk grounds.
Request: a fileConfig option such as discardOriginalAfterExtraction, which for text-delivered and OCR-derived uploads stores the extracted text, deletes the stored object, and records source: text so the existing download branch serves the extracted text as .txt — the path already implemented for files with no backing object.
With (1) in place this becomes safe for file_search uploads too, not just text-delivered ones.
More details
Why
Removes duplicate extraction work per document.
Guarantees that the embedded text and the prompt text are the same text.
Lets deployments decide whether originals are retained, rather than requiring indefinite retention for a Download-only use case.
Makes text extraction the single source of truth, which fits the direction of the unified upload work.
What features would you like to see added?
Two related changes that together let a deployment treat extracted text as the source of truth and stop retaining original binaries.
1. Embed from stored extracted text when it is available
provisionToVectorDBinpackages/api/src/files/provision/service.tscurrently fetches a download stream from storage, pipes it to a temp file, and hands that temp file touploadVectors:There is no branch that uses the already-extracted
file.text, even when it is populated.Consequences:
/textruns at upload time to populatefile.text;/embedre-parses the identical file at tool-execution time. On a custom or resource-constrained RAG API this doubles extraction cost for every searched document, and the two endpoints can be configured differently, so the text the model reads and the text that gets embedded may not match.file_searchis first called (🚚 feat: Route Attachments and Provision Tools Lazily #15763), the original must remain in storage indefinitely — potentially forever, if the file is never searched.Request: when
file.textis present and non-empty, allow embedding from that text rather than re-fetching and re-parsing the original. A config flag would be fine if chunking quality on structured documents (PDF page boundaries, tables) is a concern for the default path.2. Opt-in discarding of originals for text/OCR-derived uploads
Raised in #15941 and not covered by the preview fix in #15946, which deliberately keeps Download pointed at the original file.
For records with
llmDeliveryPath: 'text'and populatedtext, the original binary is never used to build the prompt. It is retained only for Download. On deployments that process large volumes of documents this is significant storage, and for some deployments retaining the originals is undesirable on data-handling grounds rather than disk grounds.Request: a
fileConfigoption such asdiscardOriginalAfterExtraction, which for text-delivered and OCR-derived uploads stores the extracted text, deletes the stored object, and recordssource: textso the existing download branch serves the extracted text as.txt— the path already implemented for files with no backing object.With (1) in place this becomes safe for file_search uploads too, not just text-delivered ones.
More details
Why
Related
Which components are impacted by your request?
No response
Pictures
No response
Code of Conduct