Skip to content

[Enhancement]: Embed from stored extracted text and allow discarding originals for text/OCR uploads #15957

Description

@dvejsada

What features would you like to see added?

Two related changes that together let a deployment treat extracted text as the source of truth and stop retaining original binaries.

1. Embed from stored extracted text when it is available

provisionToVectorDB in packages/api/src/files/provision/service.ts currently fetches a download stream from storage, pipes it to a temp file, and hands that temp file to uploadVectors:

stream = (await getStorageStream(file, req, signal)) ?? undefined;
if (!stream) {
  throw new Error(
    `Cannot provision file "${file.filename}" to vector DB: storage source "${file.source}" does not support download streams`,
  );
}
await pipeline(stream, fs.createWriteStream(tmpPath), { signal });

There is no branch that uses the already-extracted file.text, even when it is populated.

Consequences:

  • The same document is parsed twice. /text runs at upload time to populate file.text; /embed re-parses the identical file at tool-execution time. On a custom or resource-constrained RAG API this doubles extraction cost for every searched document, and the two endpoints can be configured differently, so the text the model reads and the text that gets embedded may not match.
  • Embedding is permanently coupled to binary retention. Because provisioning is deferred until file_search is first called (🚚 feat: Route Attachments and Provision Tools Lazily #15763), the original must remain in storage indefinitely — potentially forever, if the file is never searched.

Request: when file.text is present and non-empty, allow embedding from that text rather than re-fetching and re-parsing the original. A config flag would be fine if chunking quality on structured documents (PDF page boundaries, tables) is a concern for the default path.

2. Opt-in discarding of originals for text/OCR-derived uploads

Raised in #15941 and not covered by the preview fix in #15946, which deliberately keeps Download pointed at the original file.

For records with llmDeliveryPath: 'text' and populated text, the original binary is never used to build the prompt. It is retained only for Download. On deployments that process large volumes of documents this is significant storage, and for some deployments retaining the originals is undesirable on data-handling grounds rather than disk grounds.

Request: a fileConfig option such as discardOriginalAfterExtraction, which for text-delivered and OCR-derived uploads stores the extracted text, deletes the stored object, and records source: text so the existing download branch serves the extracted text as .txt — the path already implemented for files with no backing object.

With (1) in place this becomes safe for file_search uploads too, not just text-delivered ones.

More details

Why

  • Removes duplicate extraction work per document.
  • Guarantees that the embedded text and the prompt text are the same text.
  • Lets deployments decide whether originals are retained, rather than requiring indefinite retention for a Download-only use case.
  • Makes text extraction the single source of truth, which fits the direction of the unified upload work.

Related

Which components are impacted by your request?

No response

Pictures

No response

Code of Conduct

  • I agree to follow this project's Code of Conduct

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    ✨ enhancementNew feature or request🗺️ Content And Filescodegraph: the taxonomy area this belongs to (classifier, confidence ≥ 0.9)🗺️ File Storagecodegraph: the taxonomy area this belongs to (classifier, confidence ≥ 0.9)🗺️ Server Corecodegraph: the taxonomy area this belongs to (classifier, confidence ≥ 0.9)

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions