Skip to content

📎 fix: Keep Provider Uploads the Provider Cannot Read Inline Out of File Parts - #16473

Open
mihidumh wants to merge 6 commits into
LibreChat-AI:devfrom
mihidumh:fix/provider-doc-mime-optin
Open

mihidumh wants to merge 6 commits into
LibreChat-AI:devfrom
mihidumh:fix/provider-doc-mime-optin

Conversation

@mihidumh

@mihidumh mihidumh commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Summary

With "Upload to Provider", some files on the inherited supportedMimeTypes list went to the provider as file parts that the provider always rejects. The 400 then recurred on every later turn, because the file is re-sent from history. Azure OpenAI rejects application/sql, application/x-sh, application/xml, zip and octet-stream. Gemini rejects docx, xlsx and every textual application/* type (JSON, YAML, XML, SQL, TypeScript). The measurements are in the issue.

The inherited list stays the provider opt-in, as #13550 made it for OpenAI (CSV and XLSX to the Responses API keep working, and its test passes unchanged). Only the types that can never go inline change:

  • Archives (zip, x-zip-compressed, tar, gzip, epub, octet-stream) go to the provider only when the endpoint lists them in its own supportedMimeTypes. Upload to Code Environment and File Search are not touched.
  • Gemini (Google/Vertex, or an OpenAI-like provider whose model name contains gemini, the same detection 📄 fix: Send Textual Documents to Claude Through Gateways as Text #16055 uses for Claude) gets PDF and textual types. Behind a gateway it also gets the types the endpoint lists. On native Google/Vertex the upload allowlist does not change what the API accepts, so there a listed textual type still goes as text and a listed Office document is still left out.
  • Textual types the provider rejects as a file part go as a text part, File: "<name>"\n\n<contents>, unless the endpoint lists them: every textual type (isNativelyReadableText) on OpenAI-like Chat Completions, application/sql, x-sh and xml on the OpenAI-like Responses API too, and every textual application/* type on Gemini. OpenAI Chat Completions accepts only PDF as file.file_data (400 "Expected a base64-encoded data URL with an application/pdf MIME type" for text/plain, CSV, HTML and JSON, measured on api.openai.com with v0.8.8), while Azure OpenAI accepts those types there; a text part is read by both. text/* stays inline media on Gemini, and the Responses API keeps input_file for textual types, which OpenAI documents.

A file the model would not receive is no longer dropped silently. The document encoder returns the files it skipped in omitted, and BaseClient.processAttachments adds the provider-bound files that no encoder takes. On the current turn, the send fails with AgentAttachmentUnsupportedError (code AGENT_ATTACHMENT_UNSUPPORTED, 415), which names each file and tells the user to remove it or to upload it as text or to the code environment. It uses the same path to the user as AgentAttachmentLimitError, before the user message is saved. On history replay the turn continues: the file is left out of message_file_map, the message gets a text note that names it, and the server logs it. A child-run share with such a file is rejected in validate(), before any bytes are read. For Bedrock, a skipped non-Bedrock type is now reported too, so it rejects the turn and is not dropped.

Decoded text is bounded. The text-part fallback and the plain-text source of a native Anthropic document use fileTokenLimit per file (processTextWithTokenLimit), and fileContextCharLimit across the request (the current turn, history replay and child runs). The text is cut, not rejected, so replay cannot fail on it. This budget is separate from the extracted-text count in assertAgentAttachmentLimits.

Fixes #16472

How it works

  • isProviderDocumentCandidate (packages/api/src/files/encode/utils.ts) replaces the inline checkType test in BaseClient.processAttachments and in the child-run encoder. It uses isExplicitMimeConfig, the same inherited-list check as audio/video delivery.
  • encodeAndFormatDocuments passes isConfiguredProviderMediaType (explicit opt-in) to filterProviderDocumentFiles and formatDocumentBlock. On native Google/Vertex it passes no opt-in.
  • Omissions: DocumentResult.omitted (OmittedAttachment[]), AgentAttachmentUnsupportedError in packages/api/src/agents/attachments.ts, and processAttachments(message, attachments, fileConsumers, { historical }).
  • A text part works with the Responses API too: LangChain maps { type: 'text' } to input_text (convertMessagesToResponsesInput). I checked this with @langchain/openai 1.5.5.

An admin whose gateway accepts one of these types as a file lists it in the endpoint's supportedMimeTypes, and the file part comes back.

The Gemini and Claude checks read the model name. A gateway alias that hides the family (for example fast → vertex_ai/gemini-flash-lite-latest) is not detected, so that endpoint must list its types in supportedMimeTypes explicitly.

Change Type

  • Bug fix (non-breaking change which fixes an issue)

Testing

  • packages/api: jest src/files/encode src/agents/files gives 393 passed. jest src/files src/agents src/middleware gives 6,367 passed. The only failures are in two suites that need things this machine does not have: a Redis server (concurrency.cache_integration) and a shell profile without pyenv (hooks/executor). New cases: archives under the inherited list, explicit opt-in, SQL as text on OpenAI chat completions, Responses API, OpenRouter, Google and Vertex, XML as text on OpenAI, Gemini behind a gateway, Office documents on Gemini, PDF, CSV unchanged on Gemini, and (round 3) text/plain, CSV, HTML, JSON, YAML and TypeScript as text parts on OpenAI chat completions, the same types as input_file on the responses API, and text/plain and JSON kept as file parts when the endpoint lists them.
  • api: jest app/clients/specs/BaseClient.test.js server/controllers/agents/client.test.js gives 424 passed, including 📎 fix: Preserve Provider Document Uploads #13550's CSV/XLSX test. Review round 2 added tests for the typed error on the current turn (BaseClient, and a zip through the real AgentClient.processAttachments), the omission note on replay, archive-only and mixed child shares, a 2 MB SQL file under fileTokenLimit: 8 / fileContextCharLimit: 100, and listed SQL/DOCX/XLSX on native google and vertexai. The new BaseClient tests fail on dev without the fix. One of them replays an archive attached on an earlier message through addPreviousAttachments, the path that made a conversation fail on every later turn: the archive is skipped and the PDF from the same message is still sent.
  • I did not run the Playwright e2e suite. Its provider uploads (chat.spec.ts, unified-upload.spec.ts) use text/csv, text/plain and text/markdown on mock-model-* models; they assert the stored delivery path, not the part shape, which is now a text part on chat completions.
  • One existing test changed: "should format XLSX for Google/VertexAI as media block" now expects the XLSX to be omitted even when the endpoint lists it, because native Gemini rejects inline XLSX. The 📄 fix: Send Textual Documents to Claude Through Gateways as Text #16055 test "should keep textual documents as file parts for non-Claude OpenAI models" now uses a GPT model name instead of a Gemini one, and since round 3 it expects a text part on chat completions, with a sibling that keeps input_file on the responses API.
  • tsc --noEmit for packages/api, ESLint and Prettier are clean.
  • Known limits, from before this PR: the upload route still returns 200 and the rejection happens at send time. agents/steering/media.ts does not read omitted yet.
  • The provider behaviour was measured through a LiteLLM proxy (see the issue): Azure gpt-5-mini and Vertex gemini-flash-lite-latest.

Test Configuration:

Custom OpenAI-compatible endpoints → LiteLLM → Azure OpenAI and Gemini on Vertex AI, legacyFileUploadUX: true; v0.8.8-rc3/rc4 in production, patch against dev c8c5478.

Checklist

  • My code adheres to this project's style guidelines
  • I have performed a self-review of my own code
  • I have commented in any complex areas of my code
  • My changes do not introduce new warnings
  • I have written tests demonstrating that my changes are effective or that my feature works
  • Local unit tests pass with my changes

🤖 Generated with Claude Code

…ile Parts

The inherited supportedMimeTypes list routes files to "Upload to Provider"
(LibreChat-AI#13550), but some types on it can never go inline:

- Archives (zip, tar, gzip, epub, octet-stream) go to the provider only
  when the endpoint lists them itself, in BaseClient and child runs.
- Gemini, native or behind an OpenAI-compatible gateway (detected by model
  name, as LibreChat-AI#16055 does for Claude), gets PDF and textual types plus types
  the endpoint lists; Office documents are skipped like the Anthropic
  filter from LibreChat-AI#14535. Textual application/* types (JSON, YAML, XML, SQL)
  go to Gemini as a text part; text/* stays inline.
- application/sql, x-sh and xml go to other OpenAI-like providers as a
  text part unless the endpoint lists them: Azure OpenAI rejects them as a
  file part with a 400 that recurs on every later turn.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Copilot AI balanced review requested due to automatic review settings September 29, 2026 00:14

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Gemini handling still drops an inherited textual MIME type, and one warning recommends an ineffective configuration fix.

Review effort: Balanced
Findings: 1 Medium severity · 1 Low severity

Open (2)
What changed in this PR

Prevents unsupported provider-uploaded files from producing recurring provider errors.

Changes:

  • Filters inherited archive and Gemini-incompatible document types.
  • Converts unsupported textual file parts into text blocks.
  • Adds routing and encoder regression tests.
File Description
packages/​api/​src/​files/​encode/​utils.ts Adds provider document eligibility logic.
packages/​api/​src/​files/​encode/​processAttachments.spec.ts Tests inherited and explicit MIME routing.
packages/​api/​src/​files/​encode/​document.ts Adds Gemini filtering and text-part formatting.
packages/​api/​src/​files/​encode/​document.spec.ts Tests provider-specific document handling.
packages/​api/​src/​agents/​files/​encode.ts Applies filtering to child-run encoding.
packages/​api/​src/​agents/​files/​encode.spec.ts Tests child-run archive routing.
api/​app/​clients/​BaseClient.js Filters and warns about dropped attachments.
api/​app/​clients/​specs/​BaseClient.test.js Tests BaseClient filtering and warnings.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +211 to +213
} else if (usesGeminiDocumentCapabilities(provider, model)) {
label = 'Gemini';
isSupported = (file) => isAnthropicDocumentType(file.type) || isOptedIn(file.type ?? '');

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 5ec8bb3. The Gemini filter and sendsAsTextWithoutOptIn now use isNativelyReadableText, the same classification that routes the file to the provider. So an inherited application/vnd.coffeescript (and application/x-yaml) goes to Gemini as a text part and is no longer dropped. New spec covers it for native Google and for Gemini behind a gateway. The Claude filter from #14535 still uses isAnthropicDocumentType. I left it as it is, because this PR does not change it.

Comment thread api/app/clients/BaseClient.js Outdated
Comment on lines +1855 to +1857
logger.warn(
`[BaseClient] Not sending "${file.filename}" (${file.type}) to the provider: list this type in the endpoint's own supportedMimeTypes to send it`,
);

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed, fixed in 5ec8bb3. The warning now only says the type is not an inline document for this endpoint, and gives no configuration advice. You are right: for Claude and Bedrock, an explicitly listed archive is still filtered out.

The Gemini document filter and its text-part decision used the Anthropic
text list, which lacks application/vnd.coffeescript and application/x-yaml.
The router sends those to the provider as natively readable text, so the
filter dropped an inherited CoffeeScript file. Both now use
isNativelyReadableText, the classification that routes the file.

The BaseClient skip warning no longer tells admins to list the type: for
Claude and Bedrock an explicitly listed archive is still filtered out.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A conversation breaks when an archive from an earlier message goes out
as a file part on every later turn. The replay path (addPreviousAttachments
→ processAttachments) now has a test: the archive is skipped with a warning,
the PDF from the same message is still sent, and message_file_map holds
only the PDF.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@lia-by-librechat lia-by-librechat Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed a3155b430d89fc097d1e885749bcf1ec58b796fc again, including the full base-to-head diff, upload admission, current-turn delivery, history replay, child-run encoding, and the existing inline review threads. The added archive-history regression test is useful, and the CoffeeScript classification and warning-wording findings from the earlier review are addressed.

I am requesting changes for the three inline findings:

  1. Make unsupported attachment omissions observable instead of reporting an accepted attachment while sending no file content.
  2. Apply the existing text budgets to the new raw-file-to-text fallback, including aggregate turn/child/replay accounting.
  3. Keep native Google/Vertex capability constraints separate from upload allowlists. This configuration finding is narrower than a blanket prohibition on overrides: intentional overrides for custom gateways should remain supported.

Verification: mapped consumers with the code graph at dev commit e9d1d330b1e1429b01f49b3a6c77764e2049f47d, then inspected the actual PR source. A dependency-light probe executed the actual document encoder, BaseClient methods, MIME classifiers, and attachment-stat collector, with storage/config lookup/SDK-validation adapters supplied. It reproduced the retained-but-empty Office attachment, archive omission, unbounded SQL text, and configured native-provider media output. These were source-level checks, not live provider or end-to-end tests. git diff --check passed. The focused document/child-run/BaseClient/AgentClient Jest suites and packages/api typecheck could not start because the shared dependency install currently lacks the local Jest and TypeScript executables. GitHub reports no CI checks for this head.

These are changes requested before merging this PR, not a request to hold the release for it.

Comment on lines +233 to +235
if (skipped.length) {
console.warn(
`Skipping attachment(s) unsupported by Claude document input: ${skipped.join(', ')}`,
`Skipping attachment(s) unsupported by ${label} document input: ${skipped.join(', ')}`,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Surface rejected/omitted attachments to the user

For a provider-chosen DOCX on Google with no explicit MIME list, BaseClient.processAttachments adds the file to allFiles before this filter drops it. The actual methods return [report.docx] while message.documents is undefined and storage is never read. The upload route still reports success, and current-turn persistence can retain the attachment even though the model receives none of its content. The inherited archive path similarly ends with only a server warning, and an archive-only child share can return no file message at all.

Please reject unsupported new/current-turn provider attachments through a user-visible, typed error, or return an omission outcome that the UI can display with localized copy. Historical replay should still recover rather than throwing on every later turn, but its omission must be explicit. Preserve historical references and tool-only uploads. Add coverage for the user-visible current-turn result, restored-history omission, and an archive-only child share; a warning assertion alone does not prove that the missing input is observable.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 8c83da1, with the typed-error route.

  • encodeAndFormatDocuments now returns the files it skipped in omitted (reason unsupported_type or text_limit). BaseClient.processAttachments collects them together with the provider-bound files that no encoder takes (the archive case).
  • Current turn: processAttachments throws AgentAttachmentUnsupportedError (code AGENT_ATTACHMENT_UNSUPPORTED, 415). The message names each file, for example This model cannot read "report.docx" (…). Remove it, or upload it as text or to the code environment, and try again. isAgentAttachmentLimitError accepts it, and the code is in FATAL_AGENT_INITIALIZATION_CODES. So it takes the same path to the user as the existing attachment-limit errors, and the turn fails before the user message is saved.
  • History replay: the turn continues. The file is left out of message_file_map, the message gets a text part that names the file and the reason, and logger.warn records the omission with the message id.
  • Child share: createRunFileMessageEncoder rejects a provider-bound file that no encoder takes in validate(), before it reads bytes. An archive-only share now fails with the same error, and it no longer returns no message. An omitted list from the document encoder rejects the share too.

Tests: current-turn archive and encoder omission (BaseClient), the same archive through the real AgentClient.processAttachments, replay of an archive and of an encoder-omitted DOCX, and archive-only and mixed child shares.

Known limits, both from before this PR: the upload route still returns 200, and the rejection happens at send time, as with AgentAttachmentLimitError. agents/steering/media.ts also calls the document encoder and does not read omitted yet. One change in behaviour: a Bedrock skip is now reported, so a non-Bedrock type on the document path rejects the turn instead of being dropped.

Comment on lines +103 to +106
function formatTextDocumentBlock(filename: string, content: string): DocumentBlock {
return {
type: 'text',
text: `File: "${filename}"\n\n${Buffer.from(content, 'base64').toString('utf8')}`,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Enforce text budgets before emitting the fallback block

This decodes the complete file directly into a text block without applying fileTokenLimit or accounting for fileContextCharLimit. A legacy provider-chosen upload normally has no file.text, so collectAgentAttachmentStats counts zero extracted-text characters before this conversion. A source-level run with fileTokenLimit: 8, fileContextCharLimit: 100, and a roughly 2 MB SQL file emitted 2,000,029 characters and counted zero extracted-text characters. The byte-size limit is not a substitute for the configured text limits.

Please reuse the existing bounded-text processing and account for the resulting text in the turn's aggregate limits, including child delivery and history replay, before constructing/sending the block. Add a large-SQL regression test with deliberately small token/character budgets. This need not add another storage read.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 8c83da1. The text-part fallback and the plain-text source of a native Anthropic document now go through limitTextBlock:

  • fileTokenLimit per file (req.body.fileTokenLimit ?? fileConfig.fileTokenLimit), through processTextWithTokenLimit with countTokens, the same call that extractFileContext uses.
  • fileContextCharLimit per request. A WeakMap keyed by req holds the decoded characters already sent, so the limit covers the current turn, history replay and child runs of one request together. The text is cut to the remaining characters before tokens are counted, so a 2 MB file is never tokenized whole. A file that gets no budget is reported as text_limit (current turn: typed error; replay: note and omission).

Your case is now a test: a 2 MB SQL file with fileTokenLimit: 8 and fileContextCharLimit: 100 gives one text block of ≤ 100 characters. A second test shows that the budget is shared across two encoder calls on one request.

Two choices to call out. First, this ledger is separate from extractedTextChars in assertAgentAttachmentLimits. Decoded fallback text and extracted text are each bounded by fileContextCharLimit, so their sum can reach twice the limit. Second, the fallback cuts the text to fit, where the admission check throws. A throw here would also fire on history replay and fail every later turn, which is the failure this PR removes. If you prefer one shared counter, I can seed the ledger from the admission stats.

Comment on lines +215 to +218
isSupported = (file) =>
file.type === 'application/pdf' ||
isNativelyReadableText(file.type ?? '') ||
isOptedIn(file.type ?? '');

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Do not let native Google/Vertex upload allowlists disable safe encoding

The explicit opt-in escape hatch also applies to native Google/Vertex, not just an OpenAI-compatible gateway whose behavior an operator may intentionally override. With provider: vertexai and DOCX in that provider's allowlist, this admits the Office file and the actual encoder emits { type: 'media', mimeType: '<DOCX MIME>', ... }. With native Google and application/sql listed, optedIn also disables the text fallback at lines 147-148 and emits SQL as inline media. Those are the MIME-bearing payloads reported as rejected in #16472; allowing an upload does not change the native provider's capabilities.

Please retain the configurable gateway override, but make the known native Google/Vertex path convert these textual types to text and reject or otherwise handle unsupported Office content independently of the upload allowlist. Add explicit-allowlist SQL and Office cases for both native providers, while preserving gateway opt-in tests.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 8c83da1. On native Google/Vertex, isOptedIn is now always false, because the upload allowlist does not change what the native API accepts. As a result:

  • A textual application/* type (SQL, JSON, YAML, XML) always goes as a text part, even when the endpoint lists it.
  • Office files (DOCX, XLSX) are always omitted, and the omission is reported through the P1 path.

The opt-in is unchanged for Gemini behind an OpenAI-compatible gateway, where the operator controls what the gateway forwards. New tests: a listed XLSX and DOCX on both google and vertexai (omitted, no storage read), a listed SQL on both (text part, not media), and a listed XLSX for gemini-* through a gateway (file part). The existing test that expected a listed XLSX as media on native Google now expects the omission.

Addresses the three findings of the review on a3155b4.

Unsupported attachments are now visible. The document encoder returns the
files it skipped in `omitted` (reason `unsupported_type` or `text_limit`),
and BaseClient collects them together with the provider-bound files that no
encoder takes. On the current turn, processAttachments throws
AgentAttachmentUnsupportedError (code AGENT_ATTACHMENT_UNSUPPORTED, 415).
The error message names each file, it is an attachment error to
isAgentAttachmentLimitError, and the turn fails before the user message is
saved. On history replay the turn continues: the file is left out of
message_file_map, the message carries a text note that names it, and the
server logs the omission with the message id. The child-run encoder
rejects a share with a provider-bound file that it cannot encode, in
validate() and before it reads any bytes, so an archive-only share is no
longer answered with no file message.

Decoded file text is bounded. The text-part fallback and the plain-text
source of a native Anthropic document now go through
processTextWithTokenLimit with fileTokenLimit. fileContextCharLimit is
tracked per request, so it covers the current turn, history replay and
child runs together. The text is cut to the remaining characters before
tokens are counted, and a file that gets no budget is reported as
`text_limit`.

A native Google/Vertex upload allowlist no longer turns off safe encoding.
Listing a type for the native API does not change what that API accepts.
So on native Google/Vertex, textual application types always go as text,
and Office files are omitted whatever supportedMimeTypes says. The opt-in
still works for Gemini behind an OpenAI-compatible gateway.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@danny-avila danny-avila added the 🗺️ Content And Files codegraph: the taxonomy area this belongs to (classifier, confidence ≥ 0.9) label Sep 30, 2026
@codegraph-librechat codegraph-librechat Bot added the 🗺️ File Storage codegraph: the taxonomy area this belongs to (classifier, confidence ≥ 0.9) label Oct 1, 2026
The "Static checks" job failed on import order in attachments.ts and
document.ts. This is the output of `npm run sort-imports` for the two
files. No code change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@tommctech

Copy link
Copy Markdown

Testing against dev (v0.8.8-rc3, digest sha256:1eefa6a4c2b8…) we hit this
400-on-every-later-turn shape from a route I think sits outside this PR's scope: a file
produced by the code interpreter, on a genuine OpenAI model.
Two findings, the second of
which touches a deliberate choice in this diff.

1. Code-interpreter output reaches formatDocumentBlock with a natively-readable type

Repro: fresh conversation, agent with code execution enabled, ask for a small CSV, then
send any follow-up message.

The generated file is stored with context: "execute_code":

file_id   c3a89c96-…      type: text/csv      bytes: 252
context   execute_code
text      len=3889        textFormat: "html"

The generation turn succeeds. The next turn 400s, and keeps 400ing:

POST https://api.openai.com/v1/chat/completions -> 400 in 146ms (openai-processing-ms: 29)
model: gpt-5.4-mini — "error; not retryable"

resolveDefaultLLMDeliveryPath('text/csv') returns text, so an uploaded CSV never
reaches the document path — consistent with this issue listing Office, archive and textual
application/* types rather than text/*. A generated attachment appears not to go
through that resolution, so text/csv arrives at formatDocumentBlock and is encoded as a
file part.

2. For a genuine OpenAI model nothing filters it, and text/* is intentionally kept as a file part

In this diff filterProviderDocumentFiles branches on Bedrock / Claude / Gemini and
otherwise returns { processable: files, skipped: [] }. sendsAsTextWithoutOptIn limits
the non-Gemini text route to application/sql, application/x-sh and application/xml,
with the stated assumption:

text/*, JSON, YAML and TypeScript stay file parts.

For OpenAI Chat Completions I don't believe text/* can stay a file part. From
OpenAI's own file-inputs guide:

Chat Completions accepts only PDF files as file_data and file_id

and for anything else it directs you to "read the file in your application and send its
contents as a text content part." So data:text/csv;base64,… inside a file part is
rejected on arrival, which matches both the 29 ms rejection above and the Azure
"Invalid file data" behaviour already noted in this thread. The Responses API accepts a
much broader set; Chat Completions does not.

Minor, separate observation

The stored text for that 252-byte CSV is 3,889 characters of <!DOCTYPE html>… with
textFormat: "html" — a preview document rather than the file's own text. Any path that
falls back to stored text for a generated file would send markup instead of the CSV.


Happy to open this as its own issue if that's cleaner than widening #16473. Worth saying
that the recovery-from-history work in this PR already addresses the worst of it for us:
as it stands the thread stays dead on every subsequent turn, with no way for the user to
recover it.

@tommctech

Copy link
Copy Markdown

Follow-up to my earlier comment, with one result that narrows it: on the same instance
and build (v0.8.8-rc3, digest sha256:1eefa6a4c2b8…), the generated-file case is
OpenAI-specific. An Anthropic-backed agent passes.

Same repro, same prompt, only the agent's provider changed:

provider / model generated file follow-up turn
openAI / gpt-5.4-mini text/csv, 252 B 400, conversation dead
anthropic / claude-sonnet-5 text/csv, 60 B OK

Provider taken from the stored agent record rather than from the model's own account of
itself, and the Anthropic run genuinely produced a file — the assistant message carries
attachments=text/csv:acceptance.csv:60 and the following turn stored a normal text
block with no error.

The asymmetry looks like it sits in formatDocumentBlock. The Anthropic branch runs
getAnthropicDocumentSource, which turns an isAnthropicTextDocumentType into a
{type: "text", media_type: "text/plain"} source. The OpenAI branch has no equivalent
step and emits {type: "file", file: {file_data: "data:text/csv;base64,…"}}, which Chat
Completions rejects for anything but PDF. So the conversion that makes this safe already
exists on one path and not the other.

Separately, and possibly useful for scoping this PR: on current dev I can no longer
reproduce the problem through the upload path at all. An automated matrix of nine
types — png, pdf, txt, csv, html, json, docx, xlsx, zip — attaches each one, sends a
second turn, and all nine survive on both providers, with no Skipping attachment(s)…
warning logged. For the seven that carry readable text the model also echoes a marker
embedded in the file, so the content is demonstrably arriving rather than merely not
breaking anything. That is consistent with resolveDefaultLLMDeliveryPath doing its job
on uploads, and with the remaining exposure being where that resolution does not run.

Happy to run anything specific against either provider if it would help — the matrix is
scripted, so a re-run is cheap.

@mihidumh

mihidumh commented Oct 5, 2026

Copy link
Copy Markdown
Contributor Author

@tommctech thanks for the repro and the provider comparison.

I read the code and I agree with your trace. The branch you point at is not from this PR: it is the same on the PR base (c8c5478cf) and on current dev. In formatDocumentBlock, an OpenAI-like provider that is not Azure and does not use the Responses API gets a file part with data:<mime>;base64,… for every type. For a model that is not Gemini, this PR only adds a text route in front of it for three application/* types, so the comment "text/*, JSON, YAML and TypeScript stay file parts" describes what the PR leaves as it was. It is not a claim that Chat Completions accepts them.

Two things follow for your case:

  • This PR does not fix it. For a genuine OpenAI model filterProviderDocumentFiles returns every file as processable, so the replay path also sends the generated text/csv as a file part. The conversation stays dead with this branch too.
  • The cause has two parts, and neither is in this diff: a generated file (context: execute_code) reaches the provider document path without the delivery-path resolution that an upload gets, and the Chat Completions branch has no text conversion for a non-PDF textual type.

I would like to keep #16473 at its current scope. It is in review with changes requested, and your case needs a decision that a maintainer should make: whether to route generated files like uploads, to send every non-PDF textual type as text on Chat Completions, or both. The second option changes what every OpenAI-compatible gateway receives, so it needs its own review.

Please open it as its own issue and link it here. Your two comments already hold the repro, the stored file record and the OpenAI/Anthropic table. The observation that the stored text is the HTML preview is worth its own line in that issue, because a fallback to stored text would send markup.

Your matrix result for the upload path on current dev is useful, but it covers openAI and anthropic direct. The failures in #16472 are on two other routes: Claude behind an OpenAI-compatible gateway, and native Gemini. Those are the routes this PR filters, and your matrix does not exercise them. I will rerun the #16472 cases on those two routes against current dev and report the result here.

@mihidumh

mihidumh commented Oct 5, 2026

Copy link
Copy Markdown
Contributor Author

The rerun I promised above. The #16472 failures still reproduce on current dev.

Setup. dev at d0cbefd9d (2026-10-04), built and run locally, a new database. fileConfig.legacyFileUploadUX: true. One custom OpenAI-compatible endpoint that points at a LiteLLM proxy, with no supportedMimeTypes of its own, plus the native google endpoint on Vertex AI. Each case: a new chat, "Upload to Provider", one file, a first message that asks about the file, then a second message with no file.

Route File Turn 1 Turn 2 (no new file)
Gateway → Azure OpenAI (GPT) query.sql (application/sql) 400 400
Gateway → Gemini on Vertex report.docx 400 400
Native google (Vertex, gemini-2.5-flash) report.docx 400 400
Gateway → Claude bundle.zip 200, the file is dropped 200
  • Azure: Invalid file data: 'messages[0].content[1].file.file_data' ... but got unsupported MIME type 'application/sql'.
  • Vertex behind the gateway: Unable to submit request because it has a mimeType parameter with value application/vnd.openxmlformats-officedocument.wordprocessingml.document, which is not supported.
  • Native google: the UI shows only "Google request failed with status code 400". A message with no file on the same endpoint and model succeeds, so the attachment is the cause.
  • Claude: no 400. The server logs Skipping attachment(s) unsupported by Claude document input: "bundle.zip" (application/zip), and the model answers "I don't see an attached file". The user gets no message about it. That is the first finding of the review above, and this PR turns it into a visible rejection on the current turn.

A correction to my comment above: I named "Claude behind a gateway" as a route that gives the 400. That is wrong. The 400 routes are GPT and Gemini behind a gateway, and native Gemini. Claude behind a gateway is the silent-drop case.

Why this differs from the nine-type matrix. The precondition is legacyFileUploadUX: true with the "Upload to Provider" choice. The server then keeps the user's choice (provider) and does not resolve a delivery path from the MIME type. With the default upload UX, the same files resolve to text and never reach the document path, which matches what @tommctech measured. The code agrees: processAttachments in api/app/clients/BaseClient.js on dev still sends a provider file to the document branch when the inherited supportedMimeTypes matches it, with no isExplicitMimeConfig test, and packages/api/src/files/encode/ has no change between this PR's base (c8c5478cf) and d0cbefd9d.

So the PR is still needed for the legacy-chooser path. With this branch, row 1 goes as a text part (application/sql is in textPartApplicationTypes). Rows 2 and 3 do not reach the provider: the Gemini branch of filterProviderDocumentFiles omits the file, and the current turn is rejected with AGENT_ATTACHMENT_UNSUPPORTED (415). Row 4 gets the same visible rejection in place of the silent drop. I did not run these four cases live on the branch; the PR's specs cover them. The branch still merges into dev without a conflict.

@tommctech, can you confirm the value of fileConfig.legacyFileUploadUX on your instance? If it is not set, a rerun of your matrix with true and the "Upload to Provider" choice should show the 400s on a gateway or Gemini route.

@tommctech

Copy link
Copy Markdown

Confirmed: legacyFileUploadUX is unset on my instance — there is no fileConfig
block in librechat.yaml at all, and it resolves to undefined. So my earlier matrix ran
on the default upload UX, which is exactly the explanation you gave.

I then set it and re-ran, which adds a route your table does not cover: native OpenAI,
direct to api.openai.com/v1/chat/completions, no gateway.
Also worth noting this is
the released v0.8.8 image (librechat-ai/librechat-api:v0.8.8), not dev, so the
legacy path is affected on the release as well.

Setup: fileConfig.endpoints.agents.legacyFileUploadUX: true, agent on gpt-5.1, each
case a new conversation with one file attached and a first message asking about it. Files
are generated fixtures, each a minimal valid example of its type.

File Type Turn 1
fixture.png image/png 200
fixture.pdf application/pdf 200, content read correctly
fixture.txt text/plain 400
fixture.csv text/csv 400
fixture.html text/html 400
fixture.json application/json 400
fixture.docx Word 400
fixture.xlsx Excel 400
fixture.zip application/zip 400

All seven failures carry the same provider message, verbatim:

400 Invalid file data: 'messages[1].content[1].file.file_data'.
Expected a base64-encoded data URL with an application/pdf MIME type
(e.g. 'data:application/pdf;base64,SGVsbG8sIFdvcmxkIQ=='),
but got unsupported MIME type 'text/plain'.

Two details that may be useful:

The flag alone is sufficient — no tool resource and no UI choice are needed. All nine
uploads were plain message attachments posted to /api/files with no tool_resource, and
every one was stored with llmDeliveryPath: "provider", including the png and the
pdf. That matches resolveDefaultUploadLLMDeliveryPath, which returns "provider"
unconditionally when endpointConfig.legacyFileUploadUX === true, before any MIME
consideration. Handy if you want this reproducible from a script rather than through the
composer.

Same instance, same build, default UX: all nine pass, and for the seven that carry
readable text the model demonstrably reads the content — each fixture embeds a marker
string and the model echoes it. So the two sets of results reconcile exactly as you
described: the delivery-path resolver runs on one path and not the other.

Scope limit, so the table is not read for more than it shows: my harness stops a case at
the first failing turn, so I observed turn 1 only and did not re-confirm the turn-2
persistence you measured. And application/pdf is the one type that survives this route,
which is consistent with the error text naming it as the only accepted MIME type for
file_data on Chat Completions.

Happy to re-run any of this — it is scripted, so another route or type set is cheap.

OpenAI chat completions accepts only PDF as `file.file_data` and answers
400 "Expected a base64-encoded data URL with an application/pdf MIME type"
for `text/plain`, CSV, HTML and JSON, which recurs on every later turn
because the file is re-sent from history. Azure OpenAI accepts those
types, and the code cannot tell the two apart behind a gateway, so a
textual type now goes as a text part under the inherited list, which
both read. An endpoint that lists the type in its own `supportedMimeTypes`
keeps the file part. The responses API keeps `input_file`, which OpenAI
documents for textual types.

Measured by @tommctech on v0.8.8 with `legacyFileUploadUX: true` on
api.openai.com (LibreChat-AI#16473).
@mihidumh

mihidumh commented Oct 5, 2026

Copy link
Copy Markdown
Contributor Author

Thanks, that settles the first question. With the flag unset your matrix ran on the default UX, and the two result sets agree.

Your second table adds a route I had not measured: native OpenAI on Chat Completions. I checked it against this branch instead of a rerun. At e5fc78762 the branch did not fix 7 of your 9 rows. For a plain OpenAI provider filterProviderDocumentFiles passes every file through, and sendsAsTextWithoutOptIn routed only application/sql, application/x-sh and application/xml to a text part, so text/plain, text/csv, text/html and application/json still went as file parts, and docx, xlsx and zip reached the provider as file parts too.

The comment in the code that kept text/* and JSON as file parts came from Azure OpenAI, which accepts those types on file.file_data (measured 200 for text/plain, csv, html, json, doc and docx). Your error text shows api.openai.com accepts PDF only there. So the two OpenAI-shaped backends differ, and behind a gateway the code cannot tell them apart.

99bc9973f extends the branch for the textual rows. For an OpenAI-like provider on Chat Completions a textual type (isNativelyReadableText: text/*, JSON, YAML, TypeScript and the rest) now goes as a text part unless the endpoint's own supportedMimeTypes lists it; the opt-in keeps the file part for a backend that accepts it, such as Azure. The Responses API path is unchanged: input_file is documented for textual types, and #13550's CSV and XLSX tests still pass. That covers your txt, csv, html and json rows, and the specs now carry those four on Chat Completions and on the Responses API.

I left docx, xlsx and zip as they are. Azure accepts docx as a file part today, so rejecting every non-PDF binary for every OpenAI-shaped backend is the maintainer decision I described above, and your table is the evidence for it. Please put it in the separate issue and link it here.

@tommctech

Copy link
Copy Markdown

Opened as #16802, with both pieces: the generated-file path and the docx/xlsx/zip rows you left for the maintainer decision. The stored-text-is-the-HTML-preview observation has its own section there, since a fallback to stored text would send markup on any path that takes it.

Noted on 99bc9973f and the Azure measurement — that explains the comment I queried in my first comment here, and I was wrong to read it as simply mistaken. It was accurate for Azure and wrong for api.openai.com, which is a harder problem than either of us was describing.

Thanks for the rerun and for the detail throughout. Happy to measure anything else on the native OpenAI or Anthropic routes; it is scripted, so it costs us a couple of minutes.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

🗺️ Content And Files codegraph: the taxonomy area this belongs to (classifier, confidence ≥ 0.9) 🗺️ File Storage codegraph: the taxonomy area this belongs to (classifier, confidence ≥ 0.9)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants