Skip to content

[integration][java][python] Send multimodal user messages through OpenAI Chat Completions - #1164

Merged
wenjin272 merged 6 commits into
apache:mainfrom
Zhuoxi2000:multimodal-openai-chat-completions
Sep 29, 2026
Merged

wenjin272 merged 6 commits into
apache:mainfrom
Zhuoxi2000:multimodal-openai-chat-completions

Conversation

@Zhuoxi2000

@Zhuoxi2000 Zhuoxi2000 commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Linked issue: #1059

Purpose of change

A user ChatMessage with image, audio or document blocks now reaches OpenAI, Azure OpenAI and vLLM with its media, in Java and Python. Until now every provider sent only the text, silently dropping media. This is the first provider step of #1059 Phase 2; Ollama and explicit errors for the remaining providers follow in separate PRs.

Runtime flow

  1. The three connections share one Chat Completions converter per language: OpenAIChatCompletionsUtils.convertToOpenAIMessage (Java) and convert_to_openai_message (Python).
  2. A system, assistant or tool message is first checked for media blocks and rejected if it has any.
  3. A user message without media is sent with its text projection as a string, as before. With media, every block is mapped in order to a content part, and the first block without a part throws.

Key decisions

  • Text-only user messages keep string content, so existing requests and OpenAI-compatible servers see no change.
  • Unsupported blocks throw a new UnsupportedContentBlockException / UnsupportedContentBlockError (an IllegalArgumentException / ValueError) from the api module, rather than being dropped or converted; the later provider PRs reuse it. Its message names the block type, media type and source type, never the payload or URL.
  • Base64 images and documents are sent as data: URIs, as OpenAI documents.
  • Documents are not restricted to PDF: OpenAI rejects other types itself; compatible servers may accept more.
  • Video and audio or documents by URL are rejected: Chat Completions has no part for them (vLLM's video_url extension is left out).

Behavioral Semantics

Interaction decisions

Role Blocks Result
user text only, or none content is the text projection string
user media, every block mappable content is a part list in block order, text blocks included
user a block without a part UnsupportedContentBlockException; no request is sent
system / assistant / tool text only unchanged
system / assistant / tool any media UnsupportedContentBlockException; no request is sent

Behavioral contracts

  1. A user message without media is sent with string content equal to its text projection.
  2. A user message with media is sent as content parts in block order, one text part per TextBlock.
  3. ImageBlock becomes image_url with the URL, or data:<media_type>;base64,<data>.
  4. AudioBlock with Base64 data becomes input_audio, format wav for audio/wav (audio/wave, audio/x-wav, audio/vnd.wave) and mp3 for audio/mpeg (audio/mp3); media-type parameters and case are ignored.
  5. DocumentBlock with Base64 data becomes file with file_data as a data URI and filename set to the block's name, else document.
  6. VideoBlock, audio or documents by URL, and other audio types throw UnsupportedContentBlockException.
  7. Media in a system, assistant or tool message throws UnsupportedContentBlockException.
  8. The exception message never contains the Base64 data or the URL.

Java and Python behave identically for each contract.

Failure behavior

  • Unsupported block or role: thrown while building the request, so nothing is sent; the chat action's error strategy applies as for any provider error. The error is deterministic, so RETRY fails the same way on every attempt.
  • A model or server that rejects an accepted part (a non-vision model, a non-PDF document on OpenAI): the provider's error propagates unchanged.
  • A tool message with media fails on the media before the existing externalId check.

Tests

Contract Java: OpenAIChatCompletionsMultimodalTest Python: test_openai_multimodal.py
1 testTextOnlyUserMessageKeepsStringContent test_text_only_user_message_keeps_string_content
2, 3 testUserMediaBecomesOrderedContentParts test_user_media_becomes_ordered_content_parts
4 testAudioBecomesInputAudio test_audio_becomes_input_audio
5 testDocumentBecomesFilePart test_document_becomes_file_part
6, 8 testUnsupportedUserBlocksFailExplicitly test_unsupported_user_blocks_fail_explicitly
7 testMediaOutsideUserMessagesFails test_media_outside_user_messages_fails

Java tests assert on the SDK-serialized wire JSON; Python tests on the request dicts.

Not verified:

  • A live endpoint: the opt-in live tests (image, WAV, named and unnamed PDF) need an API key and were not run here.
  • The audio aliases and media-type parameters: mapped in code, not individually tested.
  • image_url.detail is never set, so the provider default applies.
Implementation invariants and supporting evidence
  • The converter is shared: Java OpenAICompletionsConnection, AzureOpenAIChatModelConnection and VLLMChatModelConnection (a subclass of the first); Python OpenAIChatModelConnection, AzureOpenAIChatModelConnection and VLLMChatModelConnection likewise.
  • openai-java 4.8.0 takes media parts on user messages only (contentOfArrayOfContentParts); the system, tool and assistant builders accept text parts only, which is why media is rejected by role.
  • input_audio formats in openai-java 4.8.0 and openai-python are exactly wav and mp3.
  • OpenAI's file-input guide: Chat Completions accepts file_data for PDF only, as data:application/pdf;base64,....
  • Verified locally: the Java OpenAI integration module 132/132 and the api ChatMessage* tests, spotless; Python OpenAI, Azure and vLLM plus chat-message tests 191 passed (4 skipped), ruff check and format.

API

New public types: org.apache.flink.agents.api.chat.messages.UnsupportedContentBlockException (Java) and flink_agents.api.chat_message.UnsupportedContentBlockError (Python), each with a forBlock / for_block factory that builds the shared message, so the later provider PRs report the same way.

Compatibility: text-only requests are unchanged. User media used to be dropped; it is now sent, or fails if it has no Chat Completions part. Media in a non-user message now fails. Other providers are unchanged.

Documentation

  • doc-needed
  • doc-not-needed
  • doc-included

"Multimodal Input" under OpenAI in chat_models.md, linked from Azure OpenAI and vLLM; the provider-support note that #1163 added now names these three.

Was this patch authored or co-authored using generative AI tooling?

  • Yes
  • No

If yes, include a Generated-by: <tool name and version> (<model name and version>) line, for example Generated-by: Claude Code 2.1.226 (Claude Opus 4.6), in the commit message so it reaches Git history. Repeat the same line here for reviewer visibility. See the ASF generative tooling guidance.

Generated-by: Claude Code 2.1.259 (Claude Opus 5.5)

@github-actions github-actions Bot added doc-included Your PR already contains the necessary documentation updates. fixVersion/0.4.0 priority/major Default priority of the PR or issue. and removed doc-included Your PR already contains the necessary documentation updates. labels Sep 28, 2026
@Zhuoxi2000
Zhuoxi2000 force-pushed the multimodal-openai-chat-completions branch from 196b165 to e568091 Compare September 28, 2026 05:57
@github-actions github-actions Bot added doc-included Your PR already contains the necessary documentation updates. and removed doc-included Your PR already contains the necessary documentation updates. labels Sep 28, 2026

@wenjin272 wenjin272 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for adding multimodal support!

Could we add opt-in integration tests that send real multimodal requests through ChatModelConnection.chat()? The current tests verify the serialized request shape, but not whether the provider accepts it. Covering image, audio, and PDF input—including a PDF without name to exercise the default filename—would help validate this first provider implementation. These tests can be skipped when credentials are unavailable.

**Provider support:** Built-in provider integrations currently send only the text
portion of a message. These examples construct multimodal messages; sending their
media to a model requires a provider integration that supports those block types.
**Provider support:** The OpenAI, Azure OpenAI and vLLM integrations send media

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could we narrow this statement to the OpenAI Chat Completions integration? OpenAIResponsesModelConnection still sends only the text projection, so the current wording could imply that the Responses integration also supports media.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done: the note now names the Chat Completions integration and excludes Responses.

@Zhuoxi2000

Copy link
Copy Markdown
Contributor Author

Added opt-in live tests for image, audio and PDFs; untested here without an API key.

@github-actions github-actions Bot added doc-included Your PR already contains the necessary documentation updates. and removed doc-included Your PR already contains the necessary documentation updates. labels Sep 29, 2026

@wenjin272 wenjin272 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thanks for the implementation and the live integration tests!

I ran the Java and Python live tests against Alibaba Cloud Bailian’s OpenAI-compatible endpoint. The image and both PDF tests passed in both languages. Only audio failed with qwen3.8-omni-flash: Bailian expects a Data URI for input_audio.data, whereas this implementation sends raw Base64, consistent with OpenAI’s documented format. This looks like a provider-specific compatibility difference and should not block this PR.

@wenjin272
wenjin272 merged commit 1618728 into apache:main Sep 29, 2026
30 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

doc-included Your PR already contains the necessary documentation updates. fixVersion/0.4.0 priority/major Default priority of the PR or issue.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants