[integration][java][python] Send multimodal user messages through OpenAI Chat Completions - #1164
Conversation
196b165 to
e568091
Compare
wenjin272
left a comment
There was a problem hiding this comment.
Thanks for adding multimodal support!
Could we add opt-in integration tests that send real multimodal requests through ChatModelConnection.chat()? The current tests verify the serialized request shape, but not whether the provider accepts it. Covering image, audio, and PDF input—including a PDF without name to exercise the default filename—would help validate this first provider implementation. These tests can be skipped when credentials are unavailable.
| **Provider support:** Built-in provider integrations currently send only the text | ||
| portion of a message. These examples construct multimodal messages; sending their | ||
| media to a model requires a provider integration that supports those block types. | ||
| **Provider support:** The OpenAI, Azure OpenAI and vLLM integrations send media |
There was a problem hiding this comment.
Could we narrow this statement to the OpenAI Chat Completions integration? OpenAIResponsesModelConnection still sends only the text projection, so the current wording could imply that the Responses integration also supports media.
There was a problem hiding this comment.
Done: the note now names the Chat Completions integration and excludes Responses.
…hat Completions integration in docs
|
Added opt-in live tests for image, audio and PDFs; untested here without an API key. |
wenjin272
left a comment
There was a problem hiding this comment.
LGTM, thanks for the implementation and the live integration tests!
I ran the Java and Python live tests against Alibaba Cloud Bailian’s OpenAI-compatible endpoint. The image and both PDF tests passed in both languages. Only audio failed with qwen3.8-omni-flash: Bailian expects a Data URI for input_audio.data, whereas this implementation sends raw Base64, consistent with OpenAI’s documented format. This looks like a provider-specific compatibility difference and should not block this PR.
Linked issue: #1059
Purpose of change
A user
ChatMessagewith image, audio or document blocks now reaches OpenAI, Azure OpenAI and vLLM with its media, in Java and Python. Until now every provider sent only the text, silently dropping media. This is the first provider step of #1059 Phase 2; Ollama and explicit errors for the remaining providers follow in separate PRs.Runtime flow
OpenAIChatCompletionsUtils.convertToOpenAIMessage(Java) andconvert_to_openai_message(Python).Key decisions
UnsupportedContentBlockException/UnsupportedContentBlockError(anIllegalArgumentException/ValueError) from the api module, rather than being dropped or converted; the later provider PRs reuse it. Its message names the block type, media type and source type, never the payload or URL.data:URIs, as OpenAI documents.video_urlextension is left out).Behavioral Semantics
Interaction decisions
contentis the text projection stringcontentis a part list in block order, text blocks includedUnsupportedContentBlockException; no request is sentUnsupportedContentBlockException; no request is sentBehavioral contracts
textpart perTextBlock.ImageBlockbecomesimage_urlwith the URL, ordata:<media_type>;base64,<data>.AudioBlockwith Base64 data becomesinput_audio, formatwavforaudio/wav(audio/wave,audio/x-wav,audio/vnd.wave) andmp3foraudio/mpeg(audio/mp3); media-type parameters and case are ignored.DocumentBlockwith Base64 data becomesfilewithfile_dataas a data URI andfilenameset to the block'sname, elsedocument.VideoBlock, audio or documents by URL, and other audio types throwUnsupportedContentBlockException.UnsupportedContentBlockException.Java and Python behave identically for each contract.
Failure behavior
externalIdcheck.Tests
OpenAIChatCompletionsMultimodalTesttest_openai_multimodal.pytestTextOnlyUserMessageKeepsStringContenttest_text_only_user_message_keeps_string_contenttestUserMediaBecomesOrderedContentPartstest_user_media_becomes_ordered_content_partstestAudioBecomesInputAudiotest_audio_becomes_input_audiotestDocumentBecomesFileParttest_document_becomes_file_parttestUnsupportedUserBlocksFailExplicitlytest_unsupported_user_blocks_fail_explicitlytestMediaOutsideUserMessagesFailstest_media_outside_user_messages_failsJava tests assert on the SDK-serialized wire JSON; Python tests on the request dicts.
Not verified:
image_url.detailis never set, so the provider default applies.Implementation invariants and supporting evidence
OpenAICompletionsConnection,AzureOpenAIChatModelConnectionandVLLMChatModelConnection(a subclass of the first); PythonOpenAIChatModelConnection,AzureOpenAIChatModelConnectionandVLLMChatModelConnectionlikewise.contentOfArrayOfContentParts); the system, tool and assistant builders accept text parts only, which is why media is rejected by role.input_audioformats in openai-java 4.8.0 and openai-python are exactlywavandmp3.file_datafor PDF only, asdata:application/pdf;base64,....ChatMessage*tests, spotless; Python OpenAI, Azure and vLLM plus chat-message tests 191 passed (4 skipped), ruff check and format.API
New public types:
org.apache.flink.agents.api.chat.messages.UnsupportedContentBlockException(Java) andflink_agents.api.chat_message.UnsupportedContentBlockError(Python), each with aforBlock/for_blockfactory that builds the shared message, so the later provider PRs report the same way.Compatibility: text-only requests are unchanged. User media used to be dropped; it is now sent, or fails if it has no Chat Completions part. Media in a non-user message now fails. Other providers are unchanged.
Documentation
doc-neededdoc-not-neededdoc-included"Multimodal Input" under OpenAI in
chat_models.md, linked from Azure OpenAI and vLLM; the provider-support note that #1163 added now names these three.Was this patch authored or co-authored using generative AI tooling?
If yes, include a
Generated-by: <tool name and version> (<model name and version>)line, for exampleGenerated-by: Claude Code 2.1.226 (Claude Opus 4.6), in the commit message so it reaches Git history. Repeat the same line here for reviewer visibility. See the ASF generative tooling guidance.Generated-by: Claude Code 2.1.259 (Claude Opus 5.5)