Update Qwen3-ASR-1.7B example to the library's vLLM transcriptions-route configuration - #587
Closed
HasanKhan04 wants to merge 1 commit into
Closed
HasanKhan04 wants to merge 1 commit into
HasanKhan04 wants to merge 1 commit into
Conversation
…ute configuration Replaces the pre-model-registry truss (vLLM nightly image, chat-completions endpoint, HF token secret, H100_40GB) with the configuration the Model Library now ships: vLLM 0.29, /v1/audio/transcriptions, four API-server processes, decode caps lifted, default output cap, BDN-cached weights, RTX PRO 6000. README updated to the transcriptions example. Source of truth: model-registry stt/qwen3-asr-1.7b/latency (basetenlabs/model-registry#511). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Replaces the pre-model-registry truss in
qwen/qwen-3-asr/(vLLM nightly image, chat-completions endpoint, HF token secret, H100_40GB) with the configuration the Model Library ships as of basetenlabs/model-registry#511: vLLM 0.29,/v1/audio/transcriptions, four API-server processes, decode caps lifted, default output cap, BDN-cached pinned weights, RTX PRO 6000. README updated to the transcriptions example.Why: the public library page for Qwen 3 ASR 1.7B links its README and repository to this directory, which still showed the old chat-route example and a config the library no longer deploys.
Source of truth is
stt/qwen3-asr-1.7b/latencyin model-registry. The registry's sync workflow already copies that preset into this repo, but onto themodel-registrybranch, notmain, so this directory and the synced copy will drift again the next time the preset changes unless one of the two is made canonical.Linear: FDE-3211.
🤖 Generated with Claude Code