Repository navigation
feat: expose streaming transcription, diarization and speaker embeddings - #97
Merged
Merged
Conversation
StreamingTranscriber (Nemotron multilingual on ONNX or LiteRT), SpeakerDiarizer (Sortformer 4-speaker ONNX) and SpeakerEmbedder (ReDimNet2-B6 ONNX) can be used on their own, each with pinned downloads, a readiness check and a background download set, for apps that run their own capture and segmentation. The transcriber's words now carry timings. ModelManager.endpoint routes downloads through a mirror the user chose; anything but https is refused outside loopback. Complete LiteRT FP16 decoder and joint files were rejected as truncated by the filename size floors; those floors are now per bundle. Bumps speech-core for incremental Nemotron features and word timings.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Apps that run their own capture and segmentation — a meeting recorder, a note
taker — can now load the streaming recognizer, the diarizer and the speaker
encoder on their own, without the voice pipeline. Stenograf's Android app is
the first consumer.
Depends on soniqo/speech-core#148 (incremental Nemotron features and
word timings); the submodule points at it.
What changed
StreamingTranscriber(TranscriberConfig): one Nemotron multilingual streamper instance on LiteRT (INT8/FP16) or ONNX (FP16) —
setLanguage,beginStream,pushAudio,endStream,cancelStream."auto"selectsthe model's automatic-language prompt, as speech-swift does. Results carry
word timings in seconds from
beginStream.SpeakerDiarizer(DiarizerConfig): the Sortformer 4-speaker ONNX export asone stream per recording, returning per-frame speaker probabilities.
SpeakerEmbedder(SpeakerEmbedderConfig): ReDimNet2-B6 ONNX, 192-dimensionalvoice vectors.
ModelManager:ensureTranscriberModels,ensureDiarizerModels,ensureSpeakerEmbeddingModelswith pinned revisions, byte progress andtheir own cache directories;
are*Readychecks;planned*Bytes; andendpoint, which routes every download through a mirror the user chose(https only, plain http only on loopback).
ModelDownloadWorkercan enqueue the three new sets, withincludePipeline = falsefor apps that don't use the pipeline.the filename size floors; those floors are now per bundle.
or labels.
No existing public signature changed.
Test plan
./gradlew :sdk:testDebugUnitTest: 157 tests, 28 new (config validation,model lists and URLs under a custom endpoint, readiness, worker sets).
:control-demo:testDebugUnitTest: 42 tests.:app:assembleDebug,:control-demo:assembleDebugand:sdk:assembleReleasebuild.StreamingTranscriberTest(5): an invented sentence comes back exactlywith
"auto"and"en-US"; words joined equal the text and end withinthe audio pushed.
SpeakerDiarizerTest(2): two synthetic voices take separate speakercolumns in all four turns, mean probability ≥ 0.99.
SpeakerEmbedderTest(3): same voice 0.914, other voice 0.254.libspeech_android.sofor arm64-v8a and x86_64 with everynew JNI symbol.