Skip to content

feat: expose streaming transcription, diarization and speaker embeddings - #97

Merged
ivan-digital merged 3 commits into
mainfrom
feat/meeting-transcription-models
Sep 15, 2026
Merged

ivan-digital merged 3 commits into
mainfrom
feat/meeting-transcription-models

Conversation

@ivan-digital

Copy link
Copy Markdown
Member

Summary

Apps that run their own capture and segmentation — a meeting recorder, a note
taker — can now load the streaming recognizer, the diarizer and the speaker
encoder on their own, without the voice pipeline. Stenograf's Android app is
the first consumer.

Depends on soniqo/speech-core#148 (incremental Nemotron features and
word timings); the submodule points at it.

What changed

  • StreamingTranscriber(TranscriberConfig): one Nemotron multilingual stream
    per instance on LiteRT (INT8/FP16) or ONNX (FP16) — setLanguage,
    beginStream, pushAudio, endStream, cancelStream. "auto" selects
    the model's automatic-language prompt, as speech-swift does. Results carry
    word timings in seconds from beginStream.
  • SpeakerDiarizer(DiarizerConfig): the Sortformer 4-speaker ONNX export as
    one stream per recording, returning per-frame speaker probabilities.
  • SpeakerEmbedder(SpeakerEmbedderConfig): ReDimNet2-B6 ONNX, 192-dimensional
    voice vectors.
  • ModelManager: ensureTranscriberModels, ensureDiarizerModels,
    ensureSpeakerEmbeddingModels with pinned revisions, byte progress and
    their own cache directories; are*Ready checks; planned*Bytes; and
    endpoint, which routes every download through a mirror the user chose
    (https only, plain http only on loopback).
  • ModelDownloadWorker can enqueue the three new sets, with
    includePipeline = false for apps that don't use the pipeline.
  • Complete LiteRT FP16 decoder and joint files were rejected as truncated by
    the filename size floors; those floors are now per bundle.
  • READMEs document the building blocks; no engine decides thresholds, turns
    or labels.

No existing public signature changed.

Test plan

  • ./gradlew :sdk:testDebugUnitTest: 157 tests, 28 new (config validation,
    model lists and URLs under a custom endpoint, readiness, worker sets).
  • :control-demo:testDebugUnitTest: 42 tests. :app:assembleDebug,
    :control-demo:assembleDebug and :sdk:assembleRelease build.
  • Instrumentation on an arm64 emulator (API 35) with real downloads:
    • StreamingTranscriberTest (5): an invented sentence comes back exactly
      with "auto" and "en-US"; words joined equal the text and end within
      the audio pushed.
    • SpeakerDiarizerTest (2): two synthetic voices take separate speaker
      columns in all four turns, mean probability ≥ 0.99.
    • SpeakerEmbedderTest (3): same voice 0.914, other voice 0.254.
  • The AAR carries libspeech_android.so for arm64-v8a and x86_64 with every
    new JNI symbol.

StreamingTranscriber (Nemotron multilingual on ONNX or LiteRT),
SpeakerDiarizer (Sortformer 4-speaker ONNX) and SpeakerEmbedder
(ReDimNet2-B6 ONNX) can be used on their own, each with pinned downloads,
a readiness check and a background download set, for apps that run their
own capture and segmentation. The transcriber's words now carry timings.
ModelManager.endpoint routes downloads through a mirror the user chose;
anything but https is refused outside loopback.

Complete LiteRT FP16 decoder and joint files were rejected as truncated
by the filename size floors; those floors are now per bundle.

Bumps speech-core for incremental Nemotron features and word timings.
@ivan-digital
ivan-digital merged commit e79b4f1 into main Sep 15, 2026
5 checks passed
@ivan-digital
ivan-digital deleted the feat/meeting-transcription-models branch September 15, 2026 19:22
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant