Skip to content

Latest commit

 

History

History
310 lines (239 loc) · 14.8 KB

File metadata and controls

310 lines (239 loc) · 14.8 KB

Interfaces

speech-core defines a small set of pure-virtual C++ interfaces in include/speech_core/interfaces.h. The orchestration layer (VoicePipeline, TurnDetector, SpeechQueue) depends only on these interfaces, never on concrete implementations. This lets callers plug in any backend — ONNX Runtime, CoreML, MLX, a remote API — without modifying the core.

This file is the C++ counterpart to speech-swift's docs/shared-protocols.md. The two libraries share the same conceptual surface; the Swift side calls them protocol, here they're class … = 0.

┌─────────────────────────────────────────────┐
│      speech_core (orchestration)            │
│                                             │
│  VoicePipeline ─┬─► STTInterface            │
│                 ├─► TTSInterface            │
│                 ├─► VADInterface            │
│                 ├─► TurnCompletionInterface │
│                 ├─► EnhancerInterface       │
│                 ├─► EchoCancellerInterface  │
│                 └─► LLMInterface (optional) │
└─────────────────────────────────────────────┘
              ▲
              │ implements
              │
┌─────────────┴────────────────────────────────┐
│  Reference implementations (optional, ORT)   │
│                                              │
│  SileroVad             : VADInterface        │
│  ParakeetStt           : STTInterface        │
│  KokoroTts             : TTSInterface        │
│  DeepFilterEnhancer    : EnhancerInterface   │
└──────────────────────────────────────────────┘

STTInterface — Speech-to-text

class STTInterface {
public:
    virtual ~STTInterface() = default;

    // Batch
    virtual TranscriptionResult transcribe(
        const float* audio, size_t length, int sample_rate) = 0;
    virtual int input_sample_rate() const = 0;
    virtual void cancel() {}

    // Optional streaming
    virtual bool supports_streaming() const { return false; }
    virtual void begin_stream(int sample_rate) {}
    virtual PartialResult push_chunk(const float* audio, size_t length) { return {}; }
    virtual void flush_stream() {}
    virtual TranscriptionResult end_stream() { return {}; }
    virtual void cancel_stream() {}
};

TranscriptionResult carries text, language (ISO 639-1 code, empty if not detected), confidence, start_time / end_time, and words — TimedWord{text, start_time, end_time} for backends that time words, empty otherwise. PartialResult is the streaming counterpart — the same fields minus start_time / end_time, with words covering the whole open stream. The Nemotron multilingual wrappers time each word from the encoder frames its tokens were emitted on; a word keeps its leading space, so the words concatenate to the text.

Reference implementation: ParakeetStt (Parakeet TDT v3 via ONNX Runtime).

Swift counterpart: SpeechRecognitionModel in AudioCommon/Protocols.swift.

TranscribeDiarizeInterface — Joint paragraph transcription and activity

class TranscribeDiarizeInterface {
public:
    virtual DiarizedTranscriptionResult transcribe_diarized(
        const float* audio, size_t length, int sample_rate) = 0;
    virtual int input_sample_rate() const = 0;
    virtual void cancel() {}
};

DiarizedTranscriptionResult carries plain lexical text, parsed segments{start_time,end_time,speaker,text}, and the original model raw_text for fail-closed wire validation. Segment speaker values are scoped to one result and are activity-routing metadata, not people. Applications must not persist them as identities.

Reference implementation: OnnxMossTranscribeDiarize.

TTSInterface — Text-to-speech

class TTSInterface {
public:
    virtual ~TTSInterface() = default;

    virtual void synthesize(
        const std::string& text,
        const std::string& language,
        TTSChunkCallback on_chunk) = 0;
    virtual void synthesize_with_options(
        const std::string& text,
        const std::string& language,
        const TtsSynthesisOptions& options,
        TTSChunkCallback on_chunk);
    virtual void set_voice(const std::string& voice_id) {}
    virtual int output_sample_rate() const = 0;
    virtual void cancel() {}
};

using TTSChunkCallback = std::function<void(
    const float* samples, size_t length, bool is_final)>;

The callback is invoked for each audio chunk during synthesis. is_final=true marks the last chunk. Implementations that aren't streaming-capable can emit a single chunk with is_final=true.

synthesize() is the streaming API and is not deprecated. synthesize_with_options() is the uniform delivery/post-processing API for all TTS implementations through the base default. Streaming with no post-processing delegates to synthesize(). Buffered accumulates all PCM for the submitted text input, applies the requested offline post-processing chain, then invokes the callback once with is_final=true. Any offline post-processing belongs on the complete buffered result, not on individual streaming chunks or internal text chunks. Models only need to override synthesize_with_options() when they need custom buffering around internal decoder chunks. VoicePipeline currently calls synthesize() unless a future pipeline config threads synthesis options through.

set_voice() selects a backend voice preset for subsequent synthesize() calls — Supertonic F1…F5 / M1…M5, Kokoro af_heart, ff_siwis, … An empty id restores the backend default (and, for Kokoro, re-arms the language-driven voice switch); unknown ids throw std::invalid_argument. The base default is a no-op so fixed-voice backends (Pocket, VoxCPM, …) need no override. It is not synchronized against synthesize() — a host that offers a per-call voice (speech-android's synthesize(text, language, voice)) wraps set_voice(id) → synthesize() → set_voice("") under its own lock.

Reference implementations: KokoroTts (Kokoro 82M via ONNX Runtime; sentence/chunk-split output with a final marker) and LiteRTKokoroTts (the guarded, three-stage 60-frame FP32 LiteRT bundle). Both emit 24 kHz Float32 output. The LiteRT implementation is enabled on Windows and Android, whose validated runtimes export the required TensorFlow Lite interpreter C API.

Swift counterpart: SpeechGenerationModel.

VADInterface — Voice activity detection (chunk-level)

class VADInterface {
public:
    virtual ~VADInterface() = default;

    virtual float process_chunk(const float* samples, size_t length) = 0;
    virtual void reset() = 0;
    virtual int input_sample_rate() const = 0;
    virtual size_t chunk_size() const = 0;
};

Returns a speech probability in [0, 1] per chunk. The probability stream is consumed by StreamingVAD (in speech_core/vad/streaming_vad.h), which applies hysteresis and emits SpeechStarted / SpeechEnded events. The pipeline owns the StreamingVAD; you supply the chunk-level model.

Reference implementation: SileroVad (Silero VAD v5 via ONNX Runtime, 512 samples @ 16 kHz).

Swift counterpart: StreamingVADProvider.

TurnCompletionInterface — End-of-turn classification (utterance-level)

class TurnCompletionInterface {
public:
    virtual ~TurnCompletionInterface() = default;

    virtual float turn_complete_probability(
        const float* samples, size_t length, int sample_rate) = 0;
};

Returns the probability in [0, 1] that the user has finished their turn, given the audio of the turn so far. A VAD only hears silence; a turn-completion model listens to the whole utterance, so a mid-sentence pause keeps the agent waiting while a finished sentence gets an immediate reply. TurnDetector calls it synchronously from the audio path once per VAD pause (Smart Turn looks at the last 8 s), so implementations must return within a few tens of milliseconds. Optional: attach one with VoicePipeline::set_turn_completion(); the decision threshold and the silence cap live in AgentConfig::turn_completion_threshold / turn_completion_max_silence.

Reference implementation: OnnxSmartTurn (Pipecat Smart Turn v3.2 via ONNX Runtime, last 8 s @ 16 kHz).

Swift counterpart: SmartTurnModel (TurnCompletionProvider) in speech-swift, bridged through sc_turn_completion_vtable_t.

EnhancerInterface — Speech enhancement / denoising

class EnhancerInterface {
public:
    virtual ~EnhancerInterface() = default;

    virtual void enhance(
        const float* audio, size_t length, int sample_rate,
        float* output) = 0;
    virtual int input_sample_rate() const = 0;
};

Pre-allocated output buffer. Caller is responsible for sample-rate matching.

Reference implementation: DeepFilterEnhancer (DeepFilterNet3 via ONNX Runtime, 48 kHz).

Swift counterpart: SpeechEnhancementModel.

EchoCancellerInterface — Acoustic echo cancellation

class EchoCancellerInterface {
public:
    virtual ~EchoCancellerInterface() = default;

    virtual void feed_reference(const float* samples, size_t length) = 0;
    virtual void cancel_echo(const float* input, size_t length, float* output) = 0;
    virtual int input_sample_rate() const = 0;
    virtual void reset() = 0;
};

Pipeline feeds TTS output via feed_reference() and runs cancel_echo() on mic input before VAD. Thread-safety: feed_reference() and cancel_echo() may be called concurrently.

Reference implementation: OnnxLocalVQEEchoCanceller.

FrameEchoCancellerInterface — Timestamp-aligned passive AEC

class FrameEchoCancellerInterface {
public:
    virtual int input_sample_rate() const = 0;
    virtual size_t frame_size() const = 0;
    virtual void process_frame(
        const float* microphone,
        const float* reference,
        float* output) = 0;
    virtual bool prime_delay(
        const float* microphone,
        const float* reference,
        size_t sample_count) = 0;
    virtual int current_delay_samples() const { return 0; }
    virtual float delay_confidence() const { return 0.0f; }
    virtual void reset() = 0;
};

This interface is for passive recorders that capture microphone and playback on independent timestamped streams. The caller supplies exact corresponding frames; TimestampedEchoCancellationStream provides bounded alignment, priming, and fail-closed queueing around it.

Reference implementation: OnnxLocalVQEEchoCanceller.

LLMInterface — Language model

class LLMInterface {
public:
    virtual ~LLMInterface() = default;

    virtual LLMResponse chat(
        const std::vector<Message>& messages,
        LLMTokenCallback on_token) = 0;
    virtual void set_tools(const std::vector<ToolDefinition>& tools) {}
    virtual void cancel() {}
};

Only needed for the VoicePipeline mode (full agent loop). Not needed for Echo or TranscribeOnly modes.

Diarization interfaces (batch, server-side)

Three additional interfaces support multi-speaker meeting diarization. Unlike the real-time pipeline interfaces above they operate on a whole audio buffer and are not consumed by VoicePipeline; a DiarizationPipeline (in speech_core/diarization/) composes a segmenter + embedder + clustering. The shared types SegmentationWindow, DiarizedSegment{start,end,speaker}, and DiarizerConfig live in interfaces.h alongside them.

SegmentationInterface — local speaker activity per window

class SegmentationInterface {
public:
    virtual ~SegmentationInterface() = default;
    virtual std::vector<SegmentationWindow> segment(
        const float* audio, size_t length, int sample_rate) = 0;
    virtual int input_sample_rate() const = 0;
    virtual int max_local_speakers() const = 0;
};

Reference implementation: LiteRTPyannoteSegmentation (Pyannote Segmentation 3.0 via LiteRT, streaming 1-s chunks, powerset-decoded per-speaker activity).

EmbeddingInterface — speaker embedding

class EmbeddingInterface {
public:
    virtual ~EmbeddingInterface() = default;
    virtual std::vector<float> embed(
        const float* audio, size_t length, int sample_rate) = 0;
    virtual int embedding_dim() const = 0;
    virtual int input_sample_rate() const = 0;
};

embed_short_utterance() is an optional conservative retrieval-only path. Callers must not create or update an identity from that result alone.

Reference implementations: LiteRTWeSpeakerEmbedding (WeSpeaker ResNet34-LM via LiteRT, 256-dim L2-normalised vector) and OnnxReDimNetSpeakerEmbedding (ReDimNet2-B6 via ONNX Runtime, 192-dim L2-normalised vector with a strict short retrieval probe).

DiarizerInterface — end-to-end diarization

class DiarizerInterface {
public:
    virtual ~DiarizerInterface() = default;
    virtual std::vector<DiarizedSegment> diarize(
        const float* audio, size_t length, int sample_rate,
        const DiarizerConfig& config) = 0;
};

Reference implementation: DiarizationPipeline (speech_core/diarization/diarization_pipeline.h) — composes a SegmentationInterface + EmbeddingInterface with constrained agglomerative clustering. Pure C++, ships in the core library (no ML-runtime dependency).

Design choices

  1. Direct inheritance. Models inherit interfaces directly (class SileroVad : public VADInterface) rather than going through adapter classes. speech-swift uses extension files to keep the same separation (Swift only — C++ has no equivalent). With only one consumer (speech-core itself), the indirection isn't worth its cost; revisit if a new caller wants the models without the orchestration.

  2. No vtable boilerplate in callers. Earlier the platform code (Linux speech.cpp, Android jni_bridge.cpp) hand-rolled sc_vad_t / sc_stt_t / sc_tts_t adapter structs that wrapped the model classes. With direct inheritance those adapters are unnecessary — the model pointer goes straight into VoicePipeline.

  3. ORT is optional. speech-core builds without ONNX Runtime by default. Set -DSPEECH_CORE_WITH_ONNX=ON -DORT_DIR=… to compile in the reference implementations. Consumers that bring their own backends (e.g. speech-swift uses CoreML/MLX) get a smaller artifact.

  4. No abstract InferenceBackend layer. An earlier design had InferenceBackend / OnnxBackend / LiteRT placeholder classes. The actual models bypassed them and talked to OrtApi directly, so the abstraction was dead code. Dropped. If a second backend lands, reintroduce it then.