speech-core defines a small set of pure-virtual C++ interfaces in include/speech_core/interfaces.h. The orchestration layer (VoicePipeline, TurnDetector, SpeechQueue) depends only on these interfaces, never on concrete implementations. This lets callers plug in any backend — ONNX Runtime, CoreML, MLX, a remote API — without modifying the core.
This file is the C++ counterpart to speech-swift's docs/shared-protocols.md. The two libraries share the same conceptual surface; the Swift side calls them protocol, here they're class … = 0.
┌─────────────────────────────────────────────┐
│ speech_core (orchestration) │
│ │
│ VoicePipeline ─┬─► STTInterface │
│ ├─► TTSInterface │
│ ├─► VADInterface │
│ ├─► TurnCompletionInterface │
│ ├─► EnhancerInterface │
│ ├─► EchoCancellerInterface │
│ └─► LLMInterface (optional) │
└─────────────────────────────────────────────┘
▲
│ implements
│
┌─────────────┴────────────────────────────────┐
│ Reference implementations (optional, ORT) │
│ │
│ SileroVad : VADInterface │
│ ParakeetStt : STTInterface │
│ KokoroTts : TTSInterface │
│ DeepFilterEnhancer : EnhancerInterface │
└──────────────────────────────────────────────┘
class STTInterface {
public:
virtual ~STTInterface() = default;
// Batch
virtual TranscriptionResult transcribe(
const float* audio, size_t length, int sample_rate) = 0;
virtual int input_sample_rate() const = 0;
virtual void cancel() {}
// Optional streaming
virtual bool supports_streaming() const { return false; }
virtual void begin_stream(int sample_rate) {}
virtual PartialResult push_chunk(const float* audio, size_t length) { return {}; }
virtual void flush_stream() {}
virtual TranscriptionResult end_stream() { return {}; }
virtual void cancel_stream() {}
};TranscriptionResult carries text, language (ISO 639-1 code, empty if not detected), confidence, start_time / end_time, and words — TimedWord{text, start_time, end_time} for backends that time words, empty otherwise. PartialResult is the streaming counterpart — the same fields minus start_time / end_time, with words covering the whole open stream. The Nemotron multilingual wrappers time each word from the encoder frames its tokens were emitted on; a word keeps its leading space, so the words concatenate to the text.
Reference implementation: ParakeetStt (Parakeet TDT v3 via ONNX Runtime).
Swift counterpart: SpeechRecognitionModel in AudioCommon/Protocols.swift.
class TranscribeDiarizeInterface {
public:
virtual DiarizedTranscriptionResult transcribe_diarized(
const float* audio, size_t length, int sample_rate) = 0;
virtual int input_sample_rate() const = 0;
virtual void cancel() {}
};DiarizedTranscriptionResult carries plain lexical text, parsed
segments{start_time,end_time,speaker,text}, and the original model
raw_text for fail-closed wire validation. Segment speaker values are scoped
to one result and are activity-routing metadata, not people. Applications must
not persist them as identities.
Reference implementation: OnnxMossTranscribeDiarize.
class TTSInterface {
public:
virtual ~TTSInterface() = default;
virtual void synthesize(
const std::string& text,
const std::string& language,
TTSChunkCallback on_chunk) = 0;
virtual void synthesize_with_options(
const std::string& text,
const std::string& language,
const TtsSynthesisOptions& options,
TTSChunkCallback on_chunk);
virtual void set_voice(const std::string& voice_id) {}
virtual int output_sample_rate() const = 0;
virtual void cancel() {}
};
using TTSChunkCallback = std::function<void(
const float* samples, size_t length, bool is_final)>;The callback is invoked for each audio chunk during synthesis. is_final=true marks the last chunk. Implementations that aren't streaming-capable can emit a single chunk with is_final=true.
synthesize() is the streaming API and is not deprecated. synthesize_with_options() is the uniform delivery/post-processing API for all TTS implementations through the base default. Streaming with no post-processing delegates to synthesize(). Buffered accumulates all PCM for the submitted text input, applies the requested offline post-processing chain, then invokes the callback once with is_final=true. Any offline post-processing belongs on the complete buffered result, not on individual streaming chunks or internal text chunks. Models only need to override synthesize_with_options() when they need custom buffering around internal decoder chunks. VoicePipeline currently calls synthesize() unless a future pipeline config threads synthesis options through.
set_voice() selects a backend voice preset for subsequent synthesize() calls — Supertonic F1…F5 / M1…M5, Kokoro af_heart, ff_siwis, … An empty id restores the backend default (and, for Kokoro, re-arms the language-driven voice switch); unknown ids throw std::invalid_argument. The base default is a no-op so fixed-voice backends (Pocket, VoxCPM, …) need no override. It is not synchronized against synthesize() — a host that offers a per-call voice (speech-android's synthesize(text, language, voice)) wraps set_voice(id) → synthesize() → set_voice("") under its own lock.
Reference implementations: KokoroTts (Kokoro 82M via ONNX Runtime;
sentence/chunk-split output with a final marker) and LiteRTKokoroTts (the
guarded, three-stage 60-frame FP32 LiteRT bundle). Both emit 24 kHz Float32
output. The LiteRT implementation is enabled on Windows and Android, whose
validated runtimes export the required TensorFlow Lite interpreter C API.
Swift counterpart: SpeechGenerationModel.
class VADInterface {
public:
virtual ~VADInterface() = default;
virtual float process_chunk(const float* samples, size_t length) = 0;
virtual void reset() = 0;
virtual int input_sample_rate() const = 0;
virtual size_t chunk_size() const = 0;
};Returns a speech probability in [0, 1] per chunk. The probability stream is consumed by StreamingVAD (in speech_core/vad/streaming_vad.h), which applies hysteresis and emits SpeechStarted / SpeechEnded events. The pipeline owns the StreamingVAD; you supply the chunk-level model.
Reference implementation: SileroVad (Silero VAD v5 via ONNX Runtime, 512 samples @ 16 kHz).
Swift counterpart: StreamingVADProvider.
class TurnCompletionInterface {
public:
virtual ~TurnCompletionInterface() = default;
virtual float turn_complete_probability(
const float* samples, size_t length, int sample_rate) = 0;
};Returns the probability in [0, 1] that the user has finished their turn, given the audio of the turn so far. A VAD only hears silence; a turn-completion model listens to the whole utterance, so a mid-sentence pause keeps the agent waiting while a finished sentence gets an immediate reply. TurnDetector calls it synchronously from the audio path once per VAD pause (Smart Turn looks at the last 8 s), so implementations must return within a few tens of milliseconds. Optional: attach one with VoicePipeline::set_turn_completion(); the decision threshold and the silence cap live in AgentConfig::turn_completion_threshold / turn_completion_max_silence.
Reference implementation: OnnxSmartTurn (Pipecat Smart Turn v3.2 via ONNX Runtime, last 8 s @ 16 kHz).
Swift counterpart: SmartTurnModel (TurnCompletionProvider) in speech-swift, bridged through sc_turn_completion_vtable_t.
class EnhancerInterface {
public:
virtual ~EnhancerInterface() = default;
virtual void enhance(
const float* audio, size_t length, int sample_rate,
float* output) = 0;
virtual int input_sample_rate() const = 0;
};Pre-allocated output buffer. Caller is responsible for sample-rate matching.
Reference implementation: DeepFilterEnhancer (DeepFilterNet3 via ONNX Runtime, 48 kHz).
Swift counterpart: SpeechEnhancementModel.
class EchoCancellerInterface {
public:
virtual ~EchoCancellerInterface() = default;
virtual void feed_reference(const float* samples, size_t length) = 0;
virtual void cancel_echo(const float* input, size_t length, float* output) = 0;
virtual int input_sample_rate() const = 0;
virtual void reset() = 0;
};Pipeline feeds TTS output via feed_reference() and runs cancel_echo() on
mic input before VAD. Thread-safety: feed_reference() and cancel_echo() may
be called concurrently.
Reference implementation: OnnxLocalVQEEchoCanceller.
class FrameEchoCancellerInterface {
public:
virtual int input_sample_rate() const = 0;
virtual size_t frame_size() const = 0;
virtual void process_frame(
const float* microphone,
const float* reference,
float* output) = 0;
virtual bool prime_delay(
const float* microphone,
const float* reference,
size_t sample_count) = 0;
virtual int current_delay_samples() const { return 0; }
virtual float delay_confidence() const { return 0.0f; }
virtual void reset() = 0;
};This interface is for passive recorders that capture microphone and playback
on independent timestamped streams. The caller supplies exact corresponding
frames; TimestampedEchoCancellationStream provides bounded alignment,
priming, and fail-closed queueing around it.
Reference implementation: OnnxLocalVQEEchoCanceller.
class LLMInterface {
public:
virtual ~LLMInterface() = default;
virtual LLMResponse chat(
const std::vector<Message>& messages,
LLMTokenCallback on_token) = 0;
virtual void set_tools(const std::vector<ToolDefinition>& tools) {}
virtual void cancel() {}
};Only needed for the VoicePipeline mode (full agent loop). Not needed for Echo or TranscribeOnly modes.
Three additional interfaces support multi-speaker meeting diarization. Unlike
the real-time pipeline interfaces above they operate on a whole audio buffer and
are not consumed by VoicePipeline; a DiarizationPipeline (in
speech_core/diarization/) composes a segmenter + embedder + clustering. The
shared types SegmentationWindow, DiarizedSegment{start,end,speaker}, and
DiarizerConfig live in interfaces.h alongside them.
class SegmentationInterface {
public:
virtual ~SegmentationInterface() = default;
virtual std::vector<SegmentationWindow> segment(
const float* audio, size_t length, int sample_rate) = 0;
virtual int input_sample_rate() const = 0;
virtual int max_local_speakers() const = 0;
};Reference implementation: LiteRTPyannoteSegmentation (Pyannote Segmentation 3.0 via LiteRT, streaming 1-s chunks, powerset-decoded per-speaker activity).
class EmbeddingInterface {
public:
virtual ~EmbeddingInterface() = default;
virtual std::vector<float> embed(
const float* audio, size_t length, int sample_rate) = 0;
virtual int embedding_dim() const = 0;
virtual int input_sample_rate() const = 0;
};embed_short_utterance() is an optional conservative retrieval-only path.
Callers must not create or update an identity from that result alone.
Reference implementations: LiteRTWeSpeakerEmbedding (WeSpeaker
ResNet34-LM via LiteRT, 256-dim L2-normalised vector) and
OnnxReDimNetSpeakerEmbedding (ReDimNet2-B6 via ONNX Runtime, 192-dim
L2-normalised vector with a strict short retrieval probe).
class DiarizerInterface {
public:
virtual ~DiarizerInterface() = default;
virtual std::vector<DiarizedSegment> diarize(
const float* audio, size_t length, int sample_rate,
const DiarizerConfig& config) = 0;
};Reference implementation: DiarizationPipeline (speech_core/diarization/diarization_pipeline.h) — composes a SegmentationInterface + EmbeddingInterface with constrained agglomerative clustering. Pure C++, ships in the core library (no ML-runtime dependency).
-
Direct inheritance. Models inherit interfaces directly (
class SileroVad : public VADInterface) rather than going through adapter classes. speech-swift uses extension files to keep the same separation (Swift only — C++ has no equivalent). With only one consumer (speech-core itself), the indirection isn't worth its cost; revisit if a new caller wants the models without the orchestration. -
No vtable boilerplate in callers. Earlier the platform code (Linux
speech.cpp, Androidjni_bridge.cpp) hand-rolledsc_vad_t/sc_stt_t/sc_tts_tadapter structs that wrapped the model classes. With direct inheritance those adapters are unnecessary — the model pointer goes straight intoVoicePipeline. -
ORT is optional. speech-core builds without ONNX Runtime by default. Set
-DSPEECH_CORE_WITH_ONNX=ON -DORT_DIR=…to compile in the reference implementations. Consumers that bring their own backends (e.g. speech-swift uses CoreML/MLX) get a smaller artifact. -
No abstract
InferenceBackendlayer. An earlier design hadInferenceBackend/OnnxBackend/LiteRTplaceholder classes. The actual models bypassed them and talked toOrtApidirectly, so the abstraction was dead code. Dropped. If a second backend lands, reintroduce it then.