Speech: TTS and ASR
Text-to-speech, speech recognition and voice models you can run yourself.
73 links, newest first.
- Speech: TTS and ASRPost on X
NVIDIA releases Nemotron 3 Diarization model
Nemotron 3 Diarization tracks who spoke when in overlapping speech, supports up to eight speakers, and has 100M parameters. NVIDIA says it is available on Hugging Face.
Engineers can evaluate a speaker diarization model for handling overlapping voices.
- Speech: TTS and ASRModel
Youdao releases streaming ASR and translation models
The post reviews Youdao’s Confucius4-R2T2 for real-time ASR and Confucius4-T3PO for real-time translation. It says both support streaming and can be connected in sequence.
Engineers can explore a streaming ASR-to-translation setup for live speech applications.
- Speech: TTS and ASRPost on X
Voice of Reason answers spoken math problems in speech
Kyutai Labs says Voice of Reason takes spoken math problems without transcription or a text LLM in the loop, then reasons and answers in speech. The post reports 77.1% on GSM8K, compared with 27.3% for GLM-4-Voice.
It describes a speech-native approach to spoken math reasoning without an intermediate transcription or text LLM.
- Speech: TTS and ASRRepository
GrainSpeech: compact speech synthesis with less context
GrainSpeech is a compact speech synthesis project described as using less context for more detail. Its official repository is linked alongside a project page and paper.
Engineers can review the repository and paper for details on a compact TTS approach.
- Speech: TTS and ASRPost on X
Supersonic TTS Runs on CPU and Supports 31 Languages
The author says Supersonic runs on CPU and supports 31 languages. They plan to use it to generate synthetic data.
CPU support and multilingual coverage may make it useful to evaluate for local TTS and synthetic-data workflows.
- Speech: TTS and ASRRepository
Xiaomi’s CocktailASR-1 targets a selected speaker in mixed audio
CocktailASR-1 is an end-to-end target-speaker speech recognition model that uses reference voiceprints to transcribe a selected voice in overlapping audio. The post says it supports English and Chinese.
It may help engineers evaluate speaker-conditioned ASR for crowded or overlapping audio.
- Speech: TTS and ASRRepository
RVC WebUI for Training Voice Conversion Models
RVC is an open-source WebUI project for training voice conversion models. The repository describes training a model with 10 minutes or less of voice data.
Engineers can inspect a self-hostable voice conversion workflow and its model-training requirements.
Xiaomi-CocktailASR-1 transcribes a target speaker in mixed speech
A Hugging Face Space for transcribing only a target speaker in mixed speech, using an audio clip of that speaker to isolate their voice.
Engineers can try a speaker-conditioned ASR workflow on overlapping speech.
- Speech: TTS and ASRRepository
OmniVoice: voice-cloning TTS for 600+ languages
OmniVoice is a Python voice-cloning text-to-speech project described as supporting 600+ languages. The post claims zero-shot synthesis and voice controls such as age, pitch, accent, and whispering.
Engineers can evaluate a multilingual TTS project for voice cloning and configurable speech generation.
- Speech: TTS and ASRPaper
Phoenix-VAD: Streaming Semantic Endpoint Detection
The paper introduces Phoenix-VAD, an LLM-based model for streaming semantic endpoint detection in full-duplex speech interaction. It uses a sliding-window training strategy.
Engineers building spoken dialogue systems can evaluate a semantic endpoint detection approach for streaming interactions.
- Speech: TTS and ASRPost on X
Microsoft releases MAI-Transcribe-2 speech recognition model
Microsoft’s MAI-Transcribe-2 is a speech transcription model reported at 2.0% AA-WER and about 411× real time. It supports 60 languages, speaker diarization, word-level timestamps, and clean or verbatim transcription styles.
Engineers can compare its reported accuracy, speed, and price when evaluating speech recognition services.
- Speech: TTS and ASRPost on X
VibeVoice ASR Streaming combines transcription and diarization
The post announces the release of VibeVoice ASR Streaming, describing it as a streaming model that supports transcription and diarization in one model.
Engineers evaluating speech-recognition systems may want to assess its streaming and diarization capabilities.
- Speech: TTS and ASRArticle
Photon 2.1 adds streaming speech recognition
Photon 2.1 adds automatic speech recognition for Whisper, Qwen3-ASR, and Parakeet. Moondream reports that it beat each model’s reference engine in every matched test and supports NVIDIA B200.
Engineers evaluating self-hosted ASR can compare Photon’s reported performance across models and GPU types.
- Speech: TTS and ASRPost on X
Phonon-1 speech recognition model released
Phonon-1 is a 782-million-parameter speech recognition model with a 415 MB download and Apache 2.0 licensing. Its post claims it transcribes an hour of audio in about two minutes on a MacBook Air.
Its reported size, license, and local runtime may be useful when evaluating self-hosted ASR models.
- Speech: TTS and ASRArticle
TurnBench benchmarks turn-taking in spoken dialogue
TurnBench is an open benchmark with 30 hours of triple-annotated dyadic speech, a scoring protocol grounded in conversation analysis, and results for 14 systems. It measures turn-taking in spoken dialogue.
Engineers can use the benchmark to evaluate whether voice systems handle turn endings and backchannel cues.
- Speech: TTS and ASRRepository
FireRedAudio combines speech recognition, audio understanding, and TTS
FireRedAudio uses Qwen3.5-9B as a shared backbone for ASR, audio understanding, zero-shot and instruction-based TTS, and speech editing. It uses separate continuous representations for audio understanding and generation.
Engineers can explore an open-source model covering multiple speech tasks in one audio language model.
- Speech: TTS and ASRRepository
Breeze TTS 2 releases open weights
BreezeBlue says Breeze TTS 2 is now open-weight and links to its GitHub repository and model weights on Hugging Face. The author claims it ranks first among open-weight TTS models in the Artificial Analysis Provider Voices Speech Arena.
Engineers can inspect and test the model weights and repository for self-hosted TTS use.
- Speech: TTS and ASRModel
Pocket TTS runs on Android with LiteRT
Pocket TTS, a 100M-parameter text-to-speech model, runs on-device with LiteRT. The post reports six bundled voices and real-time or faster playback on current Android devices.
Engineers can evaluate an on-device TTS option for Android apps.
- Speech: TTS and ASRPost on X
Breeze TTS 2 offers open weights and streaming generation
Breeze TTS 2 is a text-to-speech model supporting 50 languages, voice generation from text prompts, and streaming generation. Its weights are available on Hugging Face.
Engineers can evaluate an openly weighted TTS model with multilingual and streaming capabilities.
- Speech: TTS and ASRRepository
ComfyUI plugin adds VoxCPM2 inference and LoRA training
ComfyUI_VoxCPM_SM is a ComfyUI plugin for VoxCPM2 inference and LoRA fine-tuning with custom voice data. The post says it supports low-VRAM use and quantized models, including INT4, INT8, and GGUF.
Engineers can use a node-based workflow to run and customize VoxCPM2 voice generation locally.
Build with AgentLog
Send your own newsletter
CommsHarbor keeps contacts, consent and one-click unsubscribe together.
Free workspace
Open CommsHarbor









