Skip to content
EN

Speech: TTS and ASR

Text-to-speech, speech recognition and voice models you can run yourself.

73 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: Speech: TTS and ASR

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. NVIDIA releases Nemotron 3 Diarization model

    Nemotron 3 Diarization tracks who spoke when in overlapping speech, supports up to eight speakers, and has 100M parameters. NVIDIA says it is available on Hugging Face.

    Engineers can evaluate a speaker diarization model for handling overlapping voices.

  2. Youdao releases streaming ASR and translation models

    The post reviews Youdao’s Confucius4-R2T2 for real-time ASR and Confucius4-T3PO for real-time translation. It says both support streaming and can be connected in sequence.

    Engineers can explore a streaming ASR-to-translation setup for live speech applications.

  3. Voice of Reason answers spoken math problems in speech

    Kyutai Labs says Voice of Reason takes spoken math problems without transcription or a text LLM in the loop, then reasons and answers in speech. The post reports 77.1% on GSM8K, compared with 27.3% for GLM-4-Voice.

    It describes a speech-native approach to spoken math reasoning without an intermediate transcription or text LLM.

  4. GrainSpeech: compact speech synthesis with less context

    GrainSpeech is a compact speech synthesis project described as using less context for more detail. Its official repository is linked alongside a project page and paper.

    Engineers can review the repository and paper for details on a compact TTS approach.

  5. Supersonic TTS Runs on CPU and Supports 31 Languages

    The author says Supersonic runs on CPU and supports 31 languages. They plan to use it to generate synthetic data.

    CPU support and multilingual coverage may make it useful to evaluate for local TTS and synthetic-data workflows.

  6. Xiaomi’s CocktailASR-1 targets a selected speaker in mixed audio

    CocktailASR-1 is an end-to-end target-speaker speech recognition model that uses reference voiceprints to transcribe a selected voice in overlapping audio. The post says it supports English and Chinese.

    It may help engineers evaluate speaker-conditioned ASR for crowded or overlapping audio.

  7. RVC WebUI for Training Voice Conversion Models

    RVC is an open-source WebUI project for training voice conversion models. The repository describes training a model with 10 minutes or less of voice data.

    Engineers can inspect a self-hostable voice conversion workflow and its model-training requirements.

  8. Xiaomi-CocktailASR-1 transcribes a target speaker in mixed speech

    A Hugging Face Space for transcribing only a target speaker in mixed speech, using an audio clip of that speaker to isolate their voice.

    Engineers can try a speaker-conditioned ASR workflow on overlapping speech.

  9. OmniVoice: voice-cloning TTS for 600+ languages

    OmniVoice is a Python voice-cloning text-to-speech project described as supporting 600+ languages. The post claims zero-shot synthesis and voice controls such as age, pitch, accent, and whispering.

    Engineers can evaluate a multilingual TTS project for voice cloning and configurable speech generation.

  10. Phoenix-VAD: Streaming Semantic Endpoint Detection

    The paper introduces Phoenix-VAD, an LLM-based model for streaming semantic endpoint detection in full-duplex speech interaction. It uses a sliding-window training strategy.

    Engineers building spoken dialogue systems can evaluate a semantic endpoint detection approach for streaming interactions.

  11. Microsoft releases MAI-Transcribe-2 speech recognition model

    Microsoft’s MAI-Transcribe-2 is a speech transcription model reported at 2.0% AA-WER and about 411× real time. It supports 60 languages, speaker diarization, word-level timestamps, and clean or verbatim transcription styles.

    Engineers can compare its reported accuracy, speed, and price when evaluating speech recognition services.

  12. VibeVoice ASR Streaming combines transcription and diarization

    The post announces the release of VibeVoice ASR Streaming, describing it as a streaming model that supports transcription and diarization in one model.

    Engineers evaluating speech-recognition systems may want to assess its streaming and diarization capabilities.

  13. Photon 2.1 adds streaming speech recognition

    Photon 2.1 adds automatic speech recognition for Whisper, Qwen3-ASR, and Parakeet. Moondream reports that it beat each model’s reference engine in every matched test and supports NVIDIA B200.

    Engineers evaluating self-hosted ASR can compare Photon’s reported performance across models and GPU types.

  14. Phonon-1 speech recognition model released

    Phonon-1 is a 782-million-parameter speech recognition model with a 415 MB download and Apache 2.0 licensing. Its post claims it transcribes an hour of audio in about two minutes on a MacBook Air.

    Its reported size, license, and local runtime may be useful when evaluating self-hosted ASR models.

  15. TurnBench benchmarks turn-taking in spoken dialogue

    TurnBench is an open benchmark with 30 hours of triple-annotated dyadic speech, a scoring protocol grounded in conversation analysis, and results for 14 systems. It measures turn-taking in spoken dialogue.

    Engineers can use the benchmark to evaluate whether voice systems handle turn endings and backchannel cues.

  16. FireRedAudio combines speech recognition, audio understanding, and TTS

    FireRedAudio uses Qwen3.5-9B as a shared backbone for ASR, audio understanding, zero-shot and instruction-based TTS, and speech editing. It uses separate continuous representations for audio understanding and generation.

    Engineers can explore an open-source model covering multiple speech tasks in one audio language model.

  17. Breeze TTS 2 releases open weights

    BreezeBlue says Breeze TTS 2 is now open-weight and links to its GitHub repository and model weights on Hugging Face. The author claims it ranks first among open-weight TTS models in the Artificial Analysis Provider Voices Speech Arena.

    Engineers can inspect and test the model weights and repository for self-hosted TTS use.

  18. Pocket TTS runs on Android with LiteRT

    Pocket TTS, a 100M-parameter text-to-speech model, runs on-device with LiteRT. The post reports six bundled voices and real-time or faster playback on current Android devices.

    Engineers can evaluate an on-device TTS option for Android apps.

  19. Breeze TTS 2 offers open weights and streaming generation

    Breeze TTS 2 is a text-to-speech model supporting 50 languages, voice generation from text prompts, and streaming generation. Its weights are available on Hugging Face.

    Engineers can evaluate an openly weighted TTS model with multilingual and streaming capabilities.

  20. ComfyUI plugin adds VoxCPM2 inference and LoRA training

    ComfyUI_VoxCPM_SM is a ComfyUI plugin for VoxCPM2 inference and LoRA fine-tuning with custom voice data. The post says it supports low-VRAM use and quantized models, including INT4, INT8, and GGUF.

    Engineers can use a node-based workflow to run and customize VoxCPM2 voice generation locally.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor