Skip to content
EN

Speech: TTS and ASR

Text-to-speech, speech recognition and voice models you can run yourself.

73 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: Speech: TTS and ASR

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. Qwen3-TTS announced as open source

    The post announces that Qwen3-TTS is being open-sourced and mentions voice cloning and voice design. It does not include a repository or other link.

    Engineers evaluating self-hosted TTS may want to investigate the model and its voice features.

  2. Step Audio R1.1 for Real-Time Speaking Dialogue

    The post describes Step Audio R1.1 as a real-time, interactive speaking dialogue model under Apache 2.0. It claims the model can think while speaking and uses a dual-brain architecture for low latency.

    Engineers can inspect the model page and evaluate its license and architecture claims for speech applications.

  3. Pocket TTS: 100M-Parameter TTS for Laptops

    Kyutai Labs introduces Pocket TTS, a 100M-parameter text-to-speech model that the post says supports voice cloning and runs on a laptop without a GPU. The post describes it as open-source.

    Engineers can evaluate a self-hostable TTS model designed to run without a GPU.

  4. LEMAS releases a 150K-hour multilingual speech dataset and models

    LEMAS describes a 150K-hour multilingual audio suite and generative speech models. The post says the dataset covers 10 languages with word-level timestamps and names LEMAS-TTS and LEMAS-Edit.

    Engineers can explore speech data and models for multilingual TTS and speech editing.

  5. Meta releases Omnilingual ASR for 1,600+ languages

    Meta’s Omnilingual ASR is an open-source speech recognition model suite and dataset for more than 1,600 languages, including 500 described as previously unsupported.

    Engineers can explore an open-source ASR system aimed at broad multilingual coverage.

  6. Step-Audio-EditX for Prompt-Based Audio Editing

    StepFun says Step-Audio-EditX is a 3B-parameter model for audio editing, zero-shot multilingual TTS, and prompt-based control of emotion, speaking style, and vocal elements such as breaths and laughs. The post says it supports single-GPU deployment and uses Apache 2.0.

    Engineers can evaluate a single-model approach to speech generation and iterative audio editing.

  7. Inworld TTS 1 Max leads the Artificial Analysis Speech Arena

    Artificial Analysis reports that Inworld TTS 1 Max leads its Speech Arena leaderboard, where users compare generated speech and choose their preferred output without seeing the model names. Inworld says TTS Max costs $10 per million characters.

    The leaderboard describes a human-preference evaluation across four real-world speech categories.

  8. Inworld TTS 1 Max leads the Artificial Analysis Speech Arena

    The post reports that Inworld TTS 1 Max leads the Speech Arena leaderboard, which ranks TTS models through blind human comparisons. It describes both Inworld models’ 12-language support, short-audio voice cloning, and voice tags.

    The comparison and feature details may help engineers assess TTS options.

  9. LongCat-Audio-Codec for Speech LLMs

    Meituan open-sourced an audio tokenizer and detokenizer optimized for Speech LLMs. The post describes parallel semantic and acoustic tokens at 16.7 Hz, low-bitrate encoding, and a low-latency streaming decoder.

    Engineers can evaluate its audio tokenization and streaming decoder for Speech LLM applications.

  10. DiaMoE-TTS uses IPA and MoE for dialect speech synthesis

    DiaMoE-TTS is an IPA-based dialect TTS framework with a dialect-aware Mixture-of-Experts and LoRA and Conditioning Adapters for adaptation. The post reports zero-shot synthesis on unseen dialects, including Peking Opera, with a few hours of data.

    Its phonetic representation and adaptation approach may help engineers build TTS systems for dialects with limited data.

  11. Gemini 2.5 Native Audio Thinking on Big Bench Audio

    Artificial Analysis reports that Gemini 2.5 Native Audio Thinking scored 92% on its Big Bench Audio benchmark, which uses 1,000 audio questions adapted from Big Bench Hard. It reports 3.87 seconds average time to first token for the thinking model and 0.63 seconds for its non-thinking equivalent.

    The results offer a reasoning benchmark and latency comparison for speech-to-speech models.

  12. FireRedChat: Open-Source Full-Duplex Voice Interaction

    FireRedChat is a pluggable voice interaction system with cascaded and semi-cascaded implementations. The post says it supports interruption handling, endpoint detection, real-time responses, emotion perception, and expressive speech synthesis.

    Engineers can evaluate its architecture and voice interaction features for conversational assistants.

  13. Qwen3-Omni open-sources multimodal models with speech I/O

    Qwen says it has open-sourced three Qwen3-Omni models for instruction following, reasoning, and captioning. The model combines text, image, audio, and video, with speech input and output.

    Engineers can explore an open-source model that handles speech alongside other modalities.

  14. Mini-Omni-Reasoner interleaves reasoning and speech

    Mini-Omni-Reasoner uses a hierarchical Thinker–Talker architecture to interleave silent reasoning with spoken tokens. The post reports a new Spoken-Math-Problems-3M dataset and gains in arithmetic reasoning and contextual understanding.

    The approach and reported results may inform designs for speech models that reason while speaking.

  15. FireRedTTS 2 open-source text-to-speech model

    FireRedTTS 2 is an open TTS model on Hugging Face. The post describes long-form conversational speech, multilingual and cross-lingual voice cloning, low-latency output, and random timbre generation.

    Engineers can evaluate a self-hostable TTS model for multilingual speech generation and voice cloning.

  16. Qwen3-ASR-Toolkit handles long-audio transcription

    Qwen3-ASR-Toolkit is a Python CLI for the Qwen3-ASR API. It supports long audio and video transcription using VAD splitting, parallel API calls, multiple media formats, and automatic resampling.

    It provides a way to process long recordings through the Qwen3-ASR API with parallel calls and media-format support.

  17. FireRedTTS-2 Uses Streaming Tokens for Multi-Speaker Dialogue

    FireRedTTS-2 is described as a long-form streaming TTS system using a 12.5 Hz speech tokenizer, a dual-transformer architecture, and interleaved text–speech input for context-aware multi-speaker dialogue. The post claims it outperforms several systems on intelligibility, speaker turns, and…

    Its streaming design and multi-speaker format may be relevant to engineers building conversational speech generation.

  18. Deploy streaming Kyutai STT on Modal

    Modal provides an example for deploying a streaming audio transcription service with Kyutai STT.

    It offers engineers a concrete reference for running streaming speech recognition.

  19. Zonos TTS Beta and Cross-Platform Local Launcher

    Zonos announced a beta expressive TTS model with voice cloning, releasing transformer and SSM-hybrid models under Apache 2.0. The post also presents a one-click launcher for Mac, Windows, and Linux.

    Engineers can explore a locally runnable voice-cloning TTS model and its available model architectures.

  20. Hibiki: a model for streaming speech translation

    Hibiki is a decoder-only model for simultaneous French–English speech translation. Its synthetic training corpus uses translations from MADLAD, the authors’ TTS, and a simple lag rule.

    The repository offers engineers a model for streaming speech translation to evaluate and run.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor