Skip to content
EN

Speech: TTS and ASR

Text-to-speech, speech recognition and voice models you can run yourself.

73 links, newest first.

Get the weekly briefing

The best new links of the topics you pick, summarized with the source. At most one email a week.

Topics: Speech: TTS and ASR

Before the first issue we email you to confirm; leaving takes one click. Sent with CommsHarbor. Privacy

  1. Moonshine ASR model added to Transformers

    The post announces Moonshine, an ASR model integrated into Hugging Face Transformers. It says the model uses Rotary Position Embeddings and can handle audio inputs of any length.

    Engineers can evaluate another ASR model for processing long audio inputs.

  2. Piper is a local neural text-to-speech system

    Piper converts text into speech using various voices and phonemization libraries. It supports multiple languages and customizable options.

    Engineers can evaluate a local TTS option with multilingual support and configurable settings.

  3. Ultravox: a multimodal LLM for real-time voice

    Ultravox is a fast multimodal LLM for real-time voice. Its source code is available on GitHub.

    Engineers can inspect the project and evaluate its approach to real-time voice applications.

  4. e2-f5-tts Gradio App Supports Custom F5-TTS Finetunes

    The official e2-f5-tts Gradio app now supports custom F5-TTS finetunes. The post mentions finetunes for Spanish, Hungarian, German, Portuguese, and Greek.

    Engineers can use the app to try custom multilingual F5-TTS voice models.

  5. MaskGCT: Zero-Shot Text-to-Speech with a Masked Codec Transformer

    MaskGCT is a zero-shot text-to-speech model introduced in a paper and released by Amphion. The post claims it can clone a voice from five seconds of speech.

    Engineers can explore a released approach to zero-shot TTS and its voice-cloning requirements.

  6. E2-F5-TTS adds a Gradio fine-tuning UI

    The E2-F5-TTS project adds a Gradio app for fine-tuning the model with user-provided audio. It automatically transcribes audio clips.

    Engineers can explore a UI-based workflow for adapting a TTS model to their own audio.

  7. Hugging Face Speech-to-Speech voice agent toolkit

    A GitHub project for building voice agents with open-source models.

    Engineers can explore a project for assembling voice agents from open-source models.

  8. French Parler-TTS Mini Model on Hugging Face

    The post points to PHBJT’s French Parler-TTS Mini model hosted on Hugging Face. It also mentions XTTS and FishAudio, and speculates about a future multilingual Parler-TTS.

    Engineers exploring French text-to-speech models can inspect the linked model page.

  9. Moshi: an open-source speech-to-speech model

    Moshi is a 7.6B-parameter speech-to-speech model, accompanied by Mimi, a streaming audio codec. The post says the release includes model checkpoints and inference code for Candle, PyTorch, and MLX.

    Engineers can evaluate the model and run inference with different frameworks and hardware.

  10. VoiceCraft for Zero-Shot Speech Editing and TTS

    VoiceCraft is a neural codec language model for speech editing and zero-shot text-to-speech. The post says it can clone or edit an unseen voice using a few seconds of reference audio.

    Engineers can evaluate its approach to voice cloning and speech editing on varied audio.

  11. Parler-TTS Mini and Large v1 Released

    Parler-TTS Mini and Large v1 are open-source text-to-speech models trained on 45K hours of speech data. Speech characteristics such as gender, background noise, rate, pitch, and reverberation can be controlled with text prompts.

    Engineers can evaluate controllable TTS models with released weights, training data, and recipes.

  12. XEUS speech encoder released with code and checkpoints

    WavLab announced XEUS, a self-supervised speech encoder trained on over 1 million hours of speech and covering more than 4,000 languages. The post says its code, checkpoints, and data are being released.

    Engineers can evaluate a multilingual speech encoder and its accompanying data for their own systems.

  13. Speech-to-Text Provider Comparison

    Artificial Analysis compares speech-to-text models by word error rate, speed, and price, with a leaderboard for choosing a model by use case.

    Engineers can compare ASR quality, latency, and cost when evaluating providers.

Build with AgentLog

List your MCP, skill or plugin

Reach the engineers who read these briefings.

Sponsor AgentLog

Footer, sidebar or featured slot for 30 days.

From US$ 60

See the slots

Send your own newsletter

CommsHarbor keeps contacts, consent and one-click unsubscribe together.

Free workspace

Open CommsHarbor