Speech: TTS and ASR
Text-to-speech, speech recognition and voice models you can run yourself.
73 links, newest first.
- Speech: TTS and ASRPost on X
Moonshine ASR model added to Transformers
The post announces Moonshine, an ASR model integrated into Hugging Face Transformers. It says the model uses Rotary Position Embeddings and can handle audio inputs of any length.
Engineers can evaluate another ASR model for processing long audio inputs.
- Speech: TTS and ASRPost on X
Piper is a local neural text-to-speech system
Piper converts text into speech using various voices and phonemization libraries. It supports multiple languages and customizable options.
Engineers can evaluate a local TTS option with multilingual support and configurable settings.
- Speech: TTS and ASRRepository
Ultravox: a multimodal LLM for real-time voice
Ultravox is a fast multimodal LLM for real-time voice. Its source code is available on GitHub.
Engineers can inspect the project and evaluate its approach to real-time voice applications.
- Speech: TTS and ASRPost on X
e2-f5-tts Gradio App Supports Custom F5-TTS Finetunes
The official e2-f5-tts Gradio app now supports custom F5-TTS finetunes. The post mentions finetunes for Spanish, Hungarian, German, Portuguese, and Greek.
Engineers can use the app to try custom multilingual F5-TTS voice models.
- Speech: TTS and ASRPaper
MaskGCT: Zero-Shot Text-to-Speech with a Masked Codec Transformer
MaskGCT is a zero-shot text-to-speech model introduced in a paper and released by Amphion. The post claims it can clone a voice from five seconds of speech.
Engineers can explore a released approach to zero-shot TTS and its voice-cloning requirements.
- Speech: TTS and ASRPost on X
E2-F5-TTS adds a Gradio fine-tuning UI
The E2-F5-TTS project adds a Gradio app for fine-tuning the model with user-provided audio. It automatically transcribes audio clips.
Engineers can explore a UI-based workflow for adapting a TTS model to their own audio.
- Speech: TTS and ASRRepository
Hugging Face Speech-to-Speech voice agent toolkit
A GitHub project for building voice agents with open-source models.
Engineers can explore a project for assembling voice agents from open-source models.
- Speech: TTS and ASRModel
French Parler-TTS Mini Model on Hugging Face
The post points to PHBJT’s French Parler-TTS Mini model hosted on Hugging Face. It also mentions XTTS and FishAudio, and speculates about a future multilingual Parler-TTS.
Engineers exploring French text-to-speech models can inspect the linked model page.
- Speech: TTS and ASRPost on X
Moshi: an open-source speech-to-speech model
Moshi is a 7.6B-parameter speech-to-speech model, accompanied by Mimi, a streaming audio codec. The post says the release includes model checkpoints and inference code for Candle, PyTorch, and MLX.
Engineers can evaluate the model and run inference with different frameworks and hardware.
- Speech: TTS and ASRPost on X
VoiceCraft for Zero-Shot Speech Editing and TTS
VoiceCraft is a neural codec language model for speech editing and zero-shot text-to-speech. The post says it can clone or edit an unseen voice using a few seconds of reference audio.
Engineers can evaluate its approach to voice cloning and speech editing on varied audio.
- Speech: TTS and ASRPost on X
Parler-TTS Mini and Large v1 Released
Parler-TTS Mini and Large v1 are open-source text-to-speech models trained on 45K hours of speech data. Speech characteristics such as gender, background noise, rate, pitch, and reverberation can be controlled with text prompts.
Engineers can evaluate controllable TTS models with released weights, training data, and recipes.
- Speech: TTS and ASRPost on X
XEUS speech encoder released with code and checkpoints
WavLab announced XEUS, a self-supervised speech encoder trained on over 1 million hours of speech and covering more than 4,000 languages. The post says its code, checkpoints, and data are being released.
Engineers can evaluate a multilingual speech encoder and its accompanying data for their own systems.
- Speech: TTS and ASRArticle
Speech-to-Text Provider Comparison
Artificial Analysis compares speech-to-text models by word error rate, speed, and price, with a leaderboard for choosing a model by use case.
Engineers can compare ASR quality, latency, and cost when evaluating providers.
Build with AgentLog
Send your own newsletter
CommsHarbor keeps contacts, consent and one-click unsubscribe together.
Free workspace
Open CommsHarbor


