Video Transcriber
Transcribe videos from 1000+ platforms or local files offline, and search across every transcript
Open source Open in the app JSON README (API)
About
Transcribe videos from 1000+ platforms or local files offline, and search across every transcript
Details
- Kind
- MCP servers
- Topic
- AI, RAG & memory
- Publisher
- nhatvu148
- Origin
- official
- Category
- ferramentas
- Transport
- local
- Version
- 0.10.6
- Stars
- 17
- Forks
- 5
- Last push
- 2026-09-03T02:03:50Z
- Repository state
- ativo
- Language
- Rust
- License
- Apache-2.0
- Added
- 2026-08-29 04:00:53
- Updated
- 2026-08-29 04:00:53
- Origin id
io.github.nhatvu148/video-transcriber-mcp
README
# Video Transcriber MCP ๐
**High-performance video transcription MCP server using whisper.cpp (Rust)**
[](#license)
[](https://www.rust-lang.org/)
[](https://crates.io/crates/video-transcriber-mcp)
A Model Context Protocol (MCP) server that transcribes videos from **1000+ platforms** using whisper.cpp. Built with Rust for maximum performance and efficiency.
## ๐ฆ Installation
### Homebrew (macOS/Linux) - Recommended
The easiest way to install with all dependencies:
```bash
brew install nhatvu148/tap/video-transcriber-mcp
```
This automatically installs the binary along with required dependencies (cmake, yt-dlp, ffmpeg).
### Cargo Install
If you have Rust installed:
```bash
cargo install video-transcriber-mcp
```
**Note:** You'll need to manually install dependencies: `yt-dlp`, `ffmpeg`, `cmake`
### Pre-built Binaries
Download from [GitHub Releases](https://github.com/nhatvu148/video-transcriber-mcp-rs/releases/latest):
```bash
# macOS (Intel)
curl -L https://github.com/nhatvu148/video-transcriber-mcp-rs/releases/latest/download/video-transcriber-mcp-x86_64-apple-darwin.tar.gz | tar xz
sudo mv video-transcriber-mcp /usr/local/bin/
# macOS (Apple Silicon)
curl -L https://github.com/nhatvu148/video-transcriber-mcp-rs/releases/latest/download/video-transcriber-mcp-aarch64-apple-darwin.tar.gz | tar xz
sudo mv video-transcriber-mcp /usr/local/bin/
# Linux (x86_64) โ no ARM64 Linux build, see issue #13; use `cargo install`
curl -L https://github.com/nhatvu148/video-transcriber-mcp-rs/releases/latest/download/video-transcriber-mcp-x86_64-unknown-linux-gnu.tar.gz | tar xz
sudo mv video-transcriber-mcp /usr/local/bin/
# Windows: Download .zip from releases page
```
**Note:** You'll need to manually install dependencies: `yt-dlp`, `ffmpeg`
### Claude Code plugin
Installs the MCP server and a `/transcribe` skill in one step:
```bash
/plugin marketplace add nhatvu148/video-transcriber-mcp-rs
/plugin install video-transcriber@nhatvu148-tools
```
The plugin registers the MCP server for you, but it does **not** install the binary โ run one of the install commands above first, so `video-transcriber-mcp` is on your `PATH`.
## ๐ฏ Why Rust?
This version uses **whisper.cpp** (C++ implementation with Rust bindings) instead of Python's OpenAI Whisper:
| Advantage | whisper.cpp (Rust) | OpenAI Whisper (Python) |
|-----------|-------------------|------------------------|
| **Performance** | Native C++ speed | Python interpreter overhead |
| **Memory** | Lower footprint | Higher memory usage |
| **Startup** | Instant (<100ms) | Slow (~2-3s model loading) |
| **Dependencies** | Standalone binary | Requires Python + packages |
| **Portability** | Single binary | Python environment needed |
Real-world performance depends on your hardware, video length, and chosen model.
## โจ Features
- ๐ **High performance** transcription using whisper.cpp (C++ with Rust bindings)
- ๐ฅ Download from **1000+ platforms** (YouTube, Vimeo, TikTok, Twitter, etc.)
- ๐ Transcribe **local video files** (mp4, avi, mov, mkv, etc.)
- ๐ค **100% offline** transcription (privacy-first)
- ๐๏ธ **5 model sizes** (tiny, base, small, medium, large)
- ๐ **90+ languages** supported
- ๐ **Multiple output formats** (TXT, JSON, Markdown)
- ๐ **MCP integration** for Claude Code
- ๐ **Dual transport** - stdio (local) and Streamable HTTP (remote)
- โก **Native binary** - no Python or Node.js required
- ๐พ **Low memory footprint** compared to Python implementations
## โก Quick Start (Using Taskfile)
**The fastest way to get started:**
```bash
# 1. Install Task (if not already installed)
brew install go-task/tap/go-task
# 2. Complete setup (build + download model)
task setup
# 3. Run a quick test
task test:quick
# Done! ๐
```
**Available Commands:**
```bash
task setup # Complete project setup
task test:quick # Test with short video
task benchmark # Run performance benchmark
task deps:check # Check dependencies
task download:base # Download base model
task help # Show all commands
```
See [Taskfile.yml](Taskfile.yml) for all available tasks.
---
## ๐ Transport Modes
The server supports two transport modes:
### Stdio Transport (Default)
Standard I/O transport for local CLI usage with Claude Code. This is the default mode.
```bash
video-transcriber-mcp
# or explicitly:
video-transcriber-mcp --transport stdio
```
### Streamable HTTP Transport
HTTP transport for remote access. Allows the MCP server to be accessed over the network.
```bash
# Start HTTP server on default port (8080)
video-transcriber-mcp --transport http
# Custom host and port
video-transcriber-mcp --transport http --host 0.0.0.0 --port 3000
```
**Remote MCP Client Configuration:**
For HTTP transport, configure your MCP client with the URL:
```json
{
"mcpServers": {
"video-transcriber-mcp": {
"url": "http://localhost:8080/mcp"
}
}
}
```
**Benefits of HTTP Transport:**
- No local installation required for clients
- Centralized server deployment
- Automatic updates (server-side)
- Better for team environments
- Compatible with serverless platforms
### CLI Options
```bash
video-transcriber-mcp --help
Options:
-t, --transport <TRANSPORT> Transport mode [default: stdio] [possible values: stdio, http]
--host <HOST> Host address for HTTP transport [default: 127.0.0.1]
-p, --port <PORT> Port for HTTP transport [default: 8080]
-h, --help Print help
-V, --version Print version
```
---
## ๐ฆ Manual Build from Source
### Prerequisites
1. **Rust** (1.85+ for Rust 2024 edition)
```bash
curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh
```
2. **yt-dlp** (for downloading videos)
```bash
# macOS
brew install yt-dlp
# Linux
pip install yt-dlp
# Windows
winget install yt-dlp.yt-dlp
```
3. **FFmpeg** (for audio processing)
```bash
# macOS
brew install ffmpeg
# Linux
sudo apt install ffmpeg # Debian/Ubuntu
sudo dnf install ffmpeg # Fedora
# Windows
choco install ffmpeg
```
### Build from Source
```bash
# Clone the repository
git clone https://github.com/nhatvu148/video-transcriber-mcp-rs.git
cd video-transcriber-mcp-rs
# Build the project
cargo build --release
# The binary will be at: target/release/video-transcriber-mcp-rs
```
### Download Whisper Models
```bash
# Download base model (recommended for testing)
bash scripts/download-models.sh base
# Or download all models
bash scripts/download-models.sh all
```
Models are stored in `~/.cache/video-transcriber-mcp/models/`
## ๐ Quick Start
### MCP Server (for Claude Code)
Add to `~/.claude/settings.json`:
**Option 1: If installed via GitHub Release or cargo install:**
```json
{
"mcpServers": {
"video-transcriber-mcp": {
"command": "video-transcriber-mcp",
"args": [],
"env": {
"RUST_LOG": "info"
}
}
}
}
```
**Option 2: If built from source:**
```json
{
"mcpServers": {
"video-transcriber-mcp": {
"command": "/absolute/path/to/video-transcriber-mcp-rs/target/release/video-transcriber-mcp",
"args": [],
"env": {
"RUST_LOG": "info"
}
}
}
}
```
Then use in Claude Code:
**Basic transcription (uses base model by default):**
```
Please transcribe this YouTube video: https://www.youtube.com/watch?v=VIDEO_ID
```
**Transcribe with specific model:**
```
Transcribe this video using the large model for best accuracy:
https://www.youtube.com/watch?v=VIDEO_ID
```
**Transcribe local video file:**
```
Transcribe this local video file: /Users/myname/Videos/meeting.mp4
```
**Transcribe in specific language:**
```
Transcribe this Spanish video: https://www.youtube.com/watch?v=VIDEO_ID
(language: es, model: medium)
```
## ๐ Performance
### Expected Performance Characteristics
Based on whisper.cpp vs OpenAI Whisper benchmarks from the community:
**Transcription Speed** (approximate, varies by hardware):
- whisper.cpp is typically **2-6x faster** than Python Whisper
- Faster startup time (no Python interpreter overhead)
- Lower memory footprint (no Python runtime)
**Real-world factors that affect performance:**
- CPU: More cores = faster processing
- Model size: Tiny is fastest, Large is slowest but most accurate
- Video length: Longer videos take proportionally more time
- Audio complexity: Clear speech transcribes faster than noisy audio
### Want to help?
We're collecting real benchmark data! If you run both versions, please share your results:
- Hardware specs (CPU, RAM)
- Video length tested
- Model used
- Time taken for each version
Open an issue with your benchmark results to help improve this section!
## ๐๏ธ Model Comparison
| Model | Speed | Accuracy | Memory | Use Case |
|-------|-------|----------|--------|----------|
| **tiny** | โกโกโกโกโก | โญโญ | ~400 MB | Quick drafts, testing |
| **base** | โกโกโกโก | โญโญโญ | ~600 MB | General use (default) |
| **small** | โกโกโก | โญโญโญโญ | ~1.2 GB | Better accuracy |
| **medium** | โกโก | โญโญโญโญโญ | ~2.5 GB | High accuracy |
| **large** | โก | โญโญโญโญโญโญ | ~4.8 GB | Best accuracy, slowest |
## ๐ Supported Platforms
Thanks to yt-dlp, this tool supports **1000+ video platforms** including:
- **Social Media**: YouTube, TikTok, Twitter/X, Facebook, Instagram, Reddit
- **Video Hosting**: Vimeo, Dailymotion, Twitch
- **Educational**: Coursera, Udemy, Khan Academy, edX
- **News**: BBC, CNN, NBC, PBS
- **And 1000+ more!**
## ๐ Output Format
For each video, three files are generated in `~/Downloads/video-transcripts/`:
```
video-id-title.txt # Plain text transcript
video-id-title.json # JSON with metadata and timestamps
video-id-title.md # Markdown with video info
```
### Example Output
```markdown
# How to Build Fast Software
**Video:** https://www.youtube.com/watch?v=example
**Platform:** YouTube
**Channel:** Tech Channel
**Duration:** 600s
---
## Transcript
The key to building fast software is understanding...
---
*Transcribed using whisper.cpp (Rust) - Model: base*
```
## ๐ง Configuration
### Environment Variables
All environment variables are optional. The transcriber works with none of them set; they unlock authentication, remote inference, AI summaries, and the paid HTTP API.
> ๐ก The transcript **output directory** is not an env var โ pass `output_dir` to the `transcribe_video` tool (defaults to `~/Downloads/video-transcripts`). Output files are named `<video_id>-<title>.{txt,json,md}`.
#### Remote MCP access (`--transport http`)
The HTTP transport only answers requests whose `Host` header is on an
allowlist. It defaults to loopback (`localhost`, `127.0.0.1`, `::1`) as
protection against [DNS rebinding][dns-rebinding], which means a deployed
instance rejects its own public hostname with `403` until you name it:
```bash
# Comma-separated. Added on top of the loopback defaults, so local
# development and health checks keep working.
export MCP_ALLOWED_HOSTS=mcp.example.com,mcp.example.com:8080
# On Fly:
fly secrets set MCP_ALLOWED_HOSTS=your-app.fly.dev
```
Leave it unset for local use โ the server logs which hosts it accepts at
startup, so a `403` from a remote client is easy to diagnose.
> โ ๏ธ This controls **reachability, not authorization**. Anyone who can reach
> the URL can call the tools, including `transcribe_video`, which spends real
> money when remote Whisper / OpenRouter are configured. Put an
> authenticating proxy in front of a public deployment.
#### Downloading (yt-dlp cookies)
Needed only for age-restricted / members-only videos or YouTube's "Sign in to confirm you're not a bot" challenge.
```bash
# Option 1 (preferred on headless / Linux): a Netscape-format cookies file.
# Export it however you like โ e.g. a QR-login flow โ then point at it.
export YT_DLP_COOKIES=/path/to/cookies.txt
# Option 2: read cookies straight from a logged-in local browser.
# One of: chrome, brave, edge, firefox, safari, chromium, opera, vivaldi.
# Ignored when YT_DLP_COOKIES is set.
export YT_DLP_COOKIES_FROM_BROWSER=chrome
```
#### Remote Whisper (offload transcription)
```bash
# POST audio to a remote HTTP worker (e.g. a serverless GPU) instead of
# running whisper-rs locally. Endpoint must accept multipart {audio, model,
# language} and return JSON {transcript, segments[], language, duration_s}.
export REMOTE_WHISPER_URL=https://your-worker.example.com/transcribe
```
## ๐งช Development
### Build
```bash
# Debug build
cargo build
# Release build (optimized)
cargo build --release
# Run tests
cargo test
# Run with logging
RUST_LOG=debug cargo run -- --url "https://youtube.com/watch?v=example"
```
### Project Structure
```
src/
โโโ main.rs # CLI + transport selection (stdio / streamable HTTP)
โโโ lib.rs # public API for embedders
โโโ mcp/ # MCP server: tool definitions and handlers
โโโ transcriber/ # the pipeline: yt-dlp โ ffmpeg โ whisper.cpp
โโโ embeddings.rs # passage embeddings, used by `search_transcripts`
โโโ utils/ # paths
```
This crate is only the transcription pipeline and its MCP surface. The product
built on top of it โ REST API, accounts, credits, payments, AI summaries and
diagrams โ lives in a separate private crate that depends on this one as a
library, so `cargo install video-transcriber-mcp` gets you a transcription
server rather than somebody else's SaaS backend.
## ๐ค Contributing
Contributions welcome! Please:
1. Fork the repository
2. Create a feature branch
3. Make your changes
4. Add tests if applicable
5. Submit a pull request
## ๐ Acknowledgments
- [whisper.cpp](https://github.com/ggerganov/whisper.cpp) - Fast C++ implementation of Whisper
- [whisper-rs](https://codeberg.org/tazz4843/whisper-rs) - Rust bindings for whisper.cpp
- [yt-dlp](https://github.com/yt-dlp/yt-dlp) - Video downloader for 1000+ platforms
- [OpenAI Whisper](https://github.com/openai/whisper) - Original speech recognition model
- [Model Context Protocol SDK](https://github.com/modelcontextprotocol/rust-sdk) - Rust SDK for MCP
## ๐ Comparison with TypeScript Version
I built the original [video-transcriber-mcp](https://github.com/nhatvu148/video-transcriber-mcp) in TypeScript. Here's why I rewrote it in Rust:
| Aspect | TypeScript Version | **Rust Version** |
|--------|-------------------|------------------|
| Transcription Speed | 5 min for 10-min video | **50s (6x faster)** |
| Memory Usage | ~2 GB | **~800 MB (2.5x less)** |
| Startup Time | ~2s | **<100ms (20x faster)** |
| Binary Size | N/A (Node.js runtime) | **~8 MB standalone** |
| Dependencies | Node.js, Python, whisper | **Just yt-dlp, ffmpeg** |
| CPU Usage | High (Python overhead) | **Lower (native code)** |
**The Rust version is production-ready and significantly more efficient!**
## ๐ Links
- [GitHub Repository](https://github.com/nhatvu148/video-transcriber-mcp-rs)
- [TypeScript Version](https://github.com/nhatvu148/video-transcriber-mcp)
- [Model Context Protocol](https://modelcontextprotocol.io)
- [whisper.cpp](https://github.com/ggerganov/whisper.cpp)
## License
Licensed under either of
- MIT license ([LICENSE-MIT](LICENSE-MIT))
- Apache License, Version 2.0 ([LICENSE-APACHE](LICENSE-APACHE))
at your option.
## Contribution
Unless you explicitly state otherwise, any contribution intentionally submitted
for inclusion in the work by you, as defined in the Apache-2.0 license, shall be
dual licensed as above, without any additional terms or conditions.
---
**Built with โค๏ธ in Rust for maximum performance**
<sub>MCP registry ownership token โ crates.io strips HTML comments, so this line has to stay visible:</sub>
mcp-name: io.github.nhatvu148/video-transcriber-mcp