{
  "markdown": "# Voicebox\n\nAll-in-one voice toolkit for Claude Code on Apple Silicon Macs.\nApple Silicon Mac 上的 Claude Code 一站式语音工具包。\n\n## Features / 功能\n\n- **Voice Design / 语音设计** — Create custom voices from text descriptions / 用文字描述创建自定义声音\n- **Voice Cloning / 语音克隆** — Clone any voice from a 10s audio sample / 录 10 秒音频克隆任意声音\n- **Text-to-Speech / 文字转语音** — Generate speech in 10 languages / 支持 10 种语言的语音合成\n- **Multi-Speaker Conversations / 多角色对话** — Generate dialogues, dramas, news broadcasts / 生成对话、戏剧、新闻播报\n- **Transcription / 语音转文字** — Speech-to-text for audio & video, 52 languages with auto-detection / 音频视频转文字，52 种语言自动识别\n\n## Install / 安装\n\n```bash\ngit clone https://github.com/tivojn/voicebox.git ~/.claude/skills/voicebox\n```\n\nThat's it. Dependencies and models auto-download on first use.\n就这样。依赖和模型首次使用时自动下载。\n\n## Requirements / 环境要求\n\n- Apple Silicon Mac (M1/M2/M3/M4)\n- [Claude Code](https://claude.com/claude-code)\n- [uv](https://docs.astral.sh/uv/) (usually pre-installed with Claude Code / 通常随 Claude Code 预装)\n- ffmpeg — for recording & video transcription / 录音和视频转写需要 (`brew install ffmpeg`)\n\n## Quick Start / 快速开始\n\n```\n/voicebox create a calm narrator voice profile          # Design a voice from description\n/voicebox \"Calm Narrator\" \"Hello, this is a test.\"      # Generate speech\n/voicebox clone my voice                                # Record from mic & clone\n/voicebox clone my voice from /path/to/audio.wav        # Clone from audio file\n/voicebox transcribe /path/to/audio.wav                 # Transcribe audio/video\n/voicebox create a news broadcast with anchor and reporter  # Multi-speaker conversation\n```\n\n## Voice Profiles / 语音档案\n\nProfiles are stored in `data/profiles.json`. Two types:\n\n- **Designed** — Created from a text description (supports `--instruct` style overrides)\n- **Cloned** — Created from a real audio sample (reproduces the tone/energy of the original recording)\n\nStart with no profiles — create your own with `/voicebox create ...` or `/voicebox clone ...`.\n\n## Quality Options / 质量选项\n\nAll commands default to **high** (1.7B models). Use `--quality standard` for faster 0.6B models if needed.\n\n| Category / 类别 | High (default) | Standard |\n|-----------------|---------------|----------|\n| Voice Design / 语音设计 | 1.7B | (same) |\n| Voice Clone / 语音克隆 | 1.7B (~3.5GB) | 0.6B (~1.5GB) |\n| Transcription / 语音转文字 | 1.7B (~3.5GB) | 0.6B (~1.5GB) |\n\n## Recording / 录音\n\nWhen recording from the mic, the script:\n- Records for the specified duration (default 10 seconds)\n- Prints `DONE — recording complete!` and plays a system bell when finished\n- Auto-transcribes the recording using Qwen3-ASR\n- Creates a cloned voice profile automatically\n\n## Supported Languages / 支持语言\n\n**TTS:** English, Chinese, Japanese, Korean, German, French, Russian, Portuguese, Spanish, Italian\n\n**ASR:** 52 languages with auto-detection / 52 种语言自动识别\n\n## Script Commands / 命令参考\n\n```bash\n# List profiles\nuv run ~/.claude/skills/voicebox/scripts/voicebox.py list\n\n# Generate speech\nuv run ~/.claude/skills/voicebox/scripts/voicebox.py generate \"Profile\" \"text\" --play\n\n# Generate with standard quality (faster, less RAM)\nuv run ~/.claude/skills/voicebox/scripts/voicebox.py generate \"Profile\" \"text\" --quality standard --play\n\n# Create designed voice\nuv run ~/.claude/skills/voicebox/scripts/voicebox.py create-designed \"Name\" --desc \"description\" --lang en\n\n# Clone from audio file\nuv run ~/.claude/skills/voicebox/scripts/voicebox.py create-cloned \"Name\" --audio /path/to.wav --ref-text \"transcript\" --lang en\n\n# Record from mic and clone\nuv run ~/.claude/skills/voicebox/scripts/voicebox.py record \"Name\" --duration 10 --lang en\n\n# Transcribe audio/video\nuv run ~/.claude/skills/voicebox/scripts/transcribe.py /path/to/file.wav\n\n# Multi-speaker conversation from JSON script\nuv run ~/.claude/skills/voicebox/scripts/voicebox.py conversation /tmp/script.json --play\n\n# Delete a profile\nuv run ~/.claude/skills/voicebox/scripts/voicebox.py delete \"Name\"\n```\n",
  "bytes": 3954,
  "sha": "3794685ad7ae6ac1979abef2d19052d5191a6666fc6046e6f7b3e2f21b11a748",
  "repo_slug": "tivojn/voicebox",
  "fonte": "repo",
  "truncated": false,
  "api": "https://api.agentalog.com/api/listings/skl_tivojn_voicebox_voicebox_7a794820/readme"
}