voicebox-cli
Knowledge for voicebox-cli, a command-line client and TypeScript library for the voicebox text-to-speech and transcription API.
Open source Repository Open in the app JSON README (API)
About
# voicebox-cli
`voicebox-cli` is a small TypeScript command-line client for the
[voicebox](http://127.0.0.1:17493) text-to-speech API. It generates speech from
text, waits for the asynchronous job to finish, downloads the audio, and
optionally transcodes it to MP3; it also transcribes audio back to text. The
same code ships as a programmatic library and as a bundled agent skill.
- Package: `voicebox-cli` (see [package.json](../package.json))
- Binary: `voicebox-cli` → [dist/cli.js](../src/cli.ts)
- Default server: `http://127.0.0.1:17493`
- Runtime: Node.js >= 20, ES modules, run via `tsx` in development
## Folders
- [cli_commands](./cli_commands/index.md) — every `voicebox-cli` subcommand, its
flags, and the API calls it makes.
- [client_library](./client_library/index.md) — the reusable `VoiceboxClient`
HTTP client and its audio/CLI helpers.
- [data_models](./data_models/index.md) — the request and response shapes
exchanged with the API.
- [api_reference](./api_reference/index.md
Details
- Kind
- OKF bundles
- Topic
- AI, RAG & memory
- Publisher
- jeromeetienne
- Origin
- okf_github
- Category
- dados
- Version
- 0.1
- Stars
- 1
- Last push
- 2026-07-02T07:31:48Z
- Repository state
- ativo
- Language
- TypeScript
- License
- MIT
- Added
- 2026-09-09 05:03:59
- Updated
- 2026-09-09 05:03:59
- Origin id
jeromeetienne/voicebox_cli:.okf/index.md
README
# voicebox-cli
A small TypeScript command-line client for the [voicebox](http://127.0.0.1:17493) TTS API. It generates speech from text, waits for the async job to finish, downloads the audio, and optionally transcodes it to MP3 or Opus.
## Requirements
- Node.js 20+ (developed on v24)
- A running voicebox API (defaults to `http://127.0.0.1:17493`)
`ffmpeg` is **not** required on your system — the bundled [`ffmpeg-static`](https://www.npmjs.com/package/ffmpeg-static) binary is used for transcoding.
## Install
Run it directly with `npx`:
```bash
npx voicebox-cli speak "Hello, world!" --profile Test --output hello.mp3
```
Or install it globally:
```bash
npm install -g voicebox-cli
voicebox-cli speak "Hello, world!" --profile Test --output hello.mp3
```
## Usage
Once installed, the `voicebox-cli` binary is available. Run any command through `npx`:
```bash
npx voicebox-cli speak "Hello, world!" --profile Test --output outputs/hello.mp3
```
### `speak`
Synthesize speech from text with a chosen voice profile and save it to a file. The command submits the text to the API, waits for the asynchronous generation to complete, downloads the resulting audio, and writes it locally. Pick the voice with `--profile`, steer synthesis with `--language` and `--engine`, and use `--personality` to rewrite the text in the profile's character before it is spoken.
```
npx voicebox-cli speak <text> [options]
Options:
-p, --profile <profile> voice profile name or id
-o, --output <path> output file (.mp3 or .wav) (default: speech.mp3)
-e, --engine <engine> TTS engine
-l, --language <language> language code (e.g. en, fr, ja)
--personality rewrite the text in-character before TTS
--base-url <url> API base url
-h, --help display help for command
```
The output format is chosen from the file extension: `.mp3` transcodes via `ffmpeg-static`, anything else writes the raw WAV returned by the API.
Examples:
```bash
# Simplest: text + voice profile → speech.mp3 (the default output)
npx voicebox-cli speak "Hello there" --profile Test
# Save to a specific MP3 file
npx voicebox-cli speak "Hello there" --profile Test --output outputs/hello.mp3
# Save as WAV instead (any non-.mp3 extension writes raw WAV)
npx voicebox-cli speak "Hello there" --profile Test --output outputs/hello.wav
# French, with a French voice
npx voicebox-cli speak "Bonjour tout le monde !" --profile manukipu --language fr
# Rewrite the text in the profile's character before speaking
npx voicebox-cli speak "Tell me about your day." --profile donaldy --personality
# Pick a specific engine
npx voicebox-cli speak "Testing the kokoro engine." --profile Test --engine kokoro
# Point at a server on another host/port
npx voicebox-cli speak "Remote server test." --profile Test --base-url http://192.168.1.50:17493
# Using short flags
npx voicebox-cli speak "Short and sweet." -p Test -o outputs/quick.mp3
```
### `generate`
The low-level counterpart to `speak`. It targets a profile by **id** (not name) and exposes the full generation request: seed, instruction/style prompt, model size, engine, chunking for long text, crossfade, and volume normalization. It also manages the lifecycle of an existing generation — retry a failed one, regenerate from scratch, cancel an in-progress job, or wait on its status.
```
npx voicebox-cli generate run <profile-id> <text> [options]
npx voicebox-cli generate retry <id>
npx voicebox-cli generate regenerate <id>
npx voicebox-cli generate cancel <id>
npx voicebox-cli generate status <id>
run options:
-o, --output <path> output file (.mp3 or .wav) (default: outputs/generation.mp3)
-l, --language <code> language code (default: en)
--seed <n> random seed
--model-size <size> model size (e.g. 1.7B)
--instruct <text> instruction / style prompt
-e, --engine <engine> TTS engine
--personality rewrite the text in-character before TTS
--max-chunk-chars <n> max characters per chunk for long text
--crossfade-ms <n> crossfade between chunks in ms
--no-normalize do not normalize output volume
--base-url <url> API base url
```
```bash
npx voicebox-cli generate run <profile-id> "A precise, reproducible take." --seed 42 -o outputs/take.wav
```
### `profiles`
Create, inspect, and delete the voice profiles that `speak` uses, and manage the reference samples a cloned voice is built from. A profile bundles a voice's language, engine defaults, and an optional personality prompt; samples are short audio clips plus their transcripts that teach the clone how the voice sounds. `update` merges your changes with the profile's current values, so you only pass the fields you want to change.
```
npx voicebox-cli profiles <subcommand> [options]
Subcommands:
list list all profiles
get <id> show a single profile (JSON)
create <name> [options] create a profile
update <id> [options] update a profile (merges with current values)
delete <id> delete a profile
presets <engine> list preset voices for an engine
export <id> [-o <path>] export a profile to a file (default: outputs/profile.zip)
samples list <profile-id> list a profile's reference samples
samples add <profile-id> <file> <reference-text> add a sample from an audio file
samples update <sample-id> <reference-text> change a sample's transcript
samples delete <sample-id> delete a sample
```
`create` / `update` options:
```
-n, --name <name> profile name (update only)
-d, --description <text> description
-l, --language <code> language code (default: en)
--voice-type <type> voice type (e.g. cloned)
--preset-engine <engine> preset engine
--preset-voice-id <id> preset voice id
--design-prompt <text> voice design prompt
--default-engine <engine> default TTS engine
--personality <text> in-character personality prompt
--base-url <url> API base url
```
```bash
# create a profile, then clone a voice into it from a reference clip
npx voicebox-cli profiles create "Narrator" --language en --personality "calm and warm"
npx voicebox-cli profiles samples add <profile-id> sample.wav "This is my reference voice."
# list profiles, then generate with one
npx voicebox-cli profiles list
npx voicebox-cli speak "Hello there" --profile Narrator -o outputs/hello.mp3
```
### `channels`
Manage audio output channels and the voices assigned to them. A channel is a named output route that binds a set of audio devices to a set of voice profiles, letting the server play different voices through different speakers. Use these subcommands to create channels, attach output devices, and control which profiles belong to each one.
```
npx voicebox-cli channels <subcommand> [options]
Subcommands:
list list all channels
get <id> show a single channel (JSON)
create <name> [--device <id...>] create a channel
update <id> [-n <name>] [--device <id...>] update a channel
delete <id> delete a channel
voices <id> list profiles assigned to a channel
set-voices <id> <profile-ids...> assign profiles to a channel
```
```bash
npx voicebox-cli channels create "Living room" --device dev-1 --device dev-2
npx voicebox-cli channels set-voices <channel-id> <profile-id-a> <profile-id-b>
```
### `history`
Browse and manage past generations. `list` supports filtering by profile and free-text search with pagination; `get` and `stats` inspect a single item or aggregate totals; `favorite`, `delete`, and `clear-failed` manage entries; and `export` / `export-audio` save a generation's archive or audio to disk.
```
npx voicebox-cli history list [-p <profile-id>] [-s <search>] [--limit <n>] [--offset <n>]
npx voicebox-cli history get <id>
npx voicebox-cli history stats
npx voicebox-cli history favorite <id>
npx voicebox-cli history delete <id>
npx voicebox-cli history clear-failed
npx voicebox-cli history export <id> [-o <path>] # zip (default: outputs/<id>.zip)
npx voicebox-cli history export-audio <id> [-o <path>] # wav (default: outputs/<id>.wav)
```
```bash
npx voicebox-cli history list --profile <profile-id> --search "hello" --limit 20
npx voicebox-cli history export-audio <id> -o outputs/take.wav
```
### `models`
Manage the TTS models the server can use. `status` lists every model with its download and loaded state; `load` and `unload` control what sits in memory; `download`, `download-wait`, `cancel-download`, and `delete` manage what is on disk; `cache-dir` shows where models are stored; and `progress`, `migrate`, and `migrate-progress` stream live progress from download and directory-migration tasks.
```
npx voicebox-cli models status
npx voicebox-cli models load [size] # size defaults to the server default
npx voicebox-cli models unload [name] # no name unloads the default model
npx voicebox-cli models download <name> # triggers the download, returns immediately
npx voicebox-cli models download-wait <name> # downloads, shows progress, blocks until done
npx voicebox-cli models cancel-download <name>
npx voicebox-cli models delete <name>
npx voicebox-cli models cache-dir
npx voicebox-cli models progress <name> # streams SSE progress
npx voicebox-cli models migrate <destination> # streams SSE progress
npx voicebox-cli models migrate-progress # streams SSE progress
```
```bash
npx voicebox-cli models status
npx voicebox-cli models load 1.7B
npx voicebox-cli models download qwen-1.7b
# Download a transcription model and wait until it is ready to use
npx voicebox-cli models download-wait whisper-large
```
### `stories`
Assemble multi-clip stories from existing generations on a timeline and export them as one mixed audio file. The top-level subcommands manage stories themselves (`list`, `get`, `create`, `update`, `delete`, `export-audio`); the `stories items` subgroup manages the clips on a story's timeline — adding, removing, reordering, moving, trimming, adjusting per-clip volume, splitting, duplicating, and pinning a clip to a specific generation version.
```
npx voicebox-cli stories list
npx voicebox-cli stories get <id>
npx voicebox-cli stories create <name> [-d <description>]
npx voicebox-cli stories update <id> <name> [-d <description>]
npx voicebox-cli stories delete <id>
npx voicebox-cli stories export-audio <id> [-o <path>] # wav (default: outputs/<id>.wav)
npx voicebox-cli stories items add <story-id> <generation-id> [--start-time-ms <n>] [--track <n>]
npx voicebox-cli stories items remove <story-id> <item-id>
npx voicebox-cli stories items times <story-id> <gen-id:ms>... # e.g. g1:0 g2:2500
npx voicebox-cli stories items reorder <story-id> <generation-id>...
npx voicebox-cli stories items move <story-id> <item-id> <start-time-ms> [--track <n>]
npx voicebox-cli stories items trim <story-id> <item-id> <trim-start-ms> <trim-end-ms>
npx voicebox-cli stories items volume <story-id> <item-id> <volume> # linear gain 0.0-2.0
npx voicebox-cli stories items split <story-id> <item-id> <split-time-ms>
npx voicebox-cli stories items duplicate <story-id> <item-id>
npx voicebox-cli stories items version <story-id> <item-id> [version-id] # omit to clear the pin
```
Examples:
```bash
# Create a story and add two generations to its timeline
npx voicebox-cli stories create "Chapter 1" --description "The opening scene"
npx voicebox-cli stories items add <story-id> <generation-id>
npx voicebox-cli stories items add <story-id> <generation-id> --start-time-ms 3000
# Reorder by generation id, then export the mixed audio
npx voicebox-cli stories items reorder <story-id> <gen-a> <gen-b>
npx voicebox-cli stories export-audio <story-id> -o outputs/chapter1.wav
# Fine-tune a single clip
npx voicebox-cli stories items volume <story-id> <item-id> 0.8
npx voicebox-cli stories items trim <story-id> <item-id> 250 100
```
### `transcribe`
Transcribe an audio file to text. Pass the file path and, optionally, a language hint and a transcription model. By default it prints just the transcript; add `--json` to get the raw response including the audio duration. Non-WAV inputs (MP3, Opus, FLAC, and anything else ffmpeg can read) are transcoded to WAV locally before upload.
```
npx voicebox-cli transcribe <file> [options]
Options:
-l, --language <language> language hint (one of: zh, en, ja, ...; pass 'list' to see options)
-m, --model <model> transcription model (one of: whisper-base, whisper-small, whisper-medium, whisper-large, whisper-turbo; pass 'list' to see options)
--json print the raw JSON response
--base-url <url> API base url
-h, --help display help for command
```
Examples:
```bash
# Simplest: audio file → transcript printed to stdout
npx voicebox-cli transcribe outputs/take.wav
# MP3 (or any ffmpeg-readable format) is converted to WAV automatically
npx voicebox-cli transcribe outputs/take.mp3
# Give a language hint for better accuracy
npx voicebox-cli transcribe outputs/take.wav --language en
# Pick a specific transcription model (whisper-base, whisper-small, whisper-medium, whisper-large, whisper-turbo)
npx voicebox-cli transcribe outputs/take.wav --model whisper-turbo
# Pass 'list' to any enum option to print its accepted values
npx voicebox-cli transcribe outputs/take.wav --model list
# Get the raw JSON (includes the audio duration)
npx voicebox-cli transcribe outputs/take.wav --json
# Save the transcript to a text file
npx voicebox-cli transcribe outputs/take.wav > outputs/take.txt
# Point at a server on another host/port
npx voicebox-cli transcribe outputs/take.wav --base-url http://192.168.1.50:17493
# Using short flags
npx voicebox-cli transcribe outputs/take.wav -l en -m whisper-turbo
```
### `health`
Report the API's status: whether the model is loaded, which backend and GPU are in use, and any compatibility warnings. Pass `-f`/`--filesystem` to instead check that the server's storage directories exist, are writable, and have free disk space. Add `--json` to print the raw response for scripting.
```
npx voicebox-cli health [options]
Options:
-f, --filesystem check filesystem health instead
--json print the raw JSON response
--base-url <url> API base url
```
```bash
$ npx voicebox-cli health
status: healthy
model: loaded (1.7B)
gpu: MPS (Apple Silicon)
backend: mlx (cpu)
```
### `shutdown`
Gracefully shut down the API server. Because this stops the process that serves every other command, it refuses to run unless you pass `--yes` to confirm. Point it at a specific server with `--base-url`.
```
npx voicebox-cli shutdown [options]
Options:
-y, --yes skip the confirmation prompt
--base-url <url> API base url
```
```bash
npx voicebox-cli shutdown --yes
```
### `watchdog`
Control the server's parent-process watchdog. By default the server shuts itself down when the process that launched it goes away; `watchdog disable` turns that off so the server keeps running on its own. This is useful when you started the server from a short-lived launcher but want it to persist.
```
npx voicebox-cli watchdog disable [options]
Options:
--base-url <url> API base url
```
### `install`
Install the bundled voicebox **skill** into an AI agent folder so an assistant such as Claude Code knows how to drive this CLI. It copies the `SKILL.md` tree into `<agent-folder>/skills/voicebox/`. The folder defaults to the current directory; pass a `.claude` folder (or any agent folder) to target a specific one.
```
npx voicebox-cli install [agent-folder]
```
```bash
# Install into ./.claude for the current project (the usual case)
npx voicebox-cli install .claude
# Install into the current directory (the default when no folder is given)
npx voicebox-cli install
```
## Use as an AI agent skill
voicebox-cli ships a [SKILL.md](dotclaude_folder/skills/voicebox/SKILL.md) that teaches an AI coding agent (e.g. Claude Code) how to use these commands for text-to-speech and transcription. Install it into a project's agent folder with:
```bash
npx voicebox-cli install .claude
```
The bundled skill lives under `dotclaude_folder/skills/`, mirroring the `.claude/skills/` layout. When developing this repo, mirror `dotclaude_folder/` into the repo's own `.claude/` as symlinks so the skill is live while its source stays tracked:
```bash
npm run symlink:dotclaude
```
## Output formats
The API serves WAV; the CLI transcodes locally.
| Extension | Codec | Notes |
| --- | --- | --- |
| `.wav` | PCM | Uncompressed, universal |
| `.mp3` | libmp3lame | Small, widely supported |
For a royalty-free, WhatsApp/Chromium-friendly format, transcode to Opus with the bundled binary:
```bash
node_modules/ffmpeg-static/ffmpeg -i outputs/speech.mp3 -c:a libopus -b:a 96k outputs/speech.ogg
```
## Project layout
```
src/
cli.ts # Commander entry point
commands/
speak_command.ts # `speak` command
generate_command.ts # `generate` command group
profiles_command.ts # `profiles` command group
channels_command.ts # `channels` command group
history_command.ts # `history` command group
models_command.ts # `models` command group
stories_command.ts # `stories` command group
transcribe_command.ts # `transcribe` command
health_command.ts # `health` command
shutdown_command.ts # `shutdown` command
watchdog_command.ts # `watchdog` command
install_command.ts # `install` command (copies the bundled skill)
misc/
voicebox_client.ts # VoiceboxClient — /speak, /profiles, status stream, audio download
audio_convert.ts # AudioConvert — WAV → MP3 via ffmpeg-static
dotclaude_folder/
skills/
voicebox/
SKILL.md # bundled agent skill, installed by `install`
scripts/
symlink_dotclaude.sh # mirror dotclaude_folder/ into .claude/ for local dev
examples/
generate_speech.ts # library usage without the CLI
outputs/ # generated audio (git-ignored)
```
## Programmatic use
```ts
import { VoiceboxClient } from './src/misc/voicebox_client.js';
const client = new VoiceboxClient();
const generation = await client.speak({ text: 'Hello', profile: 'Test' });
const final = await client.waitForCompletion(generation.id);
const wav = await client.downloadAudio(final.id);
```
## Scripts
```bash
npm run cli # run the CLI
npm run symlink:dotclaude # mirror dotclaude_folder/ into .claude/ (local dev)
npm run typecheck # tsc against tsconfig.json
npm run build # emit dist/ via tsconfig.build.json
```
## License
[MIT](LICENSE)