travisvn/chatterbox-tts-api
Local, OpenAI-compatible text-to-speech (TTS) API using Chatterbox, enabling users to generate voice cloned speech anywhere the OpenAI API is used (e.g. Open WebUI, AnythingLLM, etc.)
About travisvn/chatterbox-tts-api
travisvn/chatterbox-tts-api is an open-source project on GitHub, mainly written in Python. Local, OpenAI-compatible text-to-speech (TTS) API using Chatterbox, enabling users to generate voice cloned speech anywhere the OpenAI API is used (e.g. It currently holds 681 stars and 154 forks with 0 open issues, and was last pushed on an unknown date (repository created unknown).
Project Overview
AI Homed tracks it on the Local & On-Device AI board.
GitHub Repository Details
README
Chatterbox TTS API
FastAPI-powered REST API for Chatterbox TTS, providing OpenAI-compatible text-to-speech endpoints with voice cloning capabilities and additional features on top of the chatterbox-tts base package.
Features
🚀 OpenAI-Compatible API - Drop-in replacement for OpenAI's TTS API
⚡ FastAPI Performance - High-performance async API with automatic documentation
🌍 Multilingual Support - Generate speech in 22 languages with language-aware voice cloning
🎨 React Frontend - Includes an optional, ready-to-use web interface
🎭 Voice Cloning - Use your own voice samples for personalized speech
🎤 Voice Library Management - Upload, manage, and use custom voices by name
📝 Smart Text Processing - Automatic chunking for long texts
📊 Real-time Status - Monitor TTS progress, statistics, and request history
🐳 Docker Ready - Full containerization with persistent voice storage
⚙️ Configurable - Extensive environment variable configuration
🎛️ Parameter Control - Real-time adjustment of speech characteristics
📚 Auto Documentation - Interactive API docs at /docs and /redoc
🔧 Type Safety - Full Pydantic validation for requests and responses
🧠 Memory Management - Advanced memory monitoring and automatic cleanup
[!NOTE]
_Support for Chatterbox Turbo coming soon_
[!IMPORTANT]
resemble-ai/chatterbox is currently broken for non-CUDA setups (see chatterbox issues)
Revert to non-multilingual by using the stable branch of this repo
> View more instructions
⚡️ Quick Start
git clone https://github.com/travisvn/chatterbox-tts-api
cd chatterbox-tts-api
uv sync
uv run main.py
[!TIP]
uv installed with curl -LsSf https://astral.sh/uv/install.sh | sh
Local Installation with Python 🐍
Option A: Using uv (Recommended - Faster & Better Dependencies)
# Clone the repository
git clone https://github.com/travisvn/chatterbox-tts-api
cd chatterbox-tts-api
Install uv if you haven't already
curl -LsSf https://astral.sh/uv/install.sh | sh
Install dependencies with uv (automatically creates venv)
uv sync
Copy and customize environment variables
cp .env.example .env
Start the API with FastAPI
uv run uvicorn app.main:app --host 0.0.0.0 --port 4123
Or use the main script
uv run main.py
💡 Why uv? Users report better compatibility with chatterbox-tts, 25-40% faster installs, and superior dependency resolution. See migration guide →
Option B: Using pip (Traditional)
# Clone the repository
git clone https://github.com/travisvn/chatterbox-tts-api
cd chatterbox-tts-api
Setup environment — using Python 3.11
python -m venv .venv
source .venv/bin/activate
Install dependencies
pip install -r requirements.txt
Copy and customize environment variables
cp .env.example .env
Add your voice sample (or use the provided one)
cp your-voice.mp3 voice-sample.mp3
Start the API with FastAPI
uvicorn app.main:app --host 0.0.0.0 --port 4123
Or use the main script
python main.py
Ran into issues? Check the troubleshooting section
🐳 Docker (Recommended)
# Clone and start with Docker Compose
git clone https://github.com/travisvn/chatterbox-tts-api
cd chatterbox-tts-api
Use Docker-optimized environment variables
cp .env.example.docker .env # Docker-specific paths, ready to use
Or: cp .env.example .env # Local development paths, needs customization
Choose your deployment method:
API Only (default)
docker compose -f docker/docker-compose.yml up -d # Standard (pip-based)
docker compose -f docker/docker-compose.uv.yml up -d # uv-optimized (faster builds)
docker compose -f docker/docker-compose.gpu.yml up -d # Standard + GPU
docker compose -f docker/docker-compose.uv.gpu.yml up -d # uv + GPU (recommended for GPU users)
docker compose -f docker/docker-compose.cpu.yml up -d # CPU-only
docker compose -f docker/docker-compose.blackwell.yml up -d # Blackwell (50XX) NVIDIA GPUs
API + Frontend (add --profile frontend to any of the above)
docker compose -f docker/docker-compose.yml --profile frontend up -d # Standard + Frontend
docker compose -f docker/docker-compose.gpu.yml --profile frontend up -d # GPU + Frontend
docker compose -f docker/docker-compose.uv.gpu.yml --profile frontend up -d # uv + GPU + Frontend
docker compose -f docker/docker-compose.blackwell.yml --profile frontend up -d # (Blackwell) uv + GPU + Frontend
Watch the logs as it initializes (the first use of TTS takes the longest)
docker logs chatterbox-tts-api -f
Test the API
curl -X POST http://localhost:4123/v1/audio/speech \
-H "Content-Type: application/json" \
-d '{"input": "Hello from Chatterbox TTS!"}' \
--output test.wav
🚀 Running with the Web UI (Full Stack)
This project includes an optional React-based web UI. Use Docker Compose profiles to easily opt in or out of the frontend:
With Docker Compose Profiles
# API only (default behavior)
docker compose -f docker/docker-compose.yml up -d
API + Frontend + Web UI (with --profile frontend)
docker compose -f docker/docker-compose.yml --profile frontend up -d
Or use the convenient helper script for fullstack:
python start.py fullstack
Same pattern works with all deployment variants:
docker compose -f docker/docker-compose.gpu.yml --profile frontend up -d # GPU + Frontend
docker compose -f docker/docker-compose.uv.yml --profile frontend up -d # uv + Frontend
docker compose -f docker/docker-compose.cpu.yml --profile frontend up -d # CPU + Frontend
Local Development
For local development, you can run the API and frontend separately:
# Start the API first (follow earlier instructions)
Then run the frontend:
cd frontend && npm install && npm run dev
Click the link provided from Vite to access the web UI.
Build for Production
Build the frontend for production deployment:
cd frontend && npm install && npm run build
You can then access it directly from your local file system at /dist/index.html.
Port Configuration
- API Only: Accessible at
http://localhost:4123(direct API access) - With Frontend: Web UI at
http://localhost:4321, API requests routed via proxy
--profile frontend, the web interface will be available at http://localhost:4321 while the API runs behind the proxy.
Screenshots of Frontend (Web UI)
🖼️ View screenshot of full frontend web UI — light mode / dark mode
API Usage
Basic Text-to-Speech (Default Voice)
This endpoint works for both the API-only and full-stack setups.
curl -X POST http://localhost:4123/v1/audio/speech \
-H "Content-Type: application/json" \
-d '{"input": "Your text here"}' \
--output speech.wav
Using Custom Parameters (JSON)
curl -X POST http://localhost:4123/v1/audio/speech \
-H "Content-Type: application/json" \
-d '{"input": "Dramatic speech!", "exaggeration": 1.2, "cfg_weight": 0.3, "temperature": 0.9}' \
--output dramatic.wav
Custom Voice Upload
Upload your own voice sample for personalized speech:
curl -X POST http://localhost:4123/v1/audio/speech/upload \
-F "input=Hello with my custom voice!" \
-F "exaggeration=0.8" \
-F "voice_file=@my_voice.mp3" \
--output custom_voice_speech.wav
With Custom Parameters and Voice Upload
curl -X POST http://localhost:4123/v1/audio/speech/upload \
-F "input=Dramatic speech!" \
-F "exaggeration=1.2" \
-F "cfg_weight=0.3" \
-F "temperature=0.9" \
-F "voice_file=@dramatic_voice.wav" \
--output dramatic.wav
Voice Library Management
Store and manage custom voices by name for reuse across requests:
# Upload a voice to the library
curl -X POST http://localhost:4123/voices \
-F "voice_file=@my_voice.wav" \
-F "voice_name=my-custom-voice"
Upload a voice with language (multilingual support)
curl -X POST http://localhost:4123/voices \
-F "voice_file=@french_voice.wav" \
-F "voice_name=french-speaker" \
-F "language=fr"
Use the voice by name in speech generation
curl -X POST http://localhost:4123/v1/audio/speech \
-H "Content-Type: application/json" \
-d '{"input": "Hello with my custom voice!", "voice": "my-custom-voice"}' \
--output custom_voice_output.wav
Generate French speech (language auto-detected from voice)
curl -X POST http://localhost:4123/v1/audio/speech \
-H "Content-Type: application/json" \
-d '{"input": "Bonjour, comment allez-vous?", "voice": "french-speaker"}' \
--output french_speech.wav
List all available voices (includes language metadata)
curl http://localhost:4123/voices
Get supported languages
curl http://localhost:4123/languages
🔧 Complete Voice Library Documentation →
🌍 Multilingual Support
Generate speech in 22 languages with language-aware voice cloning and automatic language detection.
Supported Languages
Arabic (ar) • Danish (da) • German (de) • Greek (el) • English (en) • Spanish (es) • Finnish (fi) • French (fr) • Hebrew (he) • Hindi (hi) • Italian (it) • Japanese (ja) • Korean (ko) • Malay (ms) • Dutch (nl) • Norwegian (no) • Polish (pl) • Portuguese (pt) • Russian (ru) • Swedish (sv) • Swahili (sw) • Turkish (tr)
Quick Start
# Get supported languages
curl http://localhost:4123/languages
Upload voice with language
curl -X POST http://localhost:4123/voices \
-F "voice_name=spanish_speaker" \
-F "language=es" \
-F "voice_file=@spanish_voice.wav"
Generate multilingual speech
curl -X POST http://localhost:4123/v1/audio/speech \
-H "Content-Type: application/json" \
-d '{"input": "¡Hola! ¿Cómo estás hoy?", "voice": "spanish_speaker"}' \
--output spanish_speech.wav
Key Features
- 🎯 Language Auto-Detection - Voices store language metadata, automatically used in generation
- 🌐 No API Changes - Maintains OpenAI compatibility, language determined from voice metadata
- 🔄 Configurable - Enable/disable with
USE_MULTILINGUAL_MODELenvironment variable - 📚 Voice Library Integration - Language badges and filtering in web UI
- 🧠 Smart Fallback - Defaults to English for backward compatibility
🎵 Real-time Audio Streaming
The API supports multiple streaming formats for lower latency and better user experience:
- Raw Audio Streaming: Traditional audio chunks (WAV format)
- Server-Side Events (SSE): OpenAI-compatible format with base64-encoded audio chunks
Quick Start
# Basic audio streaming
curl -X POST http://localhost:4123/v1/audio/speech/stream \
-H "Content-Type: application/json" \
-d '{"input": "This streams in real-time!"}' \
--output streaming.wav
SSE streaming (OpenAI compatible)
curl -X POST http://localhost:4123/v1/audio/speech \
-H "Content-Type: application/json" \
-H "Accept: text/event-stream" \
-d '{"input": "This streams as Server-Side Events!", "stream_format": "sse"}' \
--no-buffer
Real-time playback
curl -X POST http://localhost:4123/v1/audio/speech/stream \
-H "Content-Type: application/json" \
-d '{"input": "Play as it generates!"}' \
| ffplay -f wav -i pipe:0 -autoexit -nodisp
🚀 Complete Streaming Documentation →
For comprehensive streaming features including:
- Advanced chunking strategies (sentence, paragraph, word, fixed)
- Quality presets (fast, balanced, high)
- Configurable parameters and performance tuning
- Real-time progress monitoring
- Python, JavaScript, and cURL examples
- Integration patterns for different use cases
- ⚡ Lower latency - Start hearing audio in 1-2 seconds
- 🎯 Better UX - No waiting for complete generation
- 💾 Memory efficient - Process chunks individually
- 🎛️ Configurable - Choose speed vs quality trade-offs
🐍 Python Examples
Default Voice (JSON)
import requests
response = requests.post(
"http://localhost:4123/v1/audio/speech",
json={
"input": "Hello world!",
"exaggeration": 0.8
}
)
with open("output.wav", "wb") as f:
f.write(response.content)
Upload Voice with Language (Multilingual)
import requests
Upload a multilingual voice
with open("german_voice.wav", "rb") as voice_file:
response = requests.post(
"http://localhost:4123/voices",
data={
"voice_name": "german_speaker",
"language": "de"
},
files={
"voice_file": ("german_voice.wav", voice_file, "audio/wav")
}
)
print(f"Upload status: {response.status_code}")
Generate German speech
response = requests.post(
"http://localhost:4123/v1/audio/speech",
json={
"input": "Guten Tag! Wie geht es Ihnen?",
"voice": "german_speaker",
"exaggeration": 0.8
}
)
with open("german_output.wav", "wb") as f:
f.write(response.content)
Upload Endpoint (Default Voice)
import requests
response = requests.post(
"http://localhost:4123/v1/audio/speech/upload",
data={
"input": "Hello world!",
"exaggeration": 0.8
}
)
with open("output.wav", "wb") as f:
f.write(response.content)
Custom Voice Upload
import requests
with open("my_voice.mp3", "rb") as voice_file:
response = requests.post(
"http://localhost:4123/v1/audio/speech/upload",
data={
"input": "Hello with my custom voice!",
"exaggeration": 0.8,
"temperature": 1.0
},
files={
"voice_file": ("my_voice.mp3", voice_file, "audio/mpeg")
}
)
with open("custom_output.wav", "wb") as f:
f.write(response.content)
Basic Streaming Example
import requests
Stream audio generation in real-time
response = requests.post(
"http://localhost:4123/v1/audio/speech/stream",
json={
"input": "This will stream as it's generated!",
"exaggeration": 0.8
},
stream=True # Enable streaming mode
)
with open("streaming_output.wav", "wb") as f:
for chunk in response.iter_content(chunk_size=8192):
if chunk:
f.write(chunk)
print(f"Received chunk: {len(chunk)} bytes")
SSE Streaming Example (OpenAI Compatible)
import requests
import json
import base64
Stream audio using Server-Side Events format
response = requests.post(
"http://localhost:4123/v1/audio/speech",
json={
"input": "This streams as Server-Side Events!",
"stream_format": "sse",
"exaggeration": 0.8
},
stream=True,
headers={'Accept': 'text/event-stream'}
)
audio_chunks = []
for line in response.iter_lines(decode_unicode=True):
if line.startswith('data: '):
event_data = line[6:] # Remove 'data: ' prefix
try:
event = json.loads(event_data)
if event.get('type') == 'speech.audio.delta':
# Decode base64 audio chunk
audio_data = base64.b64decode(event['audio'])
audio_chunks.append(audio_data)
print(f"Received audio chunk: {len(audio_data)} bytes")
elif event.get('type') == 'speech.audio.done':
usage = event.get('usage', {})
print(f"Complete! Tokens: {usage.get('total_tokens', 0)}")
break
except:
continue
print(f"Received {len(audio_chunks)} audio chunks")
📚 Complete Streaming Examples & Documentation →
Including real-time playback, progress monitoring, custom voice uploads, and advanced integration patterns.
Voice File Requirements
Supported Formats:
- MP3 (.mp3)
- WAV (.wav)
- FLAC (.flac)
- M4A (.m4a)
- OGG (.ogg)
- Maximum file size: 10MB
- Recommended duration: 10-30 seconds of clear speech
- Avoid background noise for best results
- Higher quality audio produces better voice cloning
🎛️ Configuration
The project provides two environment example files:
.env.example- For local development (uses./models,./voice-sample.mp3).env.example.docker- For Docker deployment (uses/cache,/app/voice-sample.mp3)
# For local development
cp .env.example .env
For Docker deployment
cp .env.example.docker .env
Key environment variables (see the example files for full list):
| Variable | Default | Description |
| ------------------------ | -------------------- | ------------------------------ |
| PORT | 4123 | API server port |
| USE_MULTILINGUAL_MODEL | true | Enable 23-language support |
| EXAGGERATION | 0.5 | Emotion intensity (0.25-2.0) |
| CFG_WEIGHT | 0.5 | Pace control (0.0-1.0) |
| TEMPERATURE | 0.8 | Sampling randomness (0.05-5.0) |
| VOICE_SAMPLE_PATH | ./voice-sample.mp3 | Voice sample for cloning |
| DEVICE | auto | Device (auto/cuda/mps/cpu) |
🎭 Voice Cloning
Replace the default voice sample:
# Replace the default voice sample
cp your-voice.mp3 voice-sample.mp3
Or set a custom path
echo "VOICE_SAMPLE_PATH=/path/to/your/voice.mp3" >> .env
For best results:
- Use 10-30 seconds of clear speech
- Avoid background noise
- Prefer WAV or high-quality MP3
🐳 Docker Deployment
Development
docker compose -f docker/docker-compose.yml up
Production
# Create production environment
cp .env.example.docker .env
nano .env # Set production values
Deploy
docker compose -f docker/docker-compose.yml up -d
With GPU Support
# Use GPU-enabled compose file
Ensure NVIDIA Container Toolkit is installed
docker compose -f docker/docker-compose.gpu.yml up -d
📚 API Reference
API Endpoints
| Endpoint | Method | Description |
| ----------------------------- | ------ | ------------------------------------------------------------------- |
| /audio/speech | POST | Generate speech from text (complete) |
| /audio/speech/upload | POST | Generate speech with voice upload |
| /audio/speech/stream | POST | Stream speech generation (docs) |
| /audio/speech/stream/upload | POST | Stream speech with voice upload (docs) |
| /voices | GET | List voices in library (with language metadata) |
| /voices | POST | Upload voice to library (with language support) |
| /languages | GET | Get supported languages (docs) |
| /health | GET | Health check and status |
| /config | GET | Current configuration |
| /v1/models | GET | Available models (OpenAI compat) |
| /status | GET | TTS processing status & progress |
| /status/progress | GET | Real-time progress (lightweight) |
| /status/statistics | GET | Processing statistics |
| /status/history | GET | Recent request history |
| /info | GET | Complete API information |
| /docs | GET | Interactive API documentation |
| /redoc | GET | Alternative API documentation |
Parameters Reference
Speech Generation Parameters
Exaggeration (0.25-2.0)
0.3-0.4: Professional, neutral0.5: Default balanced0.7-0.8: More expressive1.0+: Very dramatic
0.2-0.3: Faster speech0.5: Default pace0.7-0.8: Slower, deliberate
0.4-0.6: More consistent0.8: Default balance1.0+: More creative/random
audio: Raw audio streaming (default)sse: Server-Side Events with base64-encoded audio chunks (OpenAI compatible)
🧠 Memory Management
The API includes advanced memory management to prevent memory leaks and optimize performance:
Memory Management Features
- Automatic Cleanup: Periodic garbage collection and tensor cleanup
- CUDA Memory Management: Automatic GPU cache clearing
- Memory Monitoring: Real-time memory usage tracking
- Manual Controls: API endpoints for manual cleanup operations
Memory Configuration
| Variable | Default | Description | | --------------------------- | ------- | ----