Blaizzy/mlx-audio

★ 7,894⑂ 0

A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.

About Blaizzy/mlx-audio

Blaizzy/mlx-audio is an open-source project on GitHub, mainly written in Python. A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework It currently holds 7,894 stars and 0 forks with 0 open issues, and was last pushed on an unknown date (repository created unknown).

Project Overview

AI Homed tracks it on the AI Audio Projects board and on the AI AI Audio Projects list.

GitHub Repository Details

Repository Blaizzy/mlx-audio · default branch - · size 0 KB · watchers 0 · source: GitHub REST API and repository README

README

MLX-Audio

https://github.com/Blaizzy/mlx-audio/blob/HEAD/Blaizzy%2Fmlx-audio | Trendshift

PyPI version Python License: MIT GitHub stars

The best audio processing library built on Apple's MLX framework, providing fast and efficient text-to-speech (TTS), speech-to-text (STT), speech-to-speech (STS), music generation, and more on Apple Silicon.

Table of Contents

Features

Installation

Using pip

pip install mlx-audio

Using uv to install only the command line tools

Latest release from pypi:
uv tool install --force mlx-audio --prerelease=allow

Latest code from github:

uv tool install --force git+https://github.com/Blaizzy/mlx-audio.git --prerelease=allow

For development or web interface:

git clone https://github.com/Blaizzy/mlx-audio.git
cd mlx-audio
pip install -e ".[dev, server]"

Quick Start

Command Line

# Basic TTS generation
mlx_audio.tts.generate --model mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-8bit --text 'Hello, world!' --voice Vivian

With a different voice and language hint

mlx_audio.tts.generate --model mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-8bit --text 'Welcome to MLX-Audio!' --voice Ryan --lang_code English

Play audio immediately

mlx_audio.tts.generate --model mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-8bit --text 'Hello!' --voice Vivian --play

Save to a specific directory

mlx_audio.tts.generate --model mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-8bit --text 'Hello!' --voice Vivian --output_path ./my_audio

Stream audio during generation

mlx_audio.tts.generate --model mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-8bit --text 'Hello!' --voice Vivian --stream

Stream audio during generation and save it to disk

mlx_audio.tts.generate --model mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-8bit --text 'Hello!' --voice Vivian --stream --save

Join multiple generated segments into one file

mlx_audio.tts.generate --model mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-8bit --text $'Hello!\nHow are you?' --voice Vivian --join_audio

By default, when generation yields multiple segments, mlx-audio saves numbered files such as audio_000.wav and audio_001.wav. Use --join_audio to save one combined file instead. When using --stream, add --save to write the streamed audio to disk.

Python API

from mlx_audio.tts.utils import load_model

Load model

model = load_model("mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-8bit")

Generate speech

for result in model.generate( "Hello from MLX-Audio!", voice="Vivian", lang_code="English", ): print(f"Generated {result.audio.shape[0]} samples") # result.audio contains the waveform as mx.array

Supported Models

Text-to-Speech (TTS)

| Model | Description | Languages | Repo | |-------|-------------|-----------|------| | Kokoro | Fast, high-quality multilingual TTS | EN, JA, ZH, FR, ES, IT, PT, HI | bf16, 8bit, 6bit, 4bit | | KittenTTS | Compact KittenTTS 0.8 models for edge-friendly TTS | EN | nano, micro, mini, collection | | Qwen3-TTS | Alibaba's multilingual TTS with voice design | ZH, EN, JA, KO, + more | mlx-community/Qwen3-TTS-12Hz-1.7B-VoiceDesign-bf16 | | Higgs Audio v3 | 4B conversational TTS with voice cloning and inline control tokens | 100 languages | bosonai/higgs-audio-v3-tts-4b | | OmniVoice | Zero-shot multilingual TTS with voice cloning, batch generation, and nonverbal tags | 646+ languages | mlx-community/OmniVoice-bf16 | | CSM / MisoTTS | Sesame-style conversational speech models with voice cloning | EN | mlx-community/csm-1b, MisoTTS bf16, MisoTTS 8bit | | Dia | Dialogue-focused TTS | EN | mlx-community/Dia-1.6B-fp16 | | OuteTTS | Efficient TTS model | EN | mlx-community/OuteTTS-1.0-0.6B-fp16 | | Spark | SparkTTS model | EN, ZH | mlx-community/Spark-TTS-0.5B-bf16 | | Chatterbox | Expressive multilingual TTS (v2/v3) | 23 languages | v3, v2 | | Soprano | High-quality TTS | EN | mlx-community/Soprano-1.1-80M-bf16 | | Ming Omni TTS (BailingMM) | Multimodal generation with voice cloning, style control, and speech/music/event generation | EN, ZH | mlx-community/Ming-omni-tts-16.8B-A3B-bf16 | | Ming Omni TTS (Dense) | Lightweight dense Ming Omni variant for voice cloning and style control | EN, ZH | mlx-community/Ming-omni-tts-0.5B-bf16 | | KugelAudio | SOTA 7B AR+Diffusion TTS for European languages | EN, DE, FR, ES, IT, PT, NL, PL, RU, UK, + 14 more | kugelaudio/kugelaudio-0-open | | Voxtral TTS | Mistral's 4B multilingual TTS (20 voices, 9 languages) | EN, FR, ES, DE, IT, PT, NL, AR, HI | mlx-community/Voxtral-4B-TTS-2603-mlx-bf16 | | rumik-oss 1 | 3B expressive multilingual Indic TTS with 22-language support, description-conditioned delivery and inline vocalizations | 22 Indic languages + EN | rumik-ai/rumik-oss-1, 8bit, 4bit | | VoxCPM2 | 2B tokenizer-free TTS with 48kHz output, voice design, voice cloning, and continuation | 30 languages | bf16, 8bit, 4bit | | LongCat-AudioDiT | SOTA diffusion TTS in waveform latent space with voice cloning | ZH, EN | mlx-community/LongCat-AudioDiT-1B-bf16 | | MeloTTS | Lightweight VITS2-based TTS with streaming | EN (more coming) | mlx-community/MeloTTS-English-MLX | | MOSS-TTS | 8B delay-pattern and local-transformer multilingual TTS with voice cloning | 31 languages | OpenMOSS-Team/MOSS-TTS-v1.5, OpenMOSS-Team/MOSS-TTS, OpenMOSS-Team/MOSS-TTS-Local-Transformer-v1.5, OpenMOSS-Team/MOSS-TTS-Local-Transformer | | MOSS-TTS-Nano | Tiny multilingual voice-cloning TTS | 20 languages | mlx-community/MOSS-TTS-Nano-100M | | Higgs Audio v2 | 3B Llama-backed TTS with real-time voice cloning | EN, ZH, KO, DE, ES | bf16 (upstream), q8, q6 |

Speech-to-Text (STT)

| Model | Description | Languages | Repo | |-------|-------------|-----------|------| | Whisper | OpenAI's robust STT model | 99+ languages | mlx-community/whisper-large-v3-turbo-asr-fp16 | | Distil-Whisper | Distilled fast Whisper variants | EN | distil-whisper/distil-large-v3 | | Qwen3-ASR | Alibaba's multilingual ASR | ZH, EN, JA, KO, + more | mlx-community/Qwen3-ASR-1.7B-8bit | | Mega-ASR | Routed Qwen3-ASR with automatic clean/base vs degraded/LoRA switching | EN (fixtures), multilingual Qwen3-ASR backbone | README | | Qwen3-ForcedAligner | Word-level audio alignment | ZH, EN, JA, KO, + more | mlx-community/Qwen3-ForcedAligner-0.6B-8bit | | MOSS-Transcribe-Diarize | Timestamped transcription with speaker labels | Multiple major languages | https://huggingface.co/OpenMOSS-Team/MOSS-Transcribe-Diarize | | Parakeet | NVIDIA's accurate STT | EN (v2), 25 EU languages (v3) | mlx-community/parakeet-tdt-0.6b-v3 | | Nemotron 3.5 ASR (streaming) | NVIDIA's cache-aware streaming FastConformer-RNNT with language-ID prompting | 40 language-locales | mlx-community/nemotron-3.5-asr-streaming-0.6b · README | | Voxtral | Mistral's speech model | Multiple | mlx-community/Voxtral-Mini-3B-2507-bf16 | | Voxtral Realtime | Mistral's 4B streaming STT | Multiple | 4bit, fp16 | | VibeVoice-ASR | Microsoft's 3B/9B ASR with diarization, timestamps, hotwords, and native chunk streaming | 10 streaming / 50+ long-form | Streaming 1.5B · Streaming 7B · Long-form · README | | Canary | NVIDIA's multilingual ASR with translation | 25 EU + RU, UK | README | | Moonshine | Useful Sensors' lightweight ASR | EN | README | | MMS | Meta's massively multilingual ASR with adapters | 1000+ | README | | Granite Speech | IBM's ASR + speech translation | EN, FR, DE, ES, PT, JA | README | | Granite Speech 5.0 TurboCTC | IBM's fast encoder-only CTC ASR | EN | README | | Qwen2-Audio | Alibaba's multimodal audio understanding (ASR, captioning, emotion, translation) | Multiple | mlx-community/Qwen2-Audio-7B-Instruct-4bit | | MOSS-Music | OpenMOSS music understanding and lyrics ASR | EN, ZH | README |

Voice Activity Detection / Speaker Diarization (VAD)

| Model | Description | Languages | Repo | |-------|-------------|-----------|------| | Silero VAD | Lightweight speech/non-speech detection with streaming state | Language-agnostic | mlx-community/silero-vad | | Sortformer v1 | NVIDIA's end-to-end speaker diarization (up to 4 speakers) | Language-agnostic | mlx-community/diar_sortformer_4spk-v1-fp32 | | Sortformer v2.1 | NVIDIA's streaming speaker diarization with AOSC compression | Language-agnostic | mlx-community/diar_streaming_sortformer_4spk-v2.1-fp32 |

See the model READMEs for API details, streaming examples, and conversion steps.

Speech-to-Speech (STS)

| Model | Description | Use Case | Repo | |-------|-------------|----------|------| | SAM-Audio | Text-guided source separation | Extract specific sounds | mlx-community/sam-audio-large | | DialogueSidon | Two-speaker separation and restoration | Separate dialogue into speaker tracks | mlx-community/DialogueSidon (FP32), mlx-community/DialogueSidon-bf16 (BF16) | | Liquid2.5-Audio* | Speech-to-Speech, Text-to-Speech and Speech-to-Text | Speech interactions | mlx-community/LFM2.5-Audio-1.5B-8bit | | MiMo-Audio | English/Chinese TTS, ASR, audio understanding and dialogue; Base few-shot speech tasks | Speech interactions and audio completion | Instruct, Base, audio tokenizer, guide | | MossFormer2 SE | Speech enhancement | Noise removal | starkdmi/MossFormer2_SE_48K_MLX | | DeepFilterNet (1/2/3) | Speech enhancement | Noise suppression | mlx-community/DeepFilterNet-mlx | | NemotronLabs VoiceChat | Full-duplex speech-to-speech with streaming transcription and function calling | Real-time voice conversation | mlx-community/NemotronLabs-VoiceChat-11B-4bit |

Music Generation

| Model | Description | Languages | Repo | |-------|-------------|-----------|------| | MiniMax Music 3 | Hierarchical AR + flow-matching song generation with lyrics and 44.1 kHz stereo output | Multilingual lyrics | BF16, 8-bit, 6-bit, 4-bit, MXFP8, MXFP4, NVFP4, guide |

Model Examples

Qwen3-TTS

Alibaba's state-of-the-art multilingual TTS with voice cloning, emotion control, and voice design capabilities.

from mlx_audio.tts.utils import load_model

model = load_model("mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-bf16") results = list(model.generate_custom_voice( text="Hello, welcome to MLX-Audio!", speaker="Vivian", language="English", ))

audio = results[0].audio # mx.array

See the Qwen3-TTS README for voice cloning, CustomVoice, VoiceDesign, and all available models.

OmniVoice

OmniVoice is a zero-shot multilingual TTS model for 646+ languages with voice cloning, batch generation, pronunciation controls, and nonverbal tags such as [laughter] and [sigh]. It uses a bidirectional Qwen3 backbone with iterative masked generation and a HiggsAudioV2 acoustic tokenizer.

from mlx_audio.tts.utils import load_model

model = load_model("mlx-community/OmniVoice-bf16")

Basic multilingual TTS

for result in model.generate( text="Hello from OmniVoice running on Apple Silicon.", language="english", duration_s=5.0, num_steps=32, ): audio = result.audio

Zero-shot voice cloning

for result in model.generate( text="This sentence uses the reference speaker.", language="english", ref_audio="reference.wav", ref_text="Transcript of the reference audio.", duration_s=5.0, ): audio = result.audio

For stable voice cloning, provide ref_text that matches the reference clip. OmniVoice also supports generate_batch() for batched TTS and inline pronunciation controls.

Ming Omni TTS (BailingMM)

mlx_audio.tts.generate \
    --model mlx-community/Ming-omni-tts-16.8B-A3B-bf16 \
    --prompt "Please generate speech based on the following description.\n" \
    --text "This is a quick Ming Omni test." \
    --lang_code en \
    --output_path audio_io \
    --file_prefix ming_basic \
    --verbose

See the Ming Omni TTS README for CLI and Python cookbook examples, and the Ming Omni Dense README for the mlx-community/Ming-omni-tts-0.5B-bf16 workflow.

Kokoro TTS

Kokoro is a fast, multilingual TTS model with 54 voice presets.

from mlx_audio.tts.utils import load_model

model = load_model("mlx-community/Kokoro-82M-bf16")

Or use a quantized variant for lower memory usage:

model = load_model("mlx-community/Kokoro-82M-8bit")

model = load_model("mlx-community/Kokoro-82M-4bit")

Generate with different voices

for result in model.generate( text="Welcome to MLX-Audio!", voice="af_heart", # American female speed=1.0, lang_code="a" # American English ): audio = result.audio

Available Voices:

Kokoro requires pip install misaki for text processing. Japanese and Mandarin may additionally require pip install misaki[ja] or pip install misaki[zh].

Language Codes: | Code | Language | Note | |------|----------|------| | a | American English | Default; requires pip install misaki | | b | British English | Requires pip install misaki | | j | Japanese | Requires pip install misaki[ja] | | z | Mandarin Chinese | Requires pip install misaki[zh] | | e | Spanish | Requires pip install misaki | | f | French | Requires pip install misaki |

CSM (Voice Cloning)

Clone any voice using a reference audio sample:

mlx_audio.tts.generate \
    --model mlx-community/csm-1b \
    --text "Hello from Sesame." \
    --ref_audio ./reference_voice.wav \
    --play

Whisper STT

from mlx_audio.stt.generate import generate_transcription

result = generate_transcription( model="mlx-community/whisper-large-v3-turbo-asr-fp16", audio="audio.wav", ) print(result.text)

Qwen3-ASR & ForcedAligner

Alibaba's multilingual speech models for transcription and word-level alignment.

from mlx_audio.stt import load

Speech recognition

model = load("mlx-community/Qwen3-ASR-0.6B-8bit") result = model.generate("audio.wav", language="English") print(result.text)

Word-level forced alignment

aligner = load("mlx-community/Qwen3-ForcedAligner-0.6B-8bit") result = aligner.generate("audio.wav", text="I have a dream", language="English") for item in result: print(f"[{item.start_time:.2f}s - {item.end_time:.2f}s] {item.text}")

See the Qwen3-ASR README for CLI usage, all models, and more examples.

Phonon-1

Fermion Research's compact English Qwen3-ASR derivatives load directly from their public transport repositories:

from mlx_audio.stt import load

model = load("FermionResearch/Phonon-1") result = model.generate("audio.wav", language="English") print(result.text)

Available builds are Phonon-1-Micro (285 MB), Phonon-1 (415 MB), and Phonon-1-Big (581 MB). See the Phonon-1 README.

VibeVoice-ASR

Microsoft's 9B parameter speech-to-text model with speaker diarization and timestamps. Supports long-form audio (up to 60 minutes) and outputs structured JSON.

from mlx_audio.stt.utils import load

model = load("mlx-community/VibeVoice-ASR-bf16")

Basic transcription

result = model.generate(audio="meeting.wav", max_tokens=8192, temperature=0.0) print(result.text)

[{"Start":0,"End":5.2,"Speaker":0,"Content":"Hello everyone, let's begin."},

{"Start":5.5,"End":9.8,"Speaker":1,"Content":"Thanks for joining today."}]

Access parsed segments

for seg in result.segments: print(f"[{seg['start_time']:.1f}-{seg['end_time']:.1f}] Speaker {seg['speaker_id']}: {seg['text']}")

Streaming transcription:

# Stream tokens as they are generated
for text in model.stream_transcribe(audio="speech.wav", max_tokens=4096):
    print(text, end="", flush=True)

With context (hotwords/metadata):

result = model.generate(
    audio="technical_talk.wav",
    context="MLX, Apple Silicon, PyTorch, Transformer",
    max_tokens=8192,
    temperature=0.0,
)

CLI usage:

# Basic transcription
python -m mlx_audio.stt.generate \
    --model mlx-community/VibeVoice-ASR-bf16 \
    --audio meeting.wav \
    --output-path output \
    --format json \
    --max-tokens 8192 \
    --verbose

With context/hotwords

python -m mlx_audio.stt.generate \ --model mlx-community/VibeVoice-ASR-bf16 \ --audio technical_talk.wav \ --output-path output \ --format json \ --max-tokens 8192 \ --context "MLX, Apple Silicon, PyTorch, Transformer" \ --verbose

Parakeet (Multilingual STT)

NVIDIA's high-accuracy speech-to-text model. Parakeet v3 supports 25 European languages.

from mlx_audio.stt.utils import load

Load the multilingual v3 model

model = load("mlx-community/parakeet-tdt-0.6b-v3")

Transcribe audio

result = model.generate("audio.wav") print(f"Text: {result.text}")

Access word-level timestamps

for sentence in result.sentences: print(f"[{sentence.start:.2f}s - {sentence.end:.2f}s] {sentence.text}")

Streaming transcription:

```python for chunk in model.generate("long_audio.wav", stre

GitHub Stars & Activity

7,894Stars
0Forks
0Open issues
PythonLanguage

GitHub Popularity

GitHub stars7,894
Forks0
Open issues0
Primary languagePython
License-
Stars gained today0
Created-
Last pushed-

Trending History

Trending statusnot on today's boards

Related AI Projects

1

huggingface / transformers

Python★ 166,221⑂ 0
2

harry0703 / MoneyPrinterTurbo

Python★ 124,071⑂ 0
3

unslothai / unsloth

Python★ 76,216⑂ 0
4

RVC-Boss / GPT-SoVITS

Python★ 61,795⑂ 0
5

calesthio / OpenMontage

Python★ 59,377⑂ 0
6

coqui-ai / TTS

Python★ 46,016⑂ 0
7

2noise / ChatTTS

Python★ 39,845⑂ 0
8

OpenBMB / VoxCPM

Python★ 37,599⑂ 0

More AI Rankings