huggingface / transformers
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
View huggingface/transformersThe audio board covers speech and sound: speech recognition, text-to-speech, voice cloning and conversion, music generation, source separation, transcription and general audio DSP. It is assembled from GitHub topic pages for text-to-speech, speech recognition and music generation and ranked by stars, which makes it a quick snapshot of where the open-source voice stack stands — long-standing toolkits next to this year's model releases. Every entry lists language, stars and forks and opens a detail page with description, license, last push date, README and related audio projects. For anyone wiring voice into a product or processing recordings at scale, the activity dates matter as much as the star counts.
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
View huggingface/transformers利用 AI 大模型和自动化工作流,根据主题或关键词一键生成高清短视频。Generate HD short videos from a topic or keyword with an automated AI workflow.
View harry0703/MoneyPrinterTurboLocal UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, DeepSeek-V4, MiniMax-H3, Gemma 4, FLUX and more.
View unslothai/unsloth1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
View RVC-Boss/GPT-SoVITSWorld's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files.
View calesthio/OpenMontage🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
View coqui-ai/TTSVoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
View OpenBMB/VoxCPMInstant voice cloning by MIT and MyShell. Audio foundation model.
View myshell-ai/OpenVoice🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time
View babysor/MockingBirdVoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.
View debpalash/VoiceStudioDeepSpeech is an open source embedded (offline, on-device) speech-to-text engine which can run in real time on devices ranging from a Raspberry Pi 4 to high power GPU servers.
View mozilla/DeepSpeechWhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
View m-bain/whisperXAn Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
View index-tts/index-ttsMulti-lingual large voice generation model, providing inference, training and deployment full-stack ability.
View QwenAudio/CosyVoiceOpen-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
View modelscope/FunASRA TTS model capable of generating ultra-realistic dialogue in one pass.
View nari-labs/diaTranslate the video from one language to another and embed dubbing & subtitles.
View jianchang512/pyvideotranskaldi-asr/kaldi is the official location of the Kaldi project.
View kaldi-asr/kaldiOffline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
View alphacep/vosk-apiState-of-the-Art Deep Learning scripts organized by models - easy to train and deploy with reproducible accuracy and performance on enterprise-grade infrastructure.
View NVIDIA/DeepLearningExamplesSpeech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection.
View k2-fsa/sherpa-onnxLightning-Fast, On-Device, Multilingual TTS — running natively via ONNX.
View supertone-oss-archive/supertonicDrench yourself in Deep Learning, Reinforcement Learning, Machine Learning, Computer Vision, and NLP by learning from these exciting lectures!!
View kmario23/deep-learning-drizzleGradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download
View abus-aikorea/voice-proEasy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System
View PaddlePaddle/PaddleSpeechUse Microsoft Edge's online text-to-speech service from Python WITHOUT needing Microsoft Edge or Windows or an API key
View rany2/edge-ttsReal-time, local speech-to-text with streaming ASR, speaker diarization, translation, and OpenAI/Deepgram-compatible APIs.
View QuentinFuxa/WhisperLiveKitOpenVINO™ is an open source toolkit for optimizing and deploying AI inference
View openvinotoolkit/openvinoAmphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get started in the field of audio
View open-mmlab/Amphion🤖 💬 Deep learning for Text to Speech (Discussion forum: https://discourse.mozilla.org/c/tts)
View mozilla/TTSOpen-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.
View QwenAudio/SenseVoiceYuE2: frontier music generation with symbolic planning, zero-shot covers, and agentic music editing.
View multimodal-art-projection/YuESpeech recognition module for Python, supporting several engines and APIs, online and offline.
View Uberi/speech_recognitionLab Materials for MIT 6.S191: Introduction to Deep Learning
View MITDeepLearning/introtodeeplearningEmotiVoice 😊: a Multi-Voice and Prompt-Controlled TTS Engine
View netease-youdao/EmotiVoiceA Deep-Learning-Based Chinese Speech Recognition System 基于深度学习的中文语音识别系统
View nl8590687/ASRT_SpeechRecognitionAn open source implementation of Microsoft's VALL-E X zero-shot TTS model. Demo is available in https://plachtaa.github.io/vallex/
View Plachtaa/VALL-E-XA text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
View Blaizzy/mlx-audioVITS: Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech
View jaywalnut310/vitsHigh-quality multi-lingual text-to-speech library by MyShell.ai. Support English, Spanish, French, Chinese, Japanese and Korean.
View myshell-ai/MeloTTSeSpeak NG is an open source speech synthesizer that supports more than hundred languages and accents.
View espeak-ng/espeak-ngTurn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation.
View Osmantic/ODSFacebook AI Research's Automatic Speech Recognition Toolkit
View flashlight/wav2letterStyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models
View yl4579/StyleTTS2This repository contains a hand-curated resources for Prompt Engineering with a focus on Generative Pre-trained Transformer (GPT), ChatGPT, PaLM etc
View promptslab/Awesome-Prompt-EngineeringFunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio UI.
View modelscope/FunClipSilero Models: pre-trained text-to-speech models made embarrassingly simple
View snakers4/silero-modelsGenerate audiobooks from EPUBs, PDFs and text with synchronized captions.
View denizsafak/abogenOpen source voice AI platform. Self-hosted alternative to Vapi and Retell. On Prem, BYOK across Speech to Speech or LLM/STT/TTS, with a visual workflow builder, MCP native and telephony support.
View dograh-hq/dograhAutomatic Speech Recognition with Speaker Diarization based on OpenAI Whisper
View MahmoudAshraf97/whisper-diarizationOpen-source AI video localization and dubbing for YouTube/Bilibili: speech recognition, subtitle translation, voice cloning, audio mixing and rendering. 开源 AI 视频翻译配音工具。
View liuzhao1225/YouDub-webuiDockerized OpenAI-compatible wrapper for Kokoro-82M text-to-speech w/multiplatform CPU, AMD, NVIDIA GPU PyTorch; multi-speaker, clone-tuning, caption timestamps, SSML, optional readalong web UI
View remsky/Kokoro-FastAPIProduction First and Production Ready End-to-End Speech Recognition Toolkit
View wenet-e2e/wenetOn-device wake word detection powered by deep learning
View Picovoice/porcupineMachine Learning and Agentic AI Resources, Practice and Research
View yanshengjia/ml-road🎵 The Ultimate Open Source Suno Alternative - Professional UI for ACE-Step 1.5 AI Music Generation. Free, local, unlimited. Stop paying for Suno!
View fspecii/ace-step-uiDiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism (SVS & TTS); AAAI 2022; Official code
View MoonInTheRiver/DiffSingerVoice Recognition to Text Tool / 一个离线运行的本地音视频转字幕工具,输出json、srt字幕、纯文字格式
View jianchang512/sttJAX implementation of OpenAI's Whisper model for up to 70x speed-up on TPU.
View sanchit-gandhi/whisper-jaxThe Ruby-native AI framework. Chats, agents, tools, images, audio, and video through one consistent API, in plain Ruby or Rails.
View crmne/ruby_llmA 100M-parameter multilingual TTS model for real-time CPU inference, voice cloning, and 48 kHz stereo generation
View OpenMOSS/MOSS-TTS-NanoFoundational model for human-like, expressive TTS
View metavoiceio/metavoice-srcDistilled variant of Whisper for speech recognition. 6x faster, 50% smaller, within 1% word error rate.
View huggingface/distil-whisperAn open-source model family for long-form speech, dialogue synthesis, voice design, sound effects, and real-time streaming TTS
View OpenMOSS/MOSS-TTS😝 TensorFlowTTS: Real-Time State-of-the-art Speech Synthesis for Tensorflow 2 (supported including English, French, Korean, Chinese, German and Easy to adapt for other languages)
View TensorSpeech/TensorFlowTTSA simple, high-quality voice conversion tool focused on ease of use and performance.
View IAHispano/ApplioOpenAI Whisper ASR Webservice API
View ahmetoner/whisper-asr-webserviceA single Gradio + React WebUI with extensions for ACE-Step, OmniVoice, Kimi Audio, Piper TTS, GPT-SoVITS, CosyVoice, XTTSv2, DIA, Kokoro, OpenVoice, ParlerTTS, Stable Audio, MMS, StyleTTS2, MAGNet
View rsxdalv/TTS-WebUIWhisper & Faster-Whisper standalone executables for those who don't want to bother with Python.
View Purfview/whisper-standalone-winAutomatic Speech Recognition (ASR), Speaker Verification, Speech Synthesis, Text-to-Speech (TTS), Language Modelling, Singing Voice Synthesis (SVS), Voice Conversion (VC)
View zzw922cn/awesome-speech-recognition-speech-synthesis-papers这是一个全自动(音频)视频翻译项目。利用Whisper识别声音,AI大模型翻译字幕,最后合并字幕视频,生成翻译后的视频。
View chenyme/Chenyme-AAVT[AutoArk] GPA (General Purpose Audio) can do ASR, TTS and voice conversion with one tiny model!
View AutoArk/GPAOpen source, local, and self-hosted Amazon Echo/Google Home competitive Voice Assistant alternative
View HeyWillow/willowThe official Python SDK for the ElevenLabs API.
View elevenlabs/elevenlabs-pythonTranscribe any audio to text, translate and edit subtitles 100% locally with a web UI. Powered by whisper models!
View pluja/whishperaeneas is a Python/C library and a set of tools to automagically synchronize audio and text (aka forced alignment)
View readbeyond/aeneasMultilingual Automatic Speech Recognition with word-level timestamps and confidence
View linto-ai/whisper-timestampedEnd-to-end Automatic Speech Recognition for Madarian and English in Tensorflow
View zzw922cn/Automatic_Speech_RecognitionAn all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, and more, with highly optimized performance.
View 0xShug0/audio.cppA realtime voice runtime that keeps Agents talking, working, and present. Real-time Voice Runtime for AI Agents
View QwenAudio/qwen-audio-agentPython library and CLI tool to interface with Google Translate's text-to-speech API
View pndurette/gTTS🐸STT - The deep learning toolkit for Speech-to-Text. Training and deploying STT models has never been so easy.
View coqui-ai/STT