1
rafska/awesome-local-llm
A curated list of awesome platforms, tools, practices and resources that helps run LLMs locally
About rafska/awesome-local-llm
rafska/awesome-local-llm is an open-source project on GitHub, mainly written in several languages. A curated list of awesome platforms, tools, practices and resources that helps run LLMs locally It currently holds 2,870 stars and 388 forks with 0 open issues, and was last pushed on an unknown date (repository created unknown).
Project Overview
AI Homed tracks it on the Local & On-Device AI board.
GitHub Repository Details
README
Awesome local LLM 
A curated list of awesome platforms, tools, practices and resources that helps run LLMs locally
Table of Contents
- Inference platforms
- Inference engines
- User Interfaces
- Large Language Models
- Explorers, Benchmarks, Leaderboards
- Model providers
- Specific models
- General purpose
- Coding
- Multimodal
- Image
- Audio
- Retrieval-Augmented Generation
- Safeguards
- Miscellaneous
- Tools
- Models
- Agent Frameworks
- Model Context Protocol
- Retrieval-Augmented Generation
- Coding Agents
- Computer Use
- Browser Automation
- Memory Management
- Testing, Evaluation and Observability
- Research
- Training and Fine-tuning
- Security and Sandboxing
- Miscellaneous
- Hardware
- Tutorials
- Models
- Prompt Engineering
- Context Engineering
- Inference
- Agents
- Retrieval-Augmented Generation
- Miscellaneous
- Communities
Inference platforms
- LM Studio - discover, download and run local LLMs
unsloth - unified web UI for training and running open models like Qwen, DeepSeek, and Gemma locally
LocalAI - the free, open-source alternative to OpenAI, Claude and others
jan - an open source alternative to ChatGPT that runs 100% offline on your computer
ChatBox - user-friendly desktop client app for AI models/LLMs
lemonade - a local LLM server with GPU and NPU Acceleration
Inference engines
ollama - get up and running with LLMs
llama.cpp - LLM inference in C/C++
vllm - a high-throughput and memory-efficient inference and serving engine for LLMs
exo - run your own AI cluster at home with everyday devices
BitNet - official inference framework for 1-bit LLMs
sglang - a fast serving framework for large language models and vision language models
TensorRT-LLM - provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs
Nano-vLLM - a lightweight vLLM implementation built from scratch
omlx - LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
koboldcpp - run GGUF models easily with a KoboldAI UI
mistral.rs - fast, flexible LLM inference
dynamo - a datacenter scale distributed inference serving framework
flashinfer - kernel library for LLM serving
mlx-lm - generate text and fine-tune large language models on Apple silicon with MLX
gpustack - simple, scalable AI model deployment on GPU clusters
LiteRT-LM - Google's production-ready, high-performance, open-source inference framework for deploying Large Language Models on edge devices
mlx-vlm - a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX
executorch - on-device AI across mobile, embedded and edge for PyTorch
mini-sglang - a lightweight yet high-performance inference framework for Large Language Models
distributed-llama - connect home devices into a powerful cluster to accelerate LLM inference
LiteRT - Google's on-device framework for high-performance ML & GenAI deployment on edge platforms, via efficient conversion, runtime, and optimization
ik_llama.cpp - llama.cpp fork with additional SOTA quants and improved performance
sonar - large-scale LLM inference engine based on vLLM
FastFlowLM - run LLMs on AMD Ryzen™ AI NPUs
tokenspeed - a speed-of-light LLM inference engine
krasis - a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware
vllm-gfx906 - vLLM for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60
llm-scaler - run LLMs on Intel Arc™ Pro B60 GPUs
User Interfaces
Open WebUI - User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
Lobe Chat - an open-source, modern design AI chat framework
Text generation web UI - LLM UI with advanced features, easy setup, and multiple backend support
SillyTavern - LLM Frontend for Power Users
Page Assist - Use your locally running AI models to assist you in your web browsing
Large Language Models
Explorers, Benchmarks, Leaderboards
- Arena - benchmark & compare the best AI models
- AI Models & API Providers Analysis - understand the AI landscape to choose the best model and provider for your use case
- SWE-rebench - a continuously evolving and decontaminated benchmark for software engineering LLMs
BullshitBench - measure whether AI models challenge nonsensical prompts instead of confidently answering them
- LLM Explorer - explore list of the open-source LLM models
- Dubesor LLM Benchmark table - small-scale manual performance comparison benchmark
- oobabooga benchmark - a list sorted by size (on disk) for each score
- CyberGym - evaluating AI agents' real-world cybersecurity capabilities at scale
vakra - a benchmark for evaluating multi-hop, multi-source tool-calling in AI agents
Model providers
- Qwen - powered by Alibaba Cloud
Mistral AI - a pioneering French artificial intelligence startup
- Tencent - a profile of a Chinese multinational technology conglomerate and holding company
- Unsloth AI - focusing on making AI more accessible to everyone (GGUFs etc.)
- bartowski - providing GGUF versions of popular LLMs
- Beijing Academy of Artificial Intelligence - a private non-profit organization engaged in AI research and development
- Open Thoughts - a team of researchers and engineers curating the best open reasoning datasets
Specific models
General purpose
- DeepSeek-V4 - a collection of the DeepSeek V4 LLMs
- Qwen3.8 - a collection of the latest generation Qwen LLMs
NVIDIA Nemotron v3 - a family of open models from NVIDIA with open weights, training data and recipes, delivering leading efficiency and accuracy for building specialized AI agents
Gemma 4 - a family of open models built by Google DeepMind, that are multimodal, handling text and image input (with audio supported on small models) and generating text output
Mistral Medium 3.5 - The first flaship models from Mistral AI handling instruction-following, reasoning, and coding in a single set of opened-weights
gpt-oss - a collection of open-weight models from OpenAI, designed for powerful reasoning, agentic tasks, and versatile developer use cases
gpt-oss-puzzle-88B - a deployment-optimized large language model developed by NVIDIA, derived from OpenAI's gpt-oss-120b
- Hunyuan - a collection of Tencent's open-source efficient LLMs designed for versatile deployment across diverse computational environments
- Phi-4 - a family of small language, multi-modal and reasoning models from Microsoft
OpenReasoning-Nemotron - a collection of models from NVIDIA, trained on 5M reasoning traces for math, code and science
- Kimi K2.5 - a collection of open-source, native multimodal agentic models from Moonshot AI that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration
- GLM-5.3 - a Z.ai's flagship model for long-horizon tasks
- Ling-3.0-flash - a native hybrid reasoning model from inclusionAI, operating with 124B total and 5.1B active parameters
- Granite 4.1 - efficient language models from IBM for multilingual generation, coding, RAG, and AI assistant workflows
- Ornith-1.5 - a collection of open-source models for agentic tasks and coding
- EXAONE-4.5 - LG's First Open-Weight Vision-Language Model for Industrial Intelligence
- Step-3.5-Flash - most capable open-source foundation model, engineered to deliver frontier reasoning and agentic capabilities with exceptional efficiency
- MiniCPM5 - a collection of SOTA on-device LLMs, small yet powerful
Coding
- Qwen3-Coder-Next - a collection of Qwen's open-weight language models designed specifically for coding agents and local development
Devstral 2 - a couple of agentic LLMs for software engineering tasks, excelling at using tools to explore codebases, edit multiple files, and power SWE Agents
Mellum 2 - an assistant model trained by JetBrain
- MiniMax-M3 - a native multimodal model with 1M context
- MiniMax-M2 - a collection of SOTA models for real-world dev & agents
- Laguna-S-2.1 - a 118B total parameter Mixture-of-Experts model with 8B activated parameters per token, designed for agentic coding and long-horizon work
- SWE-FastContext - a family of code-search models from Microsoft powering the Explore subagent for coding agents
- OmniCoder-9B - a 9-billion parameter coding agent model built by Tesslate, fine-tuned on top of Qwen3.5-9B's hybrid architecture
- NousCoder-14B - a competitive programming model post-trained on Qwen3-14B via reinforcement learning
- MusaCoder-27B - a code model developed by Moore Threads for PyTorch-to-CUDA/MUSA native kernel generation
Multimodal
- Qwen3-Omni - a collection of the natively end-to-end multilingual omni-modal foundation models from Qwen
- GLM-4.6V - a collection of open source multimodal models with native tool use from Zhipu AI
Image
- Qwen-Image - a collection of models for image generation, edit and decomposition from Qwen
- Qwen3-VL - a collection of the most powerful vision-language models in the Qwen series to date
- GLM-Image - an image generation model
- Granite Vision - multimodal models from IBM built for visual document analysis and image understanding
- HunyuanImage - a collection of image generation models from Tencent
- HunyuanVideo - a collection of video generation models from Tencent
- Vidi - a collection of models for multimodal video understanding and creation
- FastVLM - a collection of VLMs with efficient vision encoding from Apple
- MiniCPM-o & MiniCPM-V - multimodal models with leading performance
- LFM2-VL - a colection of vision-language models, designed for on-device deployment
- ClipTagger-12b - a vision-language model (VLM) designed for video understanding at massive scale
Audio
whisper-large-v3 - a state-of-the-art model for automatic speech recognition (ASR) and speech translation from OpenAI
Nemotron Speech - a collection of open, state-of-the-art, production‑ready enterprise speech models from NVIDIA for ASR, TTS, Speaker Diarization and S2SOpenAI
NVIDIA NemotronLabs VoiceChat 11B - a 11B end-to-end, real-time speech full duplex (FD) model from NVIDIA for conversational AI that jointly performs streaming speech understanding and speech generation
- Qwen3-ASR - a collection of models that support language identification and ASR for 52 languages and dialects
- Qwen3-TTS - a collection of TTS models that cover 10 major languages as well as multiple dialectal voice profiles to meet global application needs
- Granite Speech - a collection of compact and efficient speech-language models from IBM, specifically designed for multilingual automatic speech recognition (ASR) and bidirectional automatic speech translation (AST)
Voxtral-Small-24B-2507 - an enhancement of Mistral Small 3, incorporating state-of-the-art audio input capabilities while retaining best-in-class text performance
[Voxtral-Mini-4B-Realtime-2602](https://huggingface.co/mistralai/Voxtral-Mini-4B-Realti
GitHub Stars & Activity
2,870Stars
388Forks
0Open issues
-Language
GitHub Popularity
GitHub stars2,870
Forks388
Open issues0
Primary language-
License-
Stars gained today0
Created-
Last pushed-
Trending History
Trending statusnot on today's boards
Related AI Projects
2
3
4
5
6
7
8