Trending open-source projects · updated daily

Local & On-Device AI Projects

Running AI locally means the model, the weights and the data all stay on hardware you control: no API key, no per-token bill, no prompt leaving the machine. This board collects the projects that make that practical — local model runners and servers, chat interfaces and desktop assistants, mobile and on-device inference, and the quantization and compression tooling that lets a model fit into laptop or phone memory. Entries come from GitHub topic pages for local LLMs, llama.cpp-style inference, on-device AI and Ollama, ranked by stars, so a mature server sits next to a desktop client published last month. Each card lists language, stars and forks and opens a page with description, license, activity dates, README and related local AI projects. The practical question is almost always memory: check the quantization notes on the detail page before assuming a model will fit.

Trending Local & On-Device AI Projects

1

ollama / ollama

Get up and running with Kimi, GLM, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.

Go★ 181,296⑂ 17,941
View ollama/ollama
2

open-webui / open-webui

User-friendly AI Interface (Supports Ollama, OpenAI API, ...)

Python★ 152,601⑂ 22,332
View open-webui/open-webui
3

ChatGPTNextWeb / NextChat

✨ Zero-config AI chat assistant. No API key needed — sign up and instantly chat with GPT-5, Claude 4, Gemini 2.5, DeepSeek & 100+ top models. Pay-as-you-go saves you more.

TypeScript★ 88,790⑂ 59,030
View ChatGPTNextWeb/NextChat
4

HKUDS / nanobot

Ultra-lightweight, open-source, self-hosted personal AI agent framework in Python with WebUI, tools, memory, MCP, multi-agent workflows, automation, and chat apps

Python★ 48,395⑂ 8,551
View HKUDS/nanobot
5

chatboxai / chatbox

Powerful AI Client

TypeScript★ 41,812⑂ 4,252
View chatboxai/chatbox
6

chatchat-space / Langchain-Chatchat

Langchain-Chatchat(原Langchain-ChatGLM)基于 Langchain 与 ChatGLM, Qwen 与 Llama 等语言模型的 RAG 与 Agent 应用 | Langchain-Chatchat (formerly langchain-ChatGLM), local knowledge based LLM (like ChatGLM

Python★ 38,648⑂ 6,265
View chatchat-space/Langchain-Chatchat
7

1Panel-dev / 1Panel

🔥 1Panel is a modern, open-source Linux server management panel and a lightweight AI management platform.

Go★ 36,983⑂ 3,358
View 1Panel-dev/1Panel
8

Zackriya-Solutions / meetily

Privacy first, AI meeting assistant with 4x faster Parakeet/Whisper live transcription, speaker diarization, and Ollama summarization built on Rust. 100% local processing. no cloud required.

Rust★ 30,962⑂ 3,361
View Zackriya-Solutions/meetily
9

Tencent / WeKnora

Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.

Go★ 27,811⑂ 3,743
View Tencent/WeKnora
10

mozilla-ai / llamafile

Distribute and run LLMs with a single file.

C++★ 26,001⑂ 1,598
View mozilla-ai/llamafile
11

1Panel-dev / MaxKB

🔥 MaxKB is an open-source platform for building enterprise-grade agents. 强大易用的开源企业级智能体平台。

Python★ 22,844⑂ 3,157
View 1Panel-dev/MaxKB
12

dyad-sh / dyad

Local, open-source AI app builder for power users ✨ v0 / Lovable / Replit / Bolt alternative 🌟 Star if you like it!

TypeScript★ 21,587⑂ 2,639
View dyad-sh/dyad
13

lss233 / kirara-ai

🤖 可 DIY 的 多模态 AI 聊天机器人 | 🚀 快速接入 微信、 QQ、Telegram、等聊天平台 | 🦈支持DeepSeek、Grok、Claude、Ollama、Gemini、OpenAI | 工作流系统、网页搜索、AI画图、人设调教、虚拟女仆、语音对话 |

Python★ 19,027⑂ 1,836
View lss233/kirara-ai
14

AsyncFuncAI / deepwiki-open

Open Source DeepWiki: AI-Powered Wiki Generator for GitHub/Gitlab/Bitbucket Repositories. Join the discord: https://discord.gg/gMwThUMeme

Python★ 18,018⑂ 2,003
View AsyncFuncAI/deepwiki-open
15

langbot-app / LangBot

Production-grade platform for building agentic IM bots - 生产级多平台智能机器人开发平台/ Agent、知识库编排、插件系统 / Bots for Discord / Slack / LINE / Telegram / WeChat(企业微信, 企微智能机器人, 公众号) / 飞书 / 钉钉 / QQ / Matrix e.g.

Python★ 17,926⑂ 1,602
View langbot-app/LangBot
16

MODSetter / SurfSense

Air gapped, open source NotebookLM alternative. Join our Discord: https://discord.gg/ejRNvftDp9

Python★ 16,169⑂ 1,538
View MODSetter/SurfSense
17

budtmo / docker-android

Android in docker solution with noVNC supported, video recording, mcp server and AI-agent

Python★ 15,874⑂ 1,754
View budtmo/docker-android
18

lidge-jun / opencodex

Universal provider proxy for OpenAI Codex & Claude Code — use any LLM (Claude, Gemini, Grok, DeepSeek, Ollama…) with Codex CLI, App, SDK, and Claude Code

TypeScript★ 15,561⑂ 1,180
View lidge-jun/opencodex
19

GaiZhenbiao / ChuanhuChatGPT

GUI for ChatGPT API and many LLMs. Supports agents, file-based QA, GPT finetuning and query with web search. All with a neat UI.

Python★ 15,272⑂ 2,195
View GaiZhenbiao/ChuanhuChatGPT
20

tonhowtf / omniget

Udemy & Hotmart course downloader, YouTube downloader (yt-dlp GUI, 1,800+ sites) + desktop app for AI agents: Claude Code, Codex, Gemini CLI, Ollama.

Rust★ 13,986⑂ 1,227
View tonhowtf/omniget
21

Open-LLM-VTuber / Open-LLM-VTuber

Talk to any LLM with hands-free voice interaction, voice interruption, and Live2D avatar running locally across platforms

Python★ 13,850⑂ 1,653
View Open-LLM-VTuber/Open-LLM-VTuber
22

browseros-ai / BrowserOS

🌐 The open-source Agentic browser; alternative to ChatGPT Atlas, Perplexity Comet, Dia.

TypeScript★ 13,720⑂ 1,460
View browseros-ai/BrowserOS
23

langchain4j / langchain4j

LangChain4j is an idiomatic, open-source Java library for building LLM-powered applications on the JVM.

Java★ 13,131⑂ 2,561
View langchain4j/langchain4j
24

The-PR-Agent / pr-agent

🚀 PR Agent: The Original Open-Source PR Reviewer. This project is not the Qodo free tier.

Python★ 13,075⑂ 1,869
View The-PR-Agent/pr-agent
25

StarTrail-org / LEANN

[MLsys2026 Best Paper]: https://arxiv.org/abs/2506.08276. RAG on Everything with LEANN. Enjoy 97% storage savings while running a fast, accurate

Python★ 12,945⑂ 1,169
View StarTrail-org/LEANN
26

TheR1D / shell_gpt

A command-line productivity tool powered by AI large language models like GPT-5, will help you accomplish your tasks faster and more efficiently.

Python★ 12,286⑂ 977
View TheR1D/shell_gpt
27

cactus-compute / needle

Automation foundation model for tiny devices: 2-bit, 8-29 MB, tool calls, structured extraction and embeddings on phones, wearables, smart homes, robots, cars and microcontrollers.

Python★ 11,808⑂ 756
View cactus-compute/needle
28

altic-dev / FluidVoice

Fastest and only macOS Dictation app with on-device STT and custom trained AI enhancement model. Windows pre-build available! A local Wispr Flow alternative.

Swift★ 11,660⑂ 834
View altic-dev/FluidVoice
29

getumbrel / llama-gpt

A self-hosted, offline, ChatGPT-like chatbot. Powered by Llama 2. 100% private, with no data leaving your device. New: Code Llama support!

TypeScript★ 10,937⑂ 704
View getumbrel/llama-gpt
30

ollama / ollama-python

Ollama Python library

Python★ 10,541⑂ 1,176
View ollama/ollama-python
31

sigoden / aichat

All-in-one LLM CLI tool featuring Shell Assistant, Chat-REPL, RAG, AI Tools & Agents, with access to OpenAI, Claude, Gemini, Ollama, Groq, and more.

Rust★ 10,460⑂ 745
View sigoden/aichat
32

RunanywhereAI / runanywhere-sdks

Production ready toolkit to run AI locally

C++★ 10,295⑂ 380
View RunanywhereAI/runanywhere-sdks
33

xorbitsai / inference

Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified

Python★ 9,580⑂ 867
View xorbitsai/inference
34

QwenAudio / SenseVoice

Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.

C★ 9,333⑂ 827
View QwenAudio/SenseVoice
35

miurla / morphic

An AI-powered search engine with a generative UI

TypeScript★ 9,133⑂ 2,348
View miurla/morphic
36

LearningCircuit / local-deep-research

~95% on SimpleQA (e.g. Qwen3.6-27B on a 3090). Supports all local and cloud LLMs (llama.cpp, Ollama, Google, ...). 10+ search engines - arXiv, PubMed, your private documents.

Python★ 9,109⑂ 824
View LearningCircuit/local-deep-research
37

reorproject / reor

Private & local AI personal knowledge management app for high entropy people.

JavaScript★ 8,552⑂ 528
View reorproject/reor
38

qualcomm / GenieX

Run frontier LLMs and VLMs locally on Qualcomm devices across NPU, GPU, and CPU with a few lines of code

Rust★ 8,381⑂ 1,055
View qualcomm/GenieX
39

n4ze3m / page-assist

Use your locally running AI models to assist you in your web browsing

TypeScript★ 8,220⑂ 784
View n4ze3m/page-assist
40

OpenCoworkAI / open-codesign

Open-source Claude Design alternative. One-click import your Claude Code / Codex API key. Prompt → prototype / slides / PDF. Multi-model (Claude, GPT, Gemini, Kimi, GLM, Ollama).

TypeScript★ 7,951⑂ 831
View OpenCoworkAI/open-codesign
41

mnfst / awesome-free-llm-apis

List of Permanent Free LLM API (API Keys)

JavaScript★ 7,887⑂ 753
View mnfst/awesome-free-llm-apis
42

ArvinLovegood / go-stock

🦄🦄🦄AI赋能股票分析:AI加持的股票分析/选股工具。股票行情获取,AI热点资讯分析,AI资金/财务分析,涨跌报警推送。支持A股,港股,美股。支持市场整体/个股情绪分析,AI辅助选股等。数据全部保留在本地。支持DeepSeek,OpenAI, Ollama,LMStudio,AnythingLLM,硅基流动,火山方舟,阿里云百炼等平台或模型。

Go★ 7,569⑂ 1,318
View ArvinLovegood/go-stock
43

JerryZLiu / Dayflow

The automatic work journal/time tracker. Privately turns your screen into a timeline of what you actually accomplished. Open-source and local-first.

Swift★ 7,148⑂ 438
View JerryZLiu/Dayflow
44

open-multi-agent / open-multi-agent

Self-hosted TypeScript agent runtime with durable approvals and verifiable run records. Own it, approve it, audit it.

TypeScript★ 6,944⑂ 2,433
View open-multi-agent/open-multi-agent
45

MakazhanAlpamys / Soup

Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.

Python★ 6,903⑂ 1,087
View MakazhanAlpamys/Soup
46

olimorris / codecompanion.nvim

✨ AI Coding, Vim Style

Lua★ 6,864⑂ 457
View olimorris/codecompanion.nvim
47

drumih / turbo-fieldfare

Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook

Swift★ 6,785⑂ 432
View drumih/turbo-fieldfare
48

vas3k / TaxHacker

Self-hosted AI accounting app. LLM analyzer for receipts, invoices, transactions with custom prompts and categories

TypeScript★ 6,714⑂ 1,093
View vas3k/TaxHacker
49

Osmantic / ODS

Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation.

Python★ 6,613⑂ 943
View Osmantic/ODS
50

google-ai-edge / LiteRT-LM

LiteRT-LM is Google's production-ready, high-performance, open-source inference framework for deploying Large Language Models on edge devices.

C++★ 6,482⑂ 725
View google-ai-edge/LiteRT-LM
51

cactus-compute / cactus

Quantization, kernels, runtime and inference engine for mobiles, wearables, smart home and robots.

C++★ 6,033⑂ 506
View cactus-compute/cactus
52

gluonfield / enchanted

Enchanted is iOS and macOS app for chatting with private self hosted language models such as Llama2, Mistral or Vicuna using Ollama.

Swift★ 6,002⑂ 426
View gluonfield/enchanted
53

clusterzx / paperless-ai

An automated document analyzer for Paperless-ngx using OpenAI API, Ollama, Deepseek-r1, Azure and all OpenAI API compatible Services to automatically analyze and tag your documents.

JavaScript★ 5,950⑂ 331
View clusterzx/paperless-ai
54

AgentOps-AI / agentops

Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frameworks including CrewAI, Agno, OpenAI Agents SDK, Langchain, Autogen, AG2

Python★ 5,833⑂ 625
View AgentOps-AI/agentops
55

PySpur-Dev / pyspur

A visual playground for agentic workflows: Iterate over your agents 10x faster

TypeScript★ 5,785⑂ 428
View PySpur-Dev/pyspur
56

dograh-hq / dograh

Open source voice AI platform. Self-hosted alternative to Vapi and Retell. On Prem, BYOK across Speech to Speech or LLM/STT/TTS, with a visual workflow builder, MCP native and telephony support.

Python★ 5,697⑂ 1,428
View dograh-hq/dograh
57

maziyarpanahi / openmed

Local-first healthcare AI: clinical NER & HIPAA PII de-identification that runs 100% on-device. 2,200+ medical models, 21 languages, Apple MLX + Python, no cloud

Python★ 5,346⑂ 680
View maziyarpanahi/openmed
58

buxuku / SmartSub

视频转字幕、字幕翻译、AI 配音与声音克隆、字幕烧录——免费开源的一站式桌面工具。基于 Whisper / FunASR 等本地模型离线语音转文字,批量处理 + 全平台 GPU 加速,跨 Windows / macOS / Linux。Free, open-source desktop app to generate, translate

TypeScript★ 5,265⑂ 400
View buxuku/SmartSub
59

Kiln-AI / Kiln

Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.

Python★ 5,078⑂ 380
View Kiln-AI/Kiln
60

vinta / pangu.js

Opinionated paranoid text spacing in JavaScript, with on-device AI semantic judgment

TypeScript★ 4,827⑂ 315
View vinta/pangu.js
61

thunderbird / thunderbolt

AI You Control: Choose your models. Own your data. Eliminate vendor lock-in.

TypeScript★ 4,775⑂ 327
View thunderbird/thunderbolt
62

u14app / deep-research

Use any LLMs (Large Language Models) for Deep Research. Support SSE API and MCP server.

JavaScript★ 4,688⑂ 1,062
View u14app/deep-research
63

JetBrains / koog

Koog is a JVM (Java and Kotlin) framework for building predictable, fault-tolerant and enterprise-ready AI agents across all platforms – from backend services to Android and iOS, JVM

Kotlin★ 4,580⑂ 474
View JetBrains/koog
64

CaviraOSS / LongMemory

Local persistent memory store for LLM applications including claude desktop, github copilot, codex, antigravity, etc.

TypeScript★ 4,510⑂ 504
View CaviraOSS/LongMemory
65

crmne / ruby_llm

The Ruby-native AI framework. Chats, agents, tools, images, audio, and video through one consistent API, in plain Ruby or Rails.

Ruby★ 4,390⑂ 504
View crmne/ruby_llm
66

ollama / ollama-js

Ollama JavaScript library

TypeScript★ 4,370⑂ 477
View ollama/ollama-js
67

umlx5h / LLPlayer

The media player for language learning, with dual subtitles, AI-generated subtitles, real-time translation, and more!

C#★ 4,191⑂ 249
View umlx5h/LLPlayer
68

GiovanniPasq / agentic-rag-for-dummies

A modular Agentic RAG built with LangGraph — learn Retrieval-Augmented Generation Agents in minutes.

Jupyter Notebook★ 4,189⑂ 551
View GiovanniPasq/agentic-rag-for-dummies
69

langroid / langroid

Harness LLMs with Multi-Agent Programming

Python★ 4,103⑂ 401
View langroid/langroid
70

nyldn / claude-octopus

Run multiple AI models against the same research, design, or coding task. Surface disagreements before you ship.

Shell★ 4,089⑂ 378
View nyldn/claude-octopus
71

yuruotong1 / autoMate

Like Manus, Computer Use Agent(CUA) and Omniparser, we are computer-using agents.AI-driven local automation assistant that uses natural language to make computers work by themselves

Python★ 3,964⑂ 490
View yuruotong1/autoMate
72

lemony-ai / cascadeflow

Cascading runtime for AI agents. Optimize cost, latency, quality, and policy decisions inside the agent loop.

Python★ 3,946⑂ 897
View lemony-ai/cascadeflow
73

claraverse-space / ClaraVerse

Claraverse is a opesource privacy focused ecosystem to replace ChatGPT, Claude, N8N, ImageGen with your own hosted llm, keys and compute. With desktop, IOS, Android Apps.

TypeScript★ 3,900⑂ 437
View claraverse-space/ClaraVerse
74

SciSharp / LLamaSharp

A C#/.NET library to run LLM (🦙LLaMA/LLaVA) on your local device efficiently.

C#★ 3,799⑂ 508
View SciSharp/LLamaSharp
75

raullenchai / Rapid-MLX

The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing.

Python★ 3,796⑂ 416
View raullenchai/Rapid-MLX
76

gofireflyio / aiac

Artificial Intelligence Infrastructure-as-Code Generator.

Go★ 3,788⑂ 296
View gofireflyio/aiac
77

pipeshub-ai / pipeshub-ai

PipesHub is an open-source platform for securely connecting enterprise knowledge to AI. Give AI agents trusted context and your team permission-aware search with verified citations across your

Python★ 3,762⑂ 579
View pipeshub-ai/pipeshub-ai
78

OvidijusParsiunas / deep-chat

Fully customizable AI chatbot component for your website

TypeScript★ 3,716⑂ 455
View OvidijusParsiunas/deep-chat
79

twinnydotdev / twinny

The most no-nonsense, locally or API-hosted AI code completion plugin for Visual Studio Code - like GitHub Copilot but 100% free.

TypeScript★ 3,650⑂ 229
View twinnydotdev/twinny
80

deta / surf

Personal AI Notebooks. Organize files & webpages and generate notes from them. Open source, local & open data, open model choice (incl. local).

TypeScript★ 3,572⑂ 251
View deta/surf
81

rashadphz / farfalle

🔍 AI search engine - self-host with local or cloud LLMs

TypeScript★ 3,541⑂ 318
View rashadphz/farfalle
82

entropy-research / Devon

Devon: An open-source pair programmer

Python★ 3,457⑂ 282
View entropy-research/Devon
83

QiuYannnn / Local-File-Organizer

An AI-powered file management tool that ensures privacy by organizing local texts, images. Using Llama3.2 3B and Llava v1.6 models with the Nexa SDK, it intuitively scans, restructures

Python★ 3,340⑂ 331
View QiuYannnn/Local-File-Organizer
84

nicedreamzapp / claude-code-local

Run Claude Code 100% on-device with local AI on Apple Silicon. MLX-native Anthropic-API server. 6 fighters incl.

Python★ 3,321⑂ 627
View nicedreamzapp/claude-code-local
85

rag-web-ui / rag-web-ui

RAG Web UI is an intelligent dialogue system based on RAG (Retrieval-Augmented Generation) technology.

TypeScript★ 3,292⑂ 369
View rag-web-ui/rag-web-ui
86

off-grid-ai / OGAM

The Swiss Army Knife of Offline AI. Chat, see, speak, and generate images on your phone or Mac — GGUF LLMs, vision, Whisper speech-to-text, Stable Diffusion, tool calling, and local-network servers.

TypeScript★ 3,129⑂ 301
View off-grid-ai/OGAM
87

SharpAI / DeepCamera

Open-Source AI Camera Skills Platform, AI NVR & CCTV Surveillance. Local VLM video analysis with Qwen, DeepSeek, SmolVLM, LLaVA, YOLO26.

JavaScript★ 3,066⑂ 480
View SharpAI/DeepCamera
88

xyproto / algernon

Small self-contained pure-Go web server with Lua, Teal, Markdown, HTTP/2, QUIC, Redis, TypeScript, npm-less React 19, SQLite, and PostgreSQL support ++

JavaScript★ 3,027⑂ 149
View xyproto/algernon
89

sherlockchou86 / VideoPipe

A cross-platform video structuring (video analysis) framework based on CV models & mLLM.

C++★ 2,955⑂ 466
View sherlockchou86/VideoPipe
90

ax-llm / ax

The pretty much "official" DSPy framework for Typescript

TypeScript★ 2,927⑂ 195
View ax-llm/ax
91

yakami129 / VirtualWife

VirtualWife是一个虚拟数字人项目,支持B站直播,支持openai、ollama

Python★ 2,899⑂ 443
View yakami129/VirtualWife
92

rafska / awesome-local-llm

A curated list of awesome platforms, tools, practices and resources that helps run LLMs locally

★ 2,870⑂ 388
View rafska/awesome-local-llm
93

Luce-Org / lucebox

LLM speculative inference server for heterogeneous hardware & consumer GPUs

C++★ 2,868⑂ 277
View Luce-Org/lucebox
94

Mininglamp-AI / Mano-P

Mano-P: Open-source GUI-VLA agent for edge devices. #1 on OSWorld (specialized, 58.2%). Runs locally on Apple M4 Mac mini/MacBook

★ 2,776⑂ 272
View Mininglamp-AI/Mano-P
95

control-theory / gonzo

Gonzo! The Go based TUI log analysis tool

Go★ 2,769⑂ 111
View control-theory/gonzo
96

icereed / paperless-gpt

Use LLMs and LLM Vision (OCR) to handle paperless-ngx - Document Digitalization powered by AI

Go★ 2,704⑂ 209
View icereed/paperless-gpt
97

Mobile-Artificial-Intelligence / maid

Maid is a free and open source application for interfacing with llama.cpp models locally, and with Anthropic, DeepSeek, Ollama, Mistral and OpenAI models remotely.

TypeScript★ 2,695⑂ 289
View Mobile-Artificial-Intelligence/maid
98

nicepkg / aide

Conquer Any Code in VSCode: One-Click Comments, Conversions, UI-to-Code, and AI Batch Processing of Files! 在 VSCode 中征服任何代码:一键注释、转换、UI 图生成代码、AI 批量处理文件!💪

TypeScript★ 2,688⑂ 208
View nicepkg/aide
99

codedogQBY / ReadAny

AI-powered cross-platform e-book reader with semantic search, RAG chat, local vector store, notes, TTS, and WebDAV sync.

TypeScript★ 2,625⑂ 203
View codedogQBY/ReadAny
100

itayinbarr / little-coder

A harness optimized to smaller LLMs

TypeScript★ 2,606⑂ 179
View itayinbarr/little-coder

Local Model Runners & Servers

Chat UIs & Desktop Assistants

On-Device & Mobile AI

Quantization & Model Compression

More AI Trending Projects

1

cloudflare / security-audit-skill

JavaScript★ 17,418⑂ 959▲ 3,155 stars
2

affaan-m / ECC

JavaScript★ 263,267⑂ 39,395▲ 1,012 stars
3

Tencent / BrowserSkill

TypeScript★ 5,916⑂ 417▲ 612 stars
4

vercel-labs / json-render

TypeScript★ 16,987⑂ 905▲ 585 stars
5

addyosmani / agent-skills

JavaScript★ 97,396⑂ 10,274▲ 556 stars
6

anthropics / claude-code

TypeScript★ 146,908⑂ 23,988▲ 483 stars

Running an LLM locally: what you actually need

A local model needs three things: weights in a format the runtime understands, enough memory to hold them, and a runner to load them. The rule of thumb for four-bit quantization is roughly two thirds of a gigabyte per billion parameters, plus room for the context window — so a 7B model fits comfortably in 8 GB of unified memory, while a 70B model does not fit on a laptop at all without aggressive compression or offloading into system RAM. Server-grade engines change that arithmetic with paging, continuous batching and GPU offload, which is why the same weights can be impractical in one tool and comfortable in another.

Start with a runner rather than a chat client. Once the model answers from the command line, everything above it is interface.

Local chat UIs, desktop assistants and on-device inference

The desktop layer is what most people actually want: a chat window that keeps transcripts on disk, can read local files, and talks to whichever local server is running. Those clients are deliberately model-agnostic — they speak the same HTTP API as the popular runners — so the choice of client and the choice of model are independent decisions. On phones the constraints change completely: memory is small, thermals are real, and the interesting projects are the ones shipping compressed models that run in a few hundred megabytes with acceptable latency, rather than the ones streaming a full-size model over the network.

Quantization: why the same model runs on your laptop

Quantization stores weights at lower precision — four bits instead of sixteen — which cuts memory use roughly fourfold and speeds up inference on hardware with no tensor cores. It is not free: below four bits the quality loss becomes visible on reasoning tasks, and some quantized formats only load in specific runtimes. The practical guidance is boring and effective: prefer four-bit formats for local use, keep a higher-precision copy on a server if you need to compare, and test the quantized build on your own inputs instead of trusting a benchmark table.

Local AI questions

What is a local LLM? A language model whose weights run on your own hardware, so prompts and outputs never leave the machine and there is no per-token cost.

Can a local chatbot work offline? Yes, once the model has been downloaded. Runners and desktop clients on this board keep working with the network unplugged; only features that fetch from the internet stop.

What is on-device AI? Inference that happens on the device that produced the data — phone, laptop or embedded board — rather than in a hosted service. Its binding constraints are memory bandwidth and thermal budget.

How much RAM do I need to run a model locally? Roughly two thirds of a gigabyte per billion parameters at four-bit quantization, plus context. An 8B model is comfortable in 8 GB; anything past 30B wants a workstation or a GPU with real VRAM.

Is a local model as good as a hosted one? For summarising, drafting and light coding a good 8B–14B model is often enough. For hard multi-step reasoning the largest hosted models still lead, and they lead by more the longer the task runs.

Related AI Trending Lists