ollama / ollama
Get up and running with Kimi, GLM, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
View ollama/ollamaRunning AI locally means the model, the weights and the data all stay on hardware you control: no API key, no per-token bill, no prompt leaving the machine. This board collects the projects that make that practical — local model runners and servers, chat interfaces and desktop assistants, mobile and on-device inference, and the quantization and compression tooling that lets a model fit into laptop or phone memory. Entries come from GitHub topic pages for local LLMs, llama.cpp-style inference, on-device AI and Ollama, ranked by stars, so a mature server sits next to a desktop client published last month. Each card lists language, stars and forks and opens a page with description, license, activity dates, README and related local AI projects. The practical question is almost always memory: check the quantization notes on the detail page before assuming a model will fit.
Get up and running with Kimi, GLM, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
View ollama/ollamaUser-friendly AI Interface (Supports Ollama, OpenAI API, ...)
View open-webui/open-webui✨ Zero-config AI chat assistant. No API key needed — sign up and instantly chat with GPT-5, Claude 4, Gemini 2.5, DeepSeek & 100+ top models. Pay-as-you-go saves you more.
View ChatGPTNextWeb/NextChatUltra-lightweight, open-source, self-hosted personal AI agent framework in Python with WebUI, tools, memory, MCP, multi-agent workflows, automation, and chat apps
View HKUDS/nanobotLangchain-Chatchat(原Langchain-ChatGLM)基于 Langchain 与 ChatGLM, Qwen 与 Llama 等语言模型的 RAG 与 Agent 应用 | Langchain-Chatchat (formerly langchain-ChatGLM), local knowledge based LLM (like ChatGLM
View chatchat-space/Langchain-Chatchat🔥 1Panel is a modern, open-source Linux server management panel and a lightweight AI management platform.
View 1Panel-dev/1PanelPrivacy first, AI meeting assistant with 4x faster Parakeet/Whisper live transcription, speaker diarization, and Ollama summarization built on Rust. 100% local processing. no cloud required.
View Zackriya-Solutions/meetilyOpen-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.
View Tencent/WeKnora🔥 MaxKB is an open-source platform for building enterprise-grade agents. 强大易用的开源企业级智能体平台。
View 1Panel-dev/MaxKBLocal, open-source AI app builder for power users ✨ v0 / Lovable / Replit / Bolt alternative 🌟 Star if you like it!
View dyad-sh/dyad🤖 可 DIY 的 多模态 AI 聊天机器人 | 🚀 快速接入 微信、 QQ、Telegram、等聊天平台 | 🦈支持DeepSeek、Grok、Claude、Ollama、Gemini、OpenAI | 工作流系统、网页搜索、AI画图、人设调教、虚拟女仆、语音对话 |
View lss233/kirara-aiOpen Source DeepWiki: AI-Powered Wiki Generator for GitHub/Gitlab/Bitbucket Repositories. Join the discord: https://discord.gg/gMwThUMeme
View AsyncFuncAI/deepwiki-openProduction-grade platform for building agentic IM bots - 生产级多平台智能机器人开发平台/ Agent、知识库编排、插件系统 / Bots for Discord / Slack / LINE / Telegram / WeChat(企业微信, 企微智能机器人, 公众号) / 飞书 / 钉钉 / QQ / Matrix e.g.
View langbot-app/LangBotAir gapped, open source NotebookLM alternative. Join our Discord: https://discord.gg/ejRNvftDp9
View MODSetter/SurfSenseAndroid in docker solution with noVNC supported, video recording, mcp server and AI-agent
View budtmo/docker-androidUniversal provider proxy for OpenAI Codex & Claude Code — use any LLM (Claude, Gemini, Grok, DeepSeek, Ollama…) with Codex CLI, App, SDK, and Claude Code
View lidge-jun/opencodexGUI for ChatGPT API and many LLMs. Supports agents, file-based QA, GPT finetuning and query with web search. All with a neat UI.
View GaiZhenbiao/ChuanhuChatGPTUdemy & Hotmart course downloader, YouTube downloader (yt-dlp GUI, 1,800+ sites) + desktop app for AI agents: Claude Code, Codex, Gemini CLI, Ollama.
View tonhowtf/omnigetTalk to any LLM with hands-free voice interaction, voice interruption, and Live2D avatar running locally across platforms
View Open-LLM-VTuber/Open-LLM-VTuber🌐 The open-source Agentic browser; alternative to ChatGPT Atlas, Perplexity Comet, Dia.
View browseros-ai/BrowserOSLangChain4j is an idiomatic, open-source Java library for building LLM-powered applications on the JVM.
View langchain4j/langchain4j🚀 PR Agent: The Original Open-Source PR Reviewer. This project is not the Qodo free tier.
View The-PR-Agent/pr-agent[MLsys2026 Best Paper]: https://arxiv.org/abs/2506.08276. RAG on Everything with LEANN. Enjoy 97% storage savings while running a fast, accurate
View StarTrail-org/LEANNA command-line productivity tool powered by AI large language models like GPT-5, will help you accomplish your tasks faster and more efficiently.
View TheR1D/shell_gptAutomation foundation model for tiny devices: 2-bit, 8-29 MB, tool calls, structured extraction and embeddings on phones, wearables, smart homes, robots, cars and microcontrollers.
View cactus-compute/needleFastest and only macOS Dictation app with on-device STT and custom trained AI enhancement model. Windows pre-build available! A local Wispr Flow alternative.
View altic-dev/FluidVoiceA self-hosted, offline, ChatGPT-like chatbot. Powered by Llama 2. 100% private, with no data leaving your device. New: Code Llama support!
View getumbrel/llama-gptAll-in-one LLM CLI tool featuring Shell Assistant, Chat-REPL, RAG, AI Tools & Agents, with access to OpenAI, Claude, Gemini, Ollama, Groq, and more.
View sigoden/aichatProduction ready toolkit to run AI locally
View RunanywhereAI/runanywhere-sdksSwap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified
View xorbitsai/inferenceOpen-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.
View QwenAudio/SenseVoice~95% on SimpleQA (e.g. Qwen3.6-27B on a 3090). Supports all local and cloud LLMs (llama.cpp, Ollama, Google, ...). 10+ search engines - arXiv, PubMed, your private documents.
View LearningCircuit/local-deep-researchPrivate & local AI personal knowledge management app for high entropy people.
View reorproject/reorRun frontier LLMs and VLMs locally on Qualcomm devices across NPU, GPU, and CPU with a few lines of code
View qualcomm/GenieXUse your locally running AI models to assist you in your web browsing
View n4ze3m/page-assistOpen-source Claude Design alternative. One-click import your Claude Code / Codex API key. Prompt → prototype / slides / PDF. Multi-model (Claude, GPT, Gemini, Kimi, GLM, Ollama).
View OpenCoworkAI/open-codesignList of Permanent Free LLM API (API Keys)
View mnfst/awesome-free-llm-apis🦄🦄🦄AI赋能股票分析:AI加持的股票分析/选股工具。股票行情获取,AI热点资讯分析,AI资金/财务分析,涨跌报警推送。支持A股,港股,美股。支持市场整体/个股情绪分析,AI辅助选股等。数据全部保留在本地。支持DeepSeek,OpenAI, Ollama,LMStudio,AnythingLLM,硅基流动,火山方舟,阿里云百炼等平台或模型。
View ArvinLovegood/go-stockThe automatic work journal/time tracker. Privately turns your screen into a timeline of what you actually accomplished. Open-source and local-first.
View JerryZLiu/DayflowSelf-hosted TypeScript agent runtime with durable approvals and verifiable run records. Own it, approve it, audit it.
View open-multi-agent/open-multi-agentFine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.
View MakazhanAlpamys/SoupGemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
View drumih/turbo-fieldfareSelf-hosted AI accounting app. LLM analyzer for receipts, invoices, transactions with custom prompts and categories
View vas3k/TaxHackerTurn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation.
View Osmantic/ODSLiteRT-LM is Google's production-ready, high-performance, open-source inference framework for deploying Large Language Models on edge devices.
View google-ai-edge/LiteRT-LMQuantization, kernels, runtime and inference engine for mobiles, wearables, smart home and robots.
View cactus-compute/cactusEnchanted is iOS and macOS app for chatting with private self hosted language models such as Llama2, Mistral or Vicuna using Ollama.
View gluonfield/enchantedAn automated document analyzer for Paperless-ngx using OpenAI API, Ollama, Deepseek-r1, Azure and all OpenAI API compatible Services to automatically analyze and tag your documents.
View clusterzx/paperless-aiPython SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frameworks including CrewAI, Agno, OpenAI Agents SDK, Langchain, Autogen, AG2
View AgentOps-AI/agentopsA visual playground for agentic workflows: Iterate over your agents 10x faster
View PySpur-Dev/pyspurOpen source voice AI platform. Self-hosted alternative to Vapi and Retell. On Prem, BYOK across Speech to Speech or LLM/STT/TTS, with a visual workflow builder, MCP native and telephony support.
View dograh-hq/dograhLocal-first healthcare AI: clinical NER & HIPAA PII de-identification that runs 100% on-device. 2,200+ medical models, 21 languages, Apple MLX + Python, no cloud
View maziyarpanahi/openmed视频转字幕、字幕翻译、AI 配音与声音克隆、字幕烧录——免费开源的一站式桌面工具。基于 Whisper / FunASR 等本地模型离线语音转文字,批量处理 + 全平台 GPU 加速,跨 Windows / macOS / Linux。Free, open-source desktop app to generate, translate
View buxuku/SmartSubBuild, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.
View Kiln-AI/KilnOpinionated paranoid text spacing in JavaScript, with on-device AI semantic judgment
View vinta/pangu.jsAI You Control: Choose your models. Own your data. Eliminate vendor lock-in.
View thunderbird/thunderboltUse any LLMs (Large Language Models) for Deep Research. Support SSE API and MCP server.
View u14app/deep-researchKoog is a JVM (Java and Kotlin) framework for building predictable, fault-tolerant and enterprise-ready AI agents across all platforms – from backend services to Android and iOS, JVM
View JetBrains/koogLocal persistent memory store for LLM applications including claude desktop, github copilot, codex, antigravity, etc.
View CaviraOSS/LongMemoryThe Ruby-native AI framework. Chats, agents, tools, images, audio, and video through one consistent API, in plain Ruby or Rails.
View crmne/ruby_llmThe media player for language learning, with dual subtitles, AI-generated subtitles, real-time translation, and more!
View umlx5h/LLPlayerA modular Agentic RAG built with LangGraph — learn Retrieval-Augmented Generation Agents in minutes.
View GiovanniPasq/agentic-rag-for-dummiesRun multiple AI models against the same research, design, or coding task. Surface disagreements before you ship.
View nyldn/claude-octopusLike Manus, Computer Use Agent(CUA) and Omniparser, we are computer-using agents.AI-driven local automation assistant that uses natural language to make computers work by themselves
View yuruotong1/autoMateCascading runtime for AI agents. Optimize cost, latency, quality, and policy decisions inside the agent loop.
View lemony-ai/cascadeflowClaraverse is a opesource privacy focused ecosystem to replace ChatGPT, Claude, N8N, ImageGen with your own hosted llm, keys and compute. With desktop, IOS, Android Apps.
View claraverse-space/ClaraVerseA C#/.NET library to run LLM (🦙LLaMA/LLaVA) on your local device efficiently.
View SciSharp/LLamaSharpThe fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing.
View raullenchai/Rapid-MLXPipesHub is an open-source platform for securely connecting enterprise knowledge to AI. Give AI agents trusted context and your team permission-aware search with verified citations across your
View pipeshub-ai/pipeshub-aiFully customizable AI chatbot component for your website
View OvidijusParsiunas/deep-chatThe most no-nonsense, locally or API-hosted AI code completion plugin for Visual Studio Code - like GitHub Copilot but 100% free.
View twinnydotdev/twinnyPersonal AI Notebooks. Organize files & webpages and generate notes from them. Open source, local & open data, open model choice (incl. local).
View deta/surfAn AI-powered file management tool that ensures privacy by organizing local texts, images. Using Llama3.2 3B and Llava v1.6 models with the Nexa SDK, it intuitively scans, restructures
View QiuYannnn/Local-File-OrganizerRun Claude Code 100% on-device with local AI on Apple Silicon. MLX-native Anthropic-API server. 6 fighters incl.
View nicedreamzapp/claude-code-localRAG Web UI is an intelligent dialogue system based on RAG (Retrieval-Augmented Generation) technology.
View rag-web-ui/rag-web-uiThe Swiss Army Knife of Offline AI. Chat, see, speak, and generate images on your phone or Mac — GGUF LLMs, vision, Whisper speech-to-text, Stable Diffusion, tool calling, and local-network servers.
View off-grid-ai/OGAMOpen-Source AI Camera Skills Platform, AI NVR & CCTV Surveillance. Local VLM video analysis with Qwen, DeepSeek, SmolVLM, LLaVA, YOLO26.
View SharpAI/DeepCameraSmall self-contained pure-Go web server with Lua, Teal, Markdown, HTTP/2, QUIC, Redis, TypeScript, npm-less React 19, SQLite, and PostgreSQL support ++
View xyproto/algernonA cross-platform video structuring (video analysis) framework based on CV models & mLLM.
View sherlockchou86/VideoPipeA curated list of awesome platforms, tools, practices and resources that helps run LLMs locally
View rafska/awesome-local-llmLLM speculative inference server for heterogeneous hardware & consumer GPUs
View Luce-Org/luceboxMano-P: Open-source GUI-VLA agent for edge devices. #1 on OSWorld (specialized, 58.2%). Runs locally on Apple M4 Mac mini/MacBook
View Mininglamp-AI/Mano-PUse LLMs and LLM Vision (OCR) to handle paperless-ngx - Document Digitalization powered by AI
View icereed/paperless-gptMaid is a free and open source application for interfacing with llama.cpp models locally, and with Anthropic, DeepSeek, Ollama, Mistral and OpenAI models remotely.
View Mobile-Artificial-Intelligence/maidConquer Any Code in VSCode: One-Click Comments, Conversions, UI-to-Code, and AI Batch Processing of Files! 在 VSCode 中征服任何代码:一键注释、转换、UI 图生成代码、AI 批量处理文件!💪
View nicepkg/aideAI-powered cross-platform e-book reader with semantic search, RAG chat, local vector store, notes, TTS, and WebDAV sync.
View codedogQBY/ReadAnyA local model needs three things: weights in a format the runtime understands, enough memory to hold them, and a runner to load them. The rule of thumb for four-bit quantization is roughly two thirds of a gigabyte per billion parameters, plus room for the context window — so a 7B model fits comfortably in 8 GB of unified memory, while a 70B model does not fit on a laptop at all without aggressive compression or offloading into system RAM. Server-grade engines change that arithmetic with paging, continuous batching and GPU offload, which is why the same weights can be impractical in one tool and comfortable in another.
Start with a runner rather than a chat client. Once the model answers from the command line, everything above it is interface.
The desktop layer is what most people actually want: a chat window that keeps transcripts on disk, can read local files, and talks to whichever local server is running. Those clients are deliberately model-agnostic — they speak the same HTTP API as the popular runners — so the choice of client and the choice of model are independent decisions. On phones the constraints change completely: memory is small, thermals are real, and the interesting projects are the ones shipping compressed models that run in a few hundred megabytes with acceptable latency, rather than the ones streaming a full-size model over the network.
Quantization stores weights at lower precision — four bits instead of sixteen — which cuts memory use roughly fourfold and speeds up inference on hardware with no tensor cores. It is not free: below four bits the quality loss becomes visible on reasoning tasks, and some quantized formats only load in specific runtimes. The practical guidance is boring and effective: prefer four-bit formats for local use, keep a higher-precision copy on a server if you need to compare, and test the quantized build on your own inputs instead of trusting a benchmark table.
What is a local LLM? A language model whose weights run on your own hardware, so prompts and outputs never leave the machine and there is no per-token cost.
Can a local chatbot work offline? Yes, once the model has been downloaded. Runners and desktop clients on this board keep working with the network unplugged; only features that fetch from the internet stop.
What is on-device AI? Inference that happens on the device that produced the data — phone, laptop or embedded board — rather than in a hosted service. Its binding constraints are memory bandwidth and thermal budget.
How much RAM do I need to run a model locally? Roughly two thirds of a gigabyte per billion parameters at four-bit quantization, plus context. An 8B model is comfortable in 8 GB; anything past 30B wants a workstation or a GPU with real VRAM.
Is a local model as good as a hosted one? For summarising, drafting and light coding a good 8B–14B model is often enough. For hard multi-step reasoning the largest hosted models still lead, and they lead by more the longer the task runs.