mostlygeek/llama-swap
Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc
About mostlygeek/llama-swap
mostlygeek/llama-swap is an open-source project on GitHub, mainly written in Go. Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc It currently holds 5,830 stars and 490 forks with 0 open issues, and was last pushed on an unknown date (repository created unknown).
Project Overview
AI Homed tracks it on the Today's Trending board, currently at rank #44 with 17 new stars today.
GitHub Repository Details
README
llama-swap
Run multiple generative AI models on your machine and hot-swap between them on demand. llama-swap works with any OpenAI and Anthropic API compatible server and is used by thousands of people to power their local AI workflows.
Built in Go for performance and simplicity, llama-swap has zero dependencies and is incredibly easy to set up. Get started in minutes - just one binary and one configuration file.
Features:
- ✅ Easy to deploy and configure: one binary, one configuration file. no external dependencies
- ✅ On-demand model switching for many local AI servers (llama.cpp + forks, vllm, stable-diffusion.cpp, audio.cpp, ComfyUI, etc.)
- future proof, upgrade your inference servers at any time.
- ✅ OpenAI API supported endpoints:
v1/completionsv1/chat/completionsv1/responsesv1/embeddingsv1/models- list available modelsv1/audio/speech(#36)v1/audio/transcriptions(docs)v1/audio/voicesv1/images/generationsv1/images/edits- ✅ Anthropic API supported endpoints:
v1/messagesv1/messages/count_tokens- ✅ llama-server (llama.cpp) supported endpoints
v1/rerank,v1/reranking,/rerank/infill- for code infilling/completion- for completion endpoint/models- list available models. same behavior asv1/models/props- requires?model={model_id}query parameter to be provided. The autoload parameter is not supported and will be ignored.- ✅ SDAPI via stable-diffusion.cpp's server
/sdapi/v1/txt2img/sdapi/v1/img2img/sdapi/v1/loras- requiresmodelin request body to fetch the correct loras- ✅ audio.cpp supported extra endpoints
/audioapi/v1/tasks/run- ✅
/comfyui/- ComfyUI custom endpoint (#1001) for more reliable swapping - ✅ llama-swap API
/ui- web UI/upstream/:model_id- direct access to upstream server (demo)/running- list currently running models (#61)POST /api/models/unload- manually unload all running models (#58)POST /api/models/unload/:model_id- unload a specific modelGET /api/profiles- list configured profiles and the active selectionPUT /api/profiles/active- activate a profile or select none/logs- remote log monitoringGET /logsreturns buffered plain text logs.- If
Accept: text/htmlis sent,/logsredirects to/ui/. GET /logs/streamkeeps the connection open for live log streaming.- Stream endpoints send buffered history first by default; add
?no-historyto stream only new lines. GET /logs/stream/proxystreams proxy logs only.GET /logs/stream/upstreamstreams upstream process logs only.GET /logs/stream/httpstreams the HTTP access log only.GET /logs/stream/{model_id}streams logs for one model (including IDs with slashes, likeauthor/model)./health- just returns "OK"/metrics- system and GPU metrics for prometheus- ✅ API Key support - define keys to restrict access to API endpoints
- ✅ Customization
- Switch model ID routing at runtime with profiles
- Run concurrent models with a custom DSL swap matrix (#643)
- Automatic unloading of models after timeout by setting a
ttl - Docker and Podman support using
cmdandcmdStoptogether - Preload models on startup with
hooks(#235) - Apply filters to requests to control inference with
stripParams,setParamsandsetParamsByID
Web UI
llama-swap includes a real time web interface with a playground for testing out all sorts of local models:
View detailed token metrics:
Inspect request and responses:
Manually load and unload models:
Real time log streaming:
Installation
llama-swap can be installed in multiple ways
1. Docker 2. Homebrew (macOS and Linux) 3. MacPorts (macOS) 4. WinGet 5. From release binaries 6. From source
Docker Install (download images)
Two types of container images are built nightly for llama-swap:
1. A unified container with llama-server, ik-llama-server, stable-diffusion.cpp, whisper.cpp, audio.cpp and llama-swap all built from source. Available for CUDA 12, CUDA 13 and Vulkan. This one is recommended for use.
2. A legacy image, which is llama.cpp's own llama-server container with llama-swap copied in. It carries only what that base image ships, so no image generation, speech or ik-llama-server.
Unified container (Recommended)
There are three unified images. Pick the one that matches your GPU:
| tag | platforms | GPUs |
| --- | --- | --- |
| unified-cuda13 | amd64, arm64 | NVIDIA Ampere through Blackwell: A100, RTX 30xx/40xx/50xx, H100, RTX PRO, and GB10 on DGX Spark. Built with CUDA 13. |
| unified-cuda | amd64 | NVIDIA Pascal through Ada: P40, P100, GTX 10xx, RTX 20xx/30xx/40xx. Built with CUDA 12, for cards CUDA 13 dropped. |
| unified-vulkan | amd64 | AMD and other Vulkan capable GPUs. |
unified-cuda13 is a multi-arch tag, so docker pull picks the right image for
the host — including the aarch64 one a DGX Spark needs.
$ docker pull ghcr.io/mostlygeek/llama-swap:unified-cuda13
run with a custom configuration and models directory
$ docker run -it --rm --runtime nvidia -p 9292:8080 \
-v /path/to/models:/models \
-v /path/to/custom/config.yaml:/etc/llama-swap/config/config.yaml \
ghcr.io/mostlygeek/llama-swap:unified-cuda13
Configuring startup with environment variables
The unified images can be configured with LLAMA_SWAP_* environment variables
instead of command line flags, which is usually easier in a compose file or a
Kubernetes manifest. Each one maps to a llama-swap flag:
| variable | flag | default in the image |
| --- | --- | --- |
| LLAMA_SWAP_CONFIG | -config | /etc/llama-swap/config/config.yaml |
| LLAMA_SWAP_CONFIG_DIR | -config-dir | — |
| LLAMA_SWAP_LISTEN | -listen | 0.0.0.0:8080 |
| LLAMA_SWAP_TLS_CERT_FILE | -tls-cert-file | — |
| LLAMA_SWAP_TLS_KEY_FILE | -tls-key-file | — |
| LLAMA_SWAP_LISTEN_TAILCAT | -listen-tailcat | — |
| LLAMA_SWAP_WATCH_CONFIG | -watch-config | true |
# configure startup with environment variables instead of flags
$ docker run -it --rm --runtime nvidia -p 9292:8080 \
-v /path/to/models:/models \
-e LLAMA_SWAP_CONFIG=/models/llama-swap.yaml \
-e LLAMA_SWAP_LISTEN=0.0.0.0:8080 \
-e LLAMA_SWAP_WATCH_CONFIG=false \
ghcr.io/mostlygeek/llama-swap:unified-cuda13
services:
llama-swap:
image: ghcr.io/mostlygeek/llama-swap:unified-cuda13
ports:
- "9292:8080"
volumes:
- /path/to/models:/models
environment:
LLAMA_SWAP_CONFIG: /models/llama-swap.yaml
LLAMA_SWAP_LISTEN: 0.0.0.0:8080
LLAMA_SWAP_WATCH_CONFIG: "false"
Keep LLAMA_SWAP_LISTEN on 0.0.0.0 whenever you publish a port. A
container that binds its own loopback is not reachable through -p, so
localhost:8080 there gives connection refused. Bind loopback only when the
container shares the host's network, where it usefully limits llama-swap to the
host itself:
# reachable from the host only, not from the network
$ docker run -it --rm --runtime nvidia --network host \
-v /path/to/models:/models \
-e LLAMA_SWAP_LISTEN=localhost:8080 \
ghcr.io/mostlygeek/llama-swap:unified-cuda13
An unset or empty variable contributes nothing and leaves llama-swap's own
default. Booleans accept true/false, 1/0, yes/no or on/off in any
case; anything else stops the container rather than being read as "off".
-version is deliberately not mapped — use docker run -version.
Passing flags to the container still works and behaves exactly as it always has: arguments replace every default rather than adding to the variables above.
# unchanged: runs llama-swap -config /models/my.yaml, nothing else added
$ docker run ghcr.io/mostlygeek/llama-swap:unified-cuda13 -config /models/my.yaml
See docker/unified/README.md for the details.
Legacy container
This image is llama.cpp's own llama-server container
(ghcr.io/ggml-org/llama.cpp)
with the llama-swap binary copied into it. It tracks llama.cpp's nightly server
images closely, which is its only real advantage.
Prefer the unified container. The legacy image inherits whatever the
llama-server base ships and nothing else, so it has no stable-diffusion.cpp,
whisper.cpp, audio.cpp or ik-llama-server, and no LLAMA_SWAP_* environment
variable support. Use it only if staying on llama.cpp's exact server image
matters to you.
$ docker pull ghcr.io/mostlygeek/llama-swap:cuda
run with a custom configuration and models directory
$ docker run -it --rm --runtime nvidia -p 9292:8080 \
-v /path/to/models:/models \
-v /path/to/custom/config.yaml:/app/config.yaml \
ghcr.io/mostlygeek/llama-swap:cuda
more examples
# pull latest images per platform
docker pull ghcr.io/mostlygeek/llama-swap:cpu
docker pull ghcr.io/mostlygeek/llama-swap:cuda
docker pull ghcr.io/mostlygeek/llama-swap:vulkan
docker pull ghcr.io/mostlygeek/llama-swap:intel
docker pull ghcr.io/mostlygeek/llama-swap:musa
tagged llama-swap, platform and llama-server version images
docker pull ghcr.io/mostlygeek/llama-swap:v166-cuda-b6795
non-root cuda
docker pull ghcr.io/mostlygeek/llama-swap:cuda-non-root
Homebrew Install (macOS/Linux)
brew tap mostlygeek/llama-swap
brew install llama-swap
llama-swap --config path/to/config.yaml --listen localhost:8080
MacPorts (macOS)
[!NOTE]
Maintained by MacPorts community - llama-swap port. It is not an official part of llama-swap.
sudo port install llama-swap
llama-swap --config path/to/config.yaml --listen localhost:8080
WinGet Install (Windows)
[!NOTE]
WinGet is maintained by community contributor Dvd-Znf (#327). It is not an official part of llama-swap.
# install
C:\> winget install llama-swap
upgrade
C:\> winget upgrade llama-swap
Pre-built Binaries
Binaries are available on the release page for Linux, Mac, Windows and FreeBSD.
Building from source
1. Building requires Go and Node.js (for UI).
1. git clone https://github.com/mostlygeek/llama-swap.git
1. make clean all
1. look in the build/ subdirectory for the llama-swap binary
Configuration
# minimum viable config.yaml
models:
model1:
cmd: llama-server --port ${PORT} --model /path/to/model.gguf
That's all you need to get started:
1. models - holds all model configurations
2. model1 - the ID used in API calls
3. cmd - the command to run to start the server.
4. ${PORT} - an automatically assigned port number
Almost all configuration settings are optional and can be added one step at a time:
- Advanced features
matrixto run concurrent models with a custom swap logic DSLhooksto run things on startupmacrosreusable snippets- Model customization
ttlto automatically unload modelsunloadTimeoutto tune graceful unloads (manual, API andttlexpiry)aliasesto use familiar model names (e.g., "gpt-4o-mini")envto pass custom environment variables to inference serverscmdStopgracefully stop Docker/Podman containersuseModelNameto override model names sent to upstream servers${PORT}automatic port variables for dynamic port assignmentfiltersrewrite parts of requests before sending to the upstream server
You can also just ask. The Help page (in the sidebar) is an agent that calls
llama-swap's own documentation tools and answers questions about your
configuration using the real text of config.example.yaml and the knowledge
base, running entirely on a local model. Pick a tool-capable model on the
Help page and ask away — see
Writing the cmd for a model if the model
answers without calling anything (older llama-server builds need --jinja added explicitly).
Those same tools are served as an MCP endpoint at /api/mcp, so any MCP client
can ask about your configuration too. See
Connecting an MCP client.
How does llama-swap work?
When a request is made to an OpenAI compatible endpoint, llama-swap will extract the model value and load the appropriate server configuration to serve it. If the wrong upstream server is running, it will be replaced with the correct one. This is where the "swap" part comes in. The upstream server is automatically swapped to handle the request correctly.
In the most basic configuration llama-swap handles one model at a time. For more advanced use cases, using a matrix allows multiple models to be loaded at the same time. You have complete control over how your system resources are used.
Reverse Proxy Configuration (nginx)
If you deploy llama-swap behind nginx, disable response buffering for streaming endpoints. By default, nginx buffers responses which breaks Server‑Sent Events (SSE) and streaming chat completion. (#236)
Recommended nginx configuration snippets:
# SSE for UI events and logs (also covers /api/events/logs)
location /api/events {
proxy_pass http://your-llama-swap-backend;
proxy_buffering off;
proxy_cache off;
}
Streaming chat completions (stream=true)
location /v1/chat/completions {
proxy_pass http://your-llama-swap-backend;
proxy_buffering off;
proxy_cache off;
}
As a safeguard, llama-swap also sets X-Accel-Buffering: no on SSE responses. However, explicitly disabling proxy_buffering at your reverse proxy is still recommended for reliable streaming behavior.
Monitoring Logs on the CLI
# sends up to the last 10KB of logs
$ curl http://host/logs
streams combined logs
curl -Ns http://host/logs/stream
stream llama-swap's proxy status logs
curl -Ns http://host/logs/stream/proxy
stream logs from upstream processes that llama-swap loads
curl -Ns http://host/logs/stream/upstream
stream the HTTP access log, one line per request
curl -Ns http://host/logs/stream/http
stream logs only from a specific model
curl -Ns http://host/logs/stream/{model_id}
stream and filter logs with linux pipes
curl -Ns http://host/logs/stream | grep 'eval time'
appending ?no-history will disable sending buffered history first
curl -Ns 'http://host/logs/stream?no-history'
Do I need to use llama.cpp's server (llama-server)?
Any OpenAI compatible server would work. llama-swap was originally designed for llama-server and it is the best supported.
For Python based inference servers like vllm or tabbyAPI it is recommended to run them via podman or docker. This provides clean environment isolation as well as responding correctly to SIGTERM signals for proper shutdown.