ggml-org/Llama-macOS

★ 1,516⑂ 110

A cosy home for your LLMs.

About ggml-org/Llama-macOS

ggml-org/Llama-macOS is an open-source project on GitHub, mainly written in Swift. A cosy home for your LLMs. It currently holds 1,516 stars and 110 forks with 0 open issues, and was last pushed on an unknown date (repository created unknown).

Project Overview

AI Homed tracks it on the Local & On-Device AI board.

GitHub Repository Details

Repository ggml-org/Llama-macOS · default branch - · size 0 KB · watchers 0 · source: GitHub REST API and repository README

README

Llama

Llama is a macOS menu bar app for running local LLMs.

Watch a 2-minute intro 📽️


Llama


Install

brew install --cask llama-app

Or download from Releases.

How it works

When you start Llama, it runs a local server at http://localhost:9931/v1.

If you have llama.cpp installed, Llama uses it. Otherwise, it installs a prebuilt binary for your Mac. Models you've already installed via llama.cpp show up in the app automatically. You can install any GGUF model from Hugging Face, and Llama also recommends models that fit your Mac's hardware.

You can chat with any model in the built-in WebUI, connect other apps (coding agents, chat UIs, editors), or use the API directly. Models load when requested and unload when idle, so they don't take up memory when not in use.

Features

Example requests

List installed models:

curl http://localhost:9931/v1/models

Send a message to a model:

curl http://localhost:9931/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "ggml-org/gpt-oss-20b-GGUF:MXFP4",
    "messages": [{"role": "user", "content": "Hello"}]
  }'

See complete API reference in the llama.cpp server docs.

Custom model settings

The app configures each model for your Mac and writes that config to models.ini, which it regenerates on every launch. To change a setting, or add one the app doesn't set, edit ~/.config/llama/models.user.ini instead -- the app only reads that file, and merges it into the config it generates.

Use the same section header the app uses (org/repo:QUANT), and list only the keys you want to change. Keys are llama serve options without the leading dashes.

[ggml-org/gemma-4-E4B-it-GGUF:Q8_0]
temp = 0.7
ctx-size = 32768
cache-type-k = q4_1

A section for a model the app didn't find is passed through as-is, which is how you point at weights or a draft model from another repo.

[unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q2_K_XL]
model = /path/to/DeepSeek-V4-Flash-UD-Q2_K_XL.gguf
spec-type = draft-dspark
spec-draft-model = /path/to/draft.gguf

Anything you set here wins over the app's own value, including settings derived from how much memory your Mac has. If a key can't be applied, the app ignores the file and the menu says which option was at fault.

A commented template is written to this path on first launch.

Network access

By default the server is reachable only from your Mac. "Allow network access" in Settings changes that, and what it offers depends on your setup.

Tailscale — binds the server to your Tailscale address. Your other devices reach it from anywhere, Tailscale authenticates and encrypts the connection, and the server stays invisible on whatever network you're on. This is the option to use on a laptop. Shown only when Tailscale is installed and signed in.

This network — binds all interfaces (0.0.0.0), so anything on your current network can reach it. The server has no password, so use it only on a network you trust, and never together with agent mode on a network you don't own.

Experimental settings

Bind to a specific address — to pin the server to one address the app doesn't offer, set it by hand. The app leaves it alone, and falls back to localhost whenever the address isn't on any interface.

defaults write app.llama.Llama exposeToNetwork -string "192.168.1.50"

Custom server arguments — Extra CLI arguments appended to the llama serve command, for server flags the app doesn't expose (e.g. --api-key). They come after the app's own flags, so where the server honors the later occurrence they can override the app's settings. Takes effect on the next server start.

# append custom arguments to the server command
defaults write app.llama.Llama extraServerArgs -string "--api-key secret"

remove (default)

defaults delete app.llama.Llama extraServerArgs

GitHub Stars & Activity

1,516Stars
110Forks
0Open issues
SwiftLanguage

GitHub Popularity

GitHub stars1,516
Forks110
Open issues0
Primary languageSwift
License-
Stars gained today0
Created-
Last pushed-

Trending History

Trending statusnot on today's boards

Related AI Projects

1

altic-dev / FluidVoice

Swift★ 11,669⑂ 836
2

JerryZLiu / Dayflow

Swift★ 7,153⑂ 438
3

gluonfield / enchanted

Swift★ 6,002⑂ 426
4

kevinhermawan / Ollamac

Swift★ 1,911⑂ 99
5

john-rocky / CoreML-Models

Swift★ 1,870⑂ 173
6

Muesli-HQ / muesli

Swift★ 1,281⑂ 131
7

kellyvv / PhoneClaw

Swift★ 1,255⑂ 167
8

FuJacob / cotabby

Swift★ 1,030⑂ 67

More AI Rankings