StayLameBro/backburner
Your iPhone helps your Mac run a 27B model: faster prompt reading and more context over a USB-C cable
About StayLameBro/backburner
StayLameBro/backburner is an open-source project on GitHub, mainly written in Python. Your iPhone helps your Mac run a 27B model: faster prompt reading and more context over a USB-C cable It currently holds 206 stars and 21 forks with 0 open issues, and was last pushed on an unknown date (repository created unknown).
Project Overview
AI Homed tracks it on the Today's Trending board, currently at rank #97 with 0 new stars today.
GitHub Repository Details
README
Backburner
Plug your iPhone into your MacBook with a 10 Gb/s USB-C cable and it helps run Qwen3.8-27B locally:
- Faster prefill (up to 64k context). For every batch of prompt tokens the Mac runs layers 1-40 and the iPhone runs 41-64
- More context. A 24 GB Mac fits 64k tokens of 8-bit context next to the model. The iPhone holds the oldest part past that
- Same answers. Greedy output is token-identical with and without the phone (256/256 tokens at 8k and 32k; 32/32 at 140k).
llama.cpp/, StayLameBro/backburner-llama.cpp)
with its own Mac kernels (SME2, Metal fusions, DFlash2 speculative decoding). Those speed things up on the Mac alone too; the
numbers below keep the two apart.
Tested on a MacBook Pro M4 Pro (24 GB) with iPhone 17 Pro Max (A19 Pro) and iPhone 16 Pro Max (A18 Pro) phones. The app also installs on iPads with an M-series chip since v0.0.2; that is untested so far, so please post your results.
Results
Measured on 2026-10-01: Qwen3.8-27B IQ4_XS, MacBook Pro M4 Pro 24 GB, iPhone 17 Pro Max over USB-C. Raw rows are in
bench/results/*.jsonl; the scripts that produced them are in bench/.
Reading (prefill): Mac alone vs Mac + iPhone, same build
A 2,000-token file or tool result read into a saved agent session (bench/turn-bench.py, two reads per depth):
| Context already in the session | Mac alone | Mac + iPhone | | |---|---|---|---| | 16k | 109 tok/s (18.8 s) | 157 tok/s (13.1 s) | +44%, 31% less waiting | | 32k | 101 tok/s (20.3 s) | 130 tok/s (15.8 s) | +29%, 22% less waiting | | 48k | 87 tok/s (23.5 s) | 113 tok/s (18.1 s) | +30%, 23% less waiting |
A new omp session (omp's system prompt, project notes and 12 tools, 26,849 tokens, read cold; bench/session-bench.py):
| | stock llama.cpp | this fork, Mac alone | this fork + iPhone | |---|---|---|---| | first answer | 245 s | 228 s | 168 s | | later turns (1.3-1.9k-token tool results) | 17.9 s | 19.2 s | 14.5 s |
After the first time, the SSD prompt cache (scripts/proxy.py) restores that 27k-token start in 0.3-5 s.
Past 64k the Mac-alone config switches to 4-bit context: 128k at 8-bit measured ~0.3 GB over the GPU's 20 GB memory limit with the draft model loaded (2026-09-23); 8-bit between 64k and 128k was not tested. With the iPhone it stays 8-bit. There the phone changes jobs: instead of running layers 41-64 it computes attention over the old keys it holds, while the Mac runs all 64 layers (see "Who does what"). Prefill past 64k is a little faster with the phone and at higher precision: 67-73 tok/s with the iPhone at 8-bit vs 59-68 tok/s Mac alone at 4-bit (64k-96k). At 128k with the iPhone: 3 of 3 planted facts recalled (positions 1.5k, 40k, 100k), phone thermal state nominal.
Writing (decode)
The phone does not change writing speed below 64k; the fork's kernels and draft model do. Past 64k the phone's GPU and Neural Engine compute attention over the old keys for every generated token:
| | context | tok/s | |---|---|---| | stock llama.cpp (Homebrew), Mac | 27-33k | 11.3 | | this fork, Mac alone | 27-33k | 25.0 | | this fork + iPhone | 27-33k | 25.1 | | this fork + iPhone, a real omp session (36 requests) | under 16k / 16-32k / 32-49k | 29.8 / 27.5 / 24.3 (medians) | | this fork + iPhone | 128k | 12.6 (greedy, 256 tokens) |
Medium thinking, omp's request fields, the server's default sampling. Mac-alone writing speed at 128k (4-bit) was not measured on the same day, so there is no head-to-head number for it here.
Context you can hold
| | 8-bit context | |---|---| | Mac alone (24 GB) | 64k measured (128k only fits with 4-bit) | | Mac + iPhone 17 Pro Max | 196k-229k by the phone's free memory (sized at startup); tested to 128k (140k at 4-bit) |
How it works
Who does what
With the defaults and one iPhone 17 Pro Max:
| | Mac GPU | Mac CPU (SME2) | iPhone GPU (matrix units) | iPhone Neural Engine | |---|---|---|---|---| | Prefill, context up to 64k | layers 1-40 | ~30% of each big matmul | layers 41-64 (split prefill) | - | | Prefill past 64k | all 64 layers | ~30% of each big matmul | attention over the old keys it holds | - (builds its pages in the background) | | Writing, up to 64k | everything, plus the draft model | oldest keys of each attention layer past 40k | - | - | | Writing past 64k | everything else | same | attention over the old keys | part of that attention |
The Mac's own Neural Engine is not used: it shares the Mac's memory bandwidth and slowed decode by 26% when running
(docs/ANE.md).
The pieces
- Split prefill (
llama.cpp/src/llama-split.cpp; the phone's tail server is inios/Backburner). The Mac runs layers
- Phone-held context (
phone-attn/, protocol inphone-attn/phone-attn.h). Past the Mac's 64k cells, the oldest KV pages
phone-attn/pa-metal.mm). While the phone
holds keys, 512-token ubatches run as two staggered halves so Mac and phone overlap (140k: 58 -> 68 tok/s prefill). Two
phones can share the old pages (docs/TWO-PHONES.md).
- The iPhone's Neural Engine for old keys (
phone-attn/pa-ane.mm). Old keys never change, so each 16,384-key page of a
scripts/serve.sh
puts the page template on the phone the first time it sees it (scripts/phone-ane.sh) and prints "ANE pages on".
- SME2 on the Mac CPU. The M4's SME units take ~30% of the rows of each big prefill matmul while the GPU does the rest
- DFlash2 speculative decoding with recurrent-state replay for the hybrid (GDN + attention) model, lossless speculative
- SSD prompt cache (
scripts/proxy.py): a known system prompt is restored from disk instead of re-read. - Memory. The model is loaded wired (
--load-mode none) so macOS can't page it out; the token-embedding table is read
- Neural Engine results, including the two uses we measured and set aside (the Mac's ANE, and the iPhone's ANE for
docs/ANE.md.
Limits
- Small reads stay on the Mac. The phone joins a read of more than ~512 tokens (three 256-token ubatches; the last one
- Past 64k the phone does one job, not two. Its half of the model (layers 41-64) can't see the old keys it holds yet, so
- Writing speed is the Mac's below 64k. The phone only joins decoding past 64k (attention over the old keys).
- A failure turns the phone off for 60 s; the batch reruns on the Mac and the server log says so.
- Saving a session while the phone holds keys (past 64k) is implemented for the current app (the rows come back from the
- One request at a time (
-np 1).
Security
Over the USB cable, the phone app answers the Mac and nothing else. Over Wi-Fi it answers only a Mac you paired over the cable, through an encrypted, authenticated tunnel (docs/WIFI.md). Nothing is advertised on the network. Before 0.0.3 the app accepted connections over Wi-Fi too: update the app, or keep the phone's Wi-Fi off while Backburner is open. Report problems privately (Security → Report a vulnerability): SECURITY.md.
Setup
You need an Apple Silicon Mac (tested: M4 Pro, 24 GB), an iPhone 15 Pro or newer (tested: 17 Pro Max, 16 Pro Max) or an M-series iPad (untested), and a 10 Gb/s USB-C cable (the cable in the iPhone box is USB 2 and too slow). The app installs with a free Apple ID through AltStore, no developer account needed (docs/INSTALL-IPHONE.md), or builds with Xcode.
One command downloads the Mac engine and the models (~24 GB) and makes the phone's half. Safe to re-run:
curl -fsSL https://raw.githubusercontent.com/StayLameBro/backburner/main/install.sh | bash
backburner phone # once per phone: plug it in, open Backburner, copies the phone's half over the cable
backburner # OpenAI-compatible server at http://127.0.0.1:8080/v1
Or step by step:
git clone --recursive https://github.com/StayLameBro/backburner && cd backburner
1. the Mac engine
cmake -S llama.cpp -B llama.cpp/build-metal -DCMAKE_BUILD_TYPE=Release
cmake --build llama.cpp/build-metal --target llama-server llama-quantize -j
2. the models: a Qwen3.8-27B IQ4_XS GGUF at ~/Models/Qwen3.8-27B-IQ4_XS.gguf, then the draft model
huggingface-cli download z-lab/Qwen3.8-27B-DFlash2 --local-dir ~/Models/qwen38-27b-dflash2
scripts/make-drafter.sh ~/Models/qwen38-27b-dflash2 ~/Models/dflash2-v2-q4km-self16.gguf
3. the iPhone app: Backburner.ipa from the latest release via AltStore (docs/INSTALL-IPHONE.md), or build it with Xcode:
DEVELOPMENT_TEAM= UDID= scripts/build-iphone.sh
pip3 install coremltools # serve.sh builds the phone's Neural Engine page model with it, once
4. the phone's half of the model (layers 41-64, ~5.1 GB; the Mac computes the logits, so no head), copied over the cable
python3 scripts/split-gguf.py ~/Models/Qwen3.8-27B-IQ4_XS.gguf ~/Models/tail-iq4xs-L40-nohead.gguf -L 40 --no-head
scripts/phone-tail.sh L40
5. after every reboot: let the GPU keep the model wired (macOS resets this limit)
sudo sysctl iogpu.wired_limit_mb=20480
6. run: OpenAI-compatible on :8080, uses the iPhone when it's plugged in with Backburner open
scripts/serve.sh
PHONE=0 scripts/serve.sh # the Mac alone
The first serve.sh with the phone pushes the Neural Engine page model to it and relaunches the app (about a minute). After
that the startup line should read split prefill on, remote KV on and ANE pages on.
scripts/serve.sh documents each setting next to the measurement that chose it.
Reproducing the numbers
bench/turn-bench.py --build # once: a saved session at 16k / 32k / 48k (Mac alone)
bench/turn-bench.py --config mac # read 2,000-token files into each saved session
bench/turn-bench.py --config phone
bench/session-bench.py --config stock|fork-mac|fork-phone # an omp-shaped session, ~5 min each
bench/long-bench.py # past 64k (long: cold reads to 128k)
Post your results
Tried it? Post your results: your Mac, your
phone(s) and the turn-bench.py output. Other Macs, other phones, iPads and two-device setups are the numbers this README
doesn't have yet.
Contributing
Pull requests are welcome: new devices, fixes, kernels, docs. CONTRIBUTING.md has what every change needs: same answers, measured speed, and the Mac alone still working.
Status
Pre-release. Next: the phone's layers seeing the keys it holds (split prefill past 64k), a second phone in the prefill chain, an App Store build, and upstreaming what makes sense to llama.cpp. Built with a lot of help from Claude Opus 5.5.
MIT license (llama.cpp keeps its own MIT license). Created by StayLameBro. Forks are welcome; if you start a separate project from it, please give it a different name and link back here.