StayLameBro/backburner

★ 206⑂ 21

Your iPhone helps your Mac run a 27B model: faster prompt reading and more context over a USB-C cable

About StayLameBro/backburner

StayLameBro/backburner is an open-source project on GitHub, mainly written in Python. Your iPhone helps your Mac run a 27B model: faster prompt reading and more context over a USB-C cable It currently holds 206 stars and 21 forks with 0 open issues, and was last pushed on an unknown date (repository created unknown).

Project Overview

AI Homed tracks it on the Today's Trending board, currently at rank #97 with 0 new stars today.

GitHub Repository Details

Repository StayLameBro/backburner · default branch - · size 0 KB · watchers 0 · source: GitHub REST API and repository README

README

Backburner

Plug your iPhone into your MacBook with a 10 Gb/s USB-C cable and it helps run Qwen3.8-27B locally:

on its GPU, pipelined. Your agent waits less every time it reads a file or a tool result: 29-44% faster prefill at 16k-48k. and computes attention over it: its GPU during prefill, its GPU and Neural Engine while writing. The server sizes the total from the phone's free memory at startup (196k-229k tokens at 8-bit on an iPhone 17 Pro Max). Tested end to end to 128k at 8-bit and 140k at 4-bit. The engine is a llama.cpp fork (llama.cpp/, StayLameBro/backburner-llama.cpp) with its own Mac kernels (SME2, Metal fusions, DFlash2 speculative decoding). Those speed things up on the Mac alone too; the numbers below keep the two apart.

Tested on a MacBook Pro M4 Pro (24 GB) with iPhone 17 Pro Max (A19 Pro) and iPhone 16 Pro Max (A18 Pro) phones. The app also installs on iPads with an M-series chip since v0.0.2; that is untested so far, so please post your results.

Seconds of waiting for each file your agent reads

Results

Measured on 2026-10-01: Qwen3.8-27B IQ4_XS, MacBook Pro M4 Pro 24 GB, iPhone 17 Pro Max over USB-C. Raw rows are in bench/results/*.jsonl; the scripts that produced them are in bench/.

Reading (prefill): Mac alone vs Mac + iPhone, same build

A 2,000-token file or tool result read into a saved agent session (bench/turn-bench.py, two reads per depth):

| Context already in the session | Mac alone | Mac + iPhone | | |---|---|---|---| | 16k | 109 tok/s (18.8 s) | 157 tok/s (13.1 s) | +44%, 31% less waiting | | 32k | 101 tok/s (20.3 s) | 130 tok/s (15.8 s) | +29%, 22% less waiting | | 48k | 87 tok/s (23.5 s) | 113 tok/s (18.1 s) | +30%, 23% less waiting |

A new omp session (omp's system prompt, project notes and 12 tools, 26,849 tokens, read cold; bench/session-bench.py):

| | stock llama.cpp | this fork, Mac alone | this fork + iPhone | |---|---|---|---| | first answer | 245 s | 228 s | 168 s | | later turns (1.3-1.9k-token tool results) | 17.9 s | 19.2 s | 14.5 s |

After the first time, the SSD prompt cache (scripts/proxy.py) restores that 27k-token start in 0.3-5 s.

Past 64k the Mac-alone config switches to 4-bit context: 128k at 8-bit measured ~0.3 GB over the GPU's 20 GB memory limit with the draft model loaded (2026-09-23); 8-bit between 64k and 128k was not tested. With the iPhone it stays 8-bit. There the phone changes jobs: instead of running layers 41-64 it computes attention over the old keys it holds, while the Mac runs all 64 layers (see "Who does what"). Prefill past 64k is a little faster with the phone and at higher precision: 67-73 tok/s with the iPhone at 8-bit vs 59-68 tok/s Mac alone at 4-bit (64k-96k). At 128k with the iPhone: 3 of 3 planted facts recalled (positions 1.5k, 40k, 100k), phone thermal state nominal.

Writing (decode)

The phone does not change writing speed below 64k; the fork's kernels and draft model do. Past 64k the phone's GPU and Neural Engine compute attention over the old keys for every generated token:

| | context | tok/s | |---|---|---| | stock llama.cpp (Homebrew), Mac | 27-33k | 11.3 | | this fork, Mac alone | 27-33k | 25.0 | | this fork + iPhone | 27-33k | 25.1 | | this fork + iPhone, a real omp session (36 requests) | under 16k / 16-32k / 32-49k | 29.8 / 27.5 / 24.3 (medians) | | this fork + iPhone | 128k | 12.6 (greedy, 256 tokens) |

Medium thinking, omp's request fields, the server's default sampling. Mac-alone writing speed at 128k (4-bit) was not measured on the same day, so there is no head-to-head number for it here.

Context you can hold

| | 8-bit context | |---|---| | Mac alone (24 GB) | 64k measured (128k only fits with 4-bit) | | Mac + iPhone 17 Pro Max | 196k-229k by the phone's free memory (sized at startup); tested to 128k (140k at 4-bit) |

How it works

Who does what

With the defaults and one iPhone 17 Pro Max:

| | Mac GPU | Mac CPU (SME2) | iPhone GPU (matrix units) | iPhone Neural Engine | |---|---|---|---|---| | Prefill, context up to 64k | layers 1-40 | ~30% of each big matmul | layers 41-64 (split prefill) | - | | Prefill past 64k | all 64 layers | ~30% of each big matmul | attention over the old keys it holds | - (builds its pages in the background) | | Writing, up to 64k | everything, plus the draft model | oldest keys of each attention layer past 40k | - | - | | Writing past 64k | everything else | same | attention over the old keys | part of that attention |

The Mac's own Neural Engine is not used: it shares the Mac's memory bandwidth and slowed decode by 26% when running (docs/ANE.md).

The pieces

1-40 of each 256-token ubatch and streams the residual to the phone, which runs layers 41-64 on its GPU while the Mac starts the next ubatch. The phone keeps a mirror of its layers' KV rows and recurrent state; only new rows cross the cable. The last ubatch of each batch runs on the Mac so outputs stay local. On the A19 Pro the phone's layers use the GPU's matrix units (Metal 4 tensor ops): 2.4x faster than the same phone with them off. (4,096 keys each) move to the phone. Each attention step sends Q to the phone and merges its partial result (O, max, sum) with the Mac's. The phone computes it with a matrix-unit kernel on its GPU (phone-attn/pa-metal.mm). While the phone holds keys, 512-token ubatches run as two staggered halves so Mac and phone overlap (140k: 58 -> 68 tok/s prefill). Two phones can share the old pages (docs/TWO-PHONES.md). layer is compiled into a Neural Engine model with the keys and values as its weights. While writing, the Neural Engine takes part of each old-key attention call and the GPU the rest: at 140k, 279 -> 176 ms per generated token. scripts/serve.sh puts the page template on the phone the first time it sees it (scripts/phone-ane.sh) and prints "ANE pages on". (Mac-only pp2048: 121.6 -> 157.1 tok/s; 51k: 79.5 -> 92.1). Past 40k keys they also take the oldest keys of each attention layer while decoding (SME co-attention). Prior art: FusionML (arXiv 2607.22785) also splits matmuls across Apple compute units. sampling for sampled requests, and block verification (Sun et al., ICLR 2025). from the mapped file, the scheduler's worst-case buffer is mapped on demand, and freed heap goes back to macOS. Server footprint after a read: 19.7 -> 18.7 GB (2026-10-01). prefill), are in docs/ANE.md.

Limits

always runs on the Mac). Most agent steps are smaller: in a real omp session 7 of 36 requests were big enough, and they carried ~83% of the tokens read. past the Mac's 64k cells the Mac runs all 64 layers and the phone only computes the old-key attention. Doing both is the next step. It helps most at 64k-100k; deeper, the phone's GPU is already busy about two thirds of each step with old keys (140k), so a second phone is the bigger win there. phone) but not yet tested end to end.

Security

Over the USB cable, the phone app answers the Mac and nothing else. Over Wi-Fi it answers only a Mac you paired over the cable, through an encrypted, authenticated tunnel (docs/WIFI.md). Nothing is advertised on the network. Before 0.0.3 the app accepted connections over Wi-Fi too: update the app, or keep the phone's Wi-Fi off while Backburner is open. Report problems privately (Security → Report a vulnerability): SECURITY.md.

Setup

You need an Apple Silicon Mac (tested: M4 Pro, 24 GB), an iPhone 15 Pro or newer (tested: 17 Pro Max, 16 Pro Max) or an M-series iPad (untested), and a 10 Gb/s USB-C cable (the cable in the iPhone box is USB 2 and too slow). The app installs with a free Apple ID through AltStore, no developer account needed (docs/INSTALL-IPHONE.md), or builds with Xcode.

One command downloads the Mac engine and the models (~24 GB) and makes the phone's half. Safe to re-run:

curl -fsSL https://raw.githubusercontent.com/StayLameBro/backburner/main/install.sh | bash
backburner phone   # once per phone: plug it in, open Backburner, copies the phone's half over the cable
backburner         # OpenAI-compatible server at http://127.0.0.1:8080/v1

Or step by step:

git clone --recursive https://github.com/StayLameBro/backburner && cd backburner

1. the Mac engine

cmake -S llama.cpp -B llama.cpp/build-metal -DCMAKE_BUILD_TYPE=Release cmake --build llama.cpp/build-metal --target llama-server llama-quantize -j

2. the models: a Qwen3.8-27B IQ4_XS GGUF at ~/Models/Qwen3.8-27B-IQ4_XS.gguf, then the draft model

huggingface-cli download z-lab/Qwen3.8-27B-DFlash2 --local-dir ~/Models/qwen38-27b-dflash2 scripts/make-drafter.sh ~/Models/qwen38-27b-dflash2 ~/Models/dflash2-v2-q4km-self16.gguf

3. the iPhone app: Backburner.ipa from the latest release via AltStore (docs/INSTALL-IPHONE.md), or build it with Xcode:

DEVELOPMENT_TEAM= UDID= scripts/build-iphone.sh

pip3 install coremltools # serve.sh builds the phone's Neural Engine page model with it, once

4. the phone's half of the model (layers 41-64, ~5.1 GB; the Mac computes the logits, so no head), copied over the cable

python3 scripts/split-gguf.py ~/Models/Qwen3.8-27B-IQ4_XS.gguf ~/Models/tail-iq4xs-L40-nohead.gguf -L 40 --no-head scripts/phone-tail.sh L40

5. after every reboot: let the GPU keep the model wired (macOS resets this limit)

sudo sysctl iogpu.wired_limit_mb=20480

6. run: OpenAI-compatible on :8080, uses the iPhone when it's plugged in with Backburner open

scripts/serve.sh PHONE=0 scripts/serve.sh # the Mac alone

The first serve.sh with the phone pushes the Neural Engine page model to it and relaunches the app (about a minute). After that the startup line should read split prefill on, remote KV on and ANE pages on.

scripts/serve.sh documents each setting next to the measurement that chose it.

Reproducing the numbers

bench/turn-bench.py --build            # once: a saved session at 16k / 32k / 48k (Mac alone)
bench/turn-bench.py --config mac       # read 2,000-token files into each saved session
bench/turn-bench.py --config phone
bench/session-bench.py --config stock|fork-mac|fork-phone   # an omp-shaped session, ~5 min each
bench/long-bench.py                    # past 64k (long: cold reads to 128k)

Post your results

Tried it? Post your results: your Mac, your phone(s) and the turn-bench.py output. Other Macs, other phones, iPads and two-device setups are the numbers this README doesn't have yet.

Contributing

Pull requests are welcome: new devices, fixes, kernels, docs. CONTRIBUTING.md has what every change needs: same answers, measured speed, and the Mac alone still working.

Status

Pre-release. Next: the phone's layers seeing the keys it holds (split prefill past 64k), a second phone in the prefill chain, an App Store build, and upstreaming what makes sense to llama.cpp. Built with a lot of help from Claude Opus 5.5.

MIT license (llama.cpp keeps its own MIT license). Created by StayLameBro. Forks are welcome; if you start a separate project from it, please give it a different name and link back here.

GitHub Stars & Activity

206Stars
21Forks
0Open issues
PythonLanguage

GitHub Popularity

GitHub stars206
Forks21
Open issues0
Primary languagePython
License-
Stars gained today0
Created-
Last pushed-

Trending History

Daily boardrank #97 · ▲ 0 stars

Related AI Projects

1

Panniantong / Agent-Reach

Python★ 89,660⑂ 7,883▲ 1,683 stars
→
2

usestrix / strix

Python★ 66,321⑂ 7,270▲ 238 stars
→
3

calesthio / OpenMontage

Python★ 62,663⑂ 8,002▲ 328 stars
→
4

agno-agi / agno

Python★ 42,533⑂ 6,082▲ 41 stars
→
5

google / skills

Python★ 20,887⑂ 1,727▲ 290 stars
→
6

arc53 / DocsGPT

Python★ 18,306⑂ 2,177▲ 6 stars
→
7

Tracer-Cloud / opensre

Python★ 11,358⑂ 1,663▲ 27 stars
→
8

jamwithai / production-agentic-rag-course

Python★ 9,367⑂ 2,067▲ 192 stars
→

More AI Rankings