Luce-Org/lucebox

★ 2,868⑂ 277

LLM speculative inference server for heterogeneous hardware & consumer GPUs

About Luce-Org/lucebox

Luce-Org/lucebox is an open-source project on GitHub, mainly written in C++. LLM speculative inference server for heterogeneous hardware & consumer GPUs It currently holds 2,868 stars and 277 forks with 0 open issues, and was last pushed on an unknown date (repository created unknown).

Project Overview

AI Homed tracks it on the Local & On-Device AI board.

GitHub Repository Details

Repository Luce-Org/lucebox · default branch - · size 0 KB · watchers 0 · source: GitHub REST API and repository README

README

https://github.com/Luce-Org/lucebox/blob/HEAD/Lucebox

https://github.com/Luce-Org/lucebox/blob/HEAD/lucebox.com https://github.com/Luce-Org/lucebox/blob/HEAD/HuggingFace https://github.com/Luce-Org/lucebox/blob/HEAD/Discord https://github.com/Luce-Org/lucebox/blob/HEAD/Blog https://github.com/Luce-Org/lucebox/blob/HEAD/Tutorials

https://github.com/Luce-Org/lucebox/blob/HEAD/Apache 2.0 https://github.com/Luce-Org/lucebox/blob/HEAD/CUDA 12+ https://github.com/Luce-Org/lucebox/blob/HEAD/HIP 7+ https://github.com/Luce-Org/lucebox/blob/HEAD/C++17

Speculative inference for heterogeneous machines and consumer GPUs.
Custom kernels, speculative prefill and decoding, tuned for each model and hardware target.

---

Inference Engine Optimizations

| Optimization | Measured setup | Result | |---|---|---:| | DFlash2 | Qwen 3.8 27B on one R9700 | 208.1 tok/s average, 227.8 tok/s peak | | DSpark | DeepSeek V4 on Strix Halo, native top-6 | 32.7 tok/s high-acceptance median; 27.9 tok/s mixed-eval average | | PFlash + KVFlash | Laguna XS 2.1 33B at 256K on RTX 3090 | 6.1× prefill, 411 s to 67.3 s | | Luce Spark | Laguna XS.2 33B on RTX 3090 | ~100 tok/s in 14.6 GiB | | KVFlash | Laguna XS 2.1 33B at 256K on RTX 3090 | 152.3 tok/s with an 8K pool | | Heterogeneous execution | DeepSeek V4 on R9700 + Strix Halo | 86 tok/s decode; 788 tok/s prefill at 2K | | Paged attention + continuous batching | Qwen 3.8 27B + DFlash2 on R9700; DeepSeek V4 Flash AR on Strix Halo | 300.9 tok/s total at 5 clients (Qwen); 48.4 tok/s output-window at 4 clients (DeepSeek) | | Megakernel | Qwen 3.5 0.8B on RTX 3090 | 413 tok/s, 1.87 tok/J |

---

Supported Models and Drafters

Model links open the exact weights used by the measured setup. Drafter links open the published quant, or the source checkpoint when conversion is required.

| Model and optimization | Phase | Speedup | |---|:---:|:---:| | Qwen 3.5 0.8B BF16 + Megakernel | Prefill + decode | 1.9× prefill; 1.55× decode | | Qwen 3.8 27B UD-IQ4_XS + DFlash2 source, converted to Q8_0, on R9700 | Decode | 6.4× vs Lucebox AR; 3.8× vs llama.cpp with the same drafter | | Laguna XS 2.1 33B Q4_K_M + PFlash/KVFlash with Qwen3 0.6B Q8_0 | Prefill | 6.1×, 411 s to 67.3 s at 256K | | Laguna XS 2.1 33B Q4_K_M + DFlash Q4 drafter | Decode | 1.7× at 256K | | Gemma 4 26B-A4B Q4_K_M + DFlash Q8_0 drafter | Decode | 1.31× | | Gemma 4 31B IT Q4_K_M + DFlash Q8_0 drafter | Decode | 3.2× | | DeepSeek V4 Flash ROCmFPX MIX Strix + DSpark Q4RMFP4 drafter | Decode | 42 tok/s at 8K and 39 tok/s on code and math with the plain launch (PR #729) | | Ling 3.0 Flash 124B-A5.1B Q4_K_M | Decode | 34.6 tok/s median AR on DGX Spark |

Tested Machines (GPU/APU)

The engine is not tied to one reference card. NVIDIA architectures are selected by CMake; HIP builds should target the device's exact gfx architecture.

| | Architecture | Hardware | Runtime | Details | |:---:|---|---|---|---| | | RDNA4 gfx1201 | Radeon AI PRO R9700 | ROCm 7.2 | Qwen 3.8 R9700 quick start | | | RDNA3.5 gfx1151 | Ryzen AI MAX+ 395 / Strix Halo | ROCm 7.2 | DeepSeek V4 Strix profile | | | RDNA3 gfx1100 | Radeon RX 7900 XT / XTX | ROCm 6+ | DeepSeek V4 dual AMD profile | | | Ampere sm_86 | RTX 3090 | CUDA 12+ | Qwen 3.8 NVLink result and Megakernel results | | | Blackwell sm_120 | RTX 5090 | CUDA 12.8+ | Qwen 3.8 single-GPU result | | | Blackwell sm_121 | DGX Spark / GB10 | CUDA 12.9 | Qwen 3.5 NVFP4 results | | | Ada sm_89 | RTX 4090 | CUDA 12+ | Linux and WSL2 community runs | | | Turing sm_75 | RTX 2080 Ti | CUDA 12.0 | DFlash results | | | Volta sm_70, Pascal sm_61 | V100, P40 | CUDA 12.0 | CUDA quick start | | Not pictured | Blackwell sm_110 | Jetson AGX Thor | CUDA 13.0 | Thor quick start |

Single-device results

| Hardware | Model | Measured result | |---|---|---| | R9700 | Qwen 3.8 27B UD-IQ4_XS + DFlash2 source | 208.1 tok/s HumanEval average; 227.8 tok/s best request | | Strix Halo | DeepSeek V4 ROCmFPX MIX Strix + DSpark Q4RMFP4 | 42 tok/s decode and 320 tok/s prefill at 8K, 36 tok/s at 123K, 39 tok/s on code and math, 25 tok/s on prose, all six routed experts, plain launch (PR #729) | | RTX 5090 | Qwen 3.8 27B | 110.6 tok/s for a 26,758-token prompt and 1,024-token continuation (PR #637) |

Heterogeneous and parallel results

| Hardware | Configuration | Measured result | |---|---|---| | 2x RTX 3090 + NVLink | Qwen 3.8 target tensor parallel + DFlash2 | 79.7 tok/s, 2.16× autoregressive decode (PR #637) | | RX 7900 XT + Strix Halo | DeepSeek V4 with all six experts + DSpark verification width 4 | 45.0 to 47.7 tok/s decode; 111.2 tok/s prefill at 132,981 tokens (PR #604) | | R9700 + Strix Halo | DeepSeek V4 across both AMD devices | 86 tok/s decode; 788 tok/s prefill at 2K |

These runs use different prompts, quantizations, and inference policies. They show which configurations work; they are not a cross-hardware ranking.

Recommended Setups

See Recommended server setups for the model and hardware matrix, including single-GPU and mixed-GPU profiles.

The DS4 guide also documents the Strix long-context sparse-verifier profile and Qwen3-0.6B PFlash integration. PFlash is lossy prompt compression; keep it off for exact-retrieval and matched true-context benchmarks.

Client Harnesses

harness/ runs Lucebox through popular coding clients and checks server compatibility.

https://github.com/Luce-Org/lucebox/blob/HEAD/Lucebox client harness experiments on RTX 3090

| Client | Launcher | |--------|----------| | Claude Code | run_claude_code.sh | | Codex | run_codex.sh | | OpenCode | run_opencode.sh | | Hermes | run_hermes.sh | | Pi | run_pi.sh | | OpenClaw | run_openclaw.sh | | Open WebUI | run_openwebui.sh |

Set the server binary and model paths, then run a launcher:

DFLASH_SERVER_BIN=server/build/dflash_server \
DFLASH_TARGET=server/models/Qwen3.8-27B-UD-IQ4_XS.gguf \
DFLASH_DRAFT=server/models/draft/qwen38-dflash2-q8_0.gguf \
MAX_CTX=32768 \
harness/clients/run_codex.sh

See the harness guide for setup, no-draft targets, and benchmarks.

Quick Start With Docker

Prebuilt images on GHCR track main. Mount the weights and serve the OpenAI-compatible API on :8000.

| GPU | Image tag | |-----|-----------| | NVIDIA (CUDA 12+) | :cuda12 | | AMD (ROCm 6+) | :rocm |

Put the target in server/models/ and its matching drafter in server/models/draft/.

https://github.com/Luce-Org/lucebox/blob/HEAD/Lucebox prebuilt Docker images for NVIDIA and AMD

Run the image for your GPU:

# NVIDIA
docker run --rm --gpus all -p 8000:8080 \
  -v "$PWD/server/models:/opt/lucebox-hub/server/models" \
  ghcr.io/luce-org/lucebox-hub:cuda12

AMD

docker run --rm --device /dev/kfd --device /dev/dri \ --group-add video --group-add render --security-opt seccomp=unconfined \ -p 8000:8080 -v "$PWD/server/models:/opt/lucebox-hub/server/models" \ ghcr.io/luce-org/lucebox-hub:rocm

Run the Server

This quick start runs the R9700 profile above. The complete flag reference is in the server guide.

# build (ROCm 7.2+, RDNA4)
git clone --recurse-submodules https://github.com/Luce-Org/lucebox.git
cd lucebox
cmake -S server -B server/build-hip -G Ninja \
  -DCMAKE_BUILD_TYPE=Release \
  -DCMAKE_HIP_COMPILER=/opt/rocm/lib/llvm/bin/clang++ \
  -DDFLASH27B_GPU_BACKEND=hip \
  -DDFLASH27B_HIP_ARCHITECTURES=gfx1201 \
  -DGGML_HIP_MMQ_MFMA=ON \
  -DGGML_HIP_NO_VMM=ON
cmake --build server/build-hip --target dflash_server -j"$(nproc)"

target and DFlash2 drafter

mkdir -p models huggingface-cli download unsloth/Qwen3.8-27B-GGUF \ Qwen3.8-27B-UD-IQ4_XS.gguf --local-dir models huggingface-cli download incoai/Qwen3.8-27B-DFlash2 --local-dir models/dflash2 python server/scripts/convert_dflash_to_gguf.py \ models/dflash2/model.safetensors models/qwen38-dflash2-f16.gguf python server/scripts/quantize_dflash_draft.py \ models/qwen38-dflash2-f16.gguf models/qwen38-dflash2-q8_0.gguf --scheme q8_0

launch the measured profile

./server/build-hip/dflash_server models/Qwen3.8-27B-UD-IQ4_XS.gguf \ --draft models/qwen38-dflash2-q8_0.gguf \ --draft-block-size 16 --max-ctx 131072 \ --cache-type-k q8_0 --cache-type-v q8_0 \ --port 8216

curl -s http://127.0.0.1:8216/v1/chat/completions \ -H 'Content-Type: application/json' \ -d '{"messages":[{"role":"user","content":"Write a Python LRU cache."}], "max_tokens":256,"temperature":0}'

To serve up to N concurrent requests, use this launch command with the same Qwen target and DFlash2 drafter. Set N to the desired concurrency (5 below). Qwen automatically sizes the shared KV pool from available GPU memory.

N=5
./server/build-hip/dflash_server models/Qwen3.8-27B-UD-IQ4_XS.gguf \
  --draft models/qwen38-dflash2-q8_0.gguf \
  --draft-block-size 16 --max-ctx 16384 \
  --paged-attention --max-concurrency "$N" \
  --cache-type-k q8_0 --cache-type-v q8_0 \
  --port 8216

See Continuous batching in Lucebox for Qwen and DeepSeek V4 Flash results, latency measurements, and launch settings.

Documentation

| Topic | Guide | |---|---| | Recommended model and hardware profiles | Recommended setups | | Runtime parameters | Server parameter reference | | OpenAI Chat Completions, Responses, and Anthropic Messages | API reference | | CUDA, HIP, and mixed-device placement | Mixed-backend guide | | DeepSeek V4 single-device and heterogeneous profiles | DeepSeek V4 guide | | Environment variables | Environment reference | | Server internals | Architecture | | Client integration and qualification | Harness guide | | Server engine components | Engine components |

Benchmarks stay with each implementation: DFlash, PFlash, Spark, KVFlash, and Megakernel.

---

Tutorials

Video tutorials for each optimization and the harness setup.

| | | | |:-:|:-:|:-:| | Luce Spark
▶ YouTube | Luce DFlash
▶ YouTube | Luce Turboquant
▶ YouTube | | OpenClaw harness setup
▶ YouTube | Luce PFlash
▶ YouTube | Luce Megakernel
▶ YouTube | | Luce KVFlash
▶ YouTube | | |

---

The Lucebox Machine

Local AI should be the default, not a privilege. Private data, no per-token bill, no vendor lock-in. Lucebox pairs the R9700 with Strix Halo and ships this open engine ready to run.

https://github.com/Luce-Org/lucebox/blob/HEAD/Lucebox local AI PC

See the hardware and current benchmarks at lucebox.com.

---

Request for Contributions

We welcome focused contributions to CUDA and HIP kernels, speculative inference, support for more consumer GPUs and APUs, performance benchmarks, and client harnesses.

---

Citation

@software{lucebox_2026,
  title  = {Lucebox: Speculative inference for heterogeneous consumer hardware},
  author = {Lucebox},
  url    = {https://github.com/Luce-Org/lucebox},
  year   = {2026}
}

---

Community

---

Apache 2.0 · Lucebox.com

GitHub Stars & Activity

2,868Stars
277Forks
0Open issues
C++Language

GitHub Popularity

GitHub stars2,868
Forks277
Open issues0
Primary languageC++
License-
Stars gained today0
Created-
Last pushed-

Trending History

Trending statusnot on today's boards

Related AI Projects

1

mozilla-ai / llamafile

C++★ 26,001⑂ 1,598
2

RunanywhereAI / runanywhere-sdks

C++★ 10,295⑂ 380
3

google-ai-edge / LiteRT-LM

C++★ 6,482⑂ 725
4

cactus-compute / cactus

C++★ 6,033⑂ 506
5

sherlockchou86 / VideoPipe

C++★ 2,955⑂ 466
6

ollama / ollama

Go★ 181,296⑂ 17,941
7

open-webui / open-webui

Python★ 152,601⑂ 22,332
8

ChatGPTNextWeb / NextChat

TypeScript★ 88,790⑂ 59,030

More AI Rankings