google-ai-edge/LiteRT-LM
LiteRT-LM is Google's production-ready, high-performance, open-source inference framework for deploying Large Language Models on edge devices.
About google-ai-edge/LiteRT-LM
google-ai-edge/LiteRT-LM is an open-source project on GitHub, mainly written in C++. LiteRT-LM is Google's production-ready, high-performance, open-source inference framework for deploying Large Language Models on edge devices. It currently holds 6,482 stars and 725 forks with 0 open issues, and was last pushed on an unknown date (repository created unknown).
Project Overview
AI Homed tracks it on the Local & On-Device AI board.
GitHub Repository Details
README
LiteRT-LM
LiteRT-LM is Google's production-ready orchestration layer to run LLMs with LiteRT, engineered for high-performance, cross-platform execution.
🔗 Product Website | 🌐✨ Web Demo
🔥 What's New: v0.16.0
This release is a quick follow up to v0.15.0 (which brought Apple Foundation
Framework integration, CLI configuration, and JavaScript API Updates).
- 📦 C API Prebuilts: Added the first versioned C API shared library
- 🚀 Experimental YNNPACK Delegate: Added the experimental
👉 Try Gemma4-E4B with MTP on Linux, macOS, Windows or Raspberry Pi with the LiteRT-LM CLI:
litert-lm run \
--from-huggingface-repo=litert-community/gemma-4-E4B-it-litert-lm \
gemma-4-E4B-it.litertlm \
--backend=gpu \
--enable-speculative-decoding=true \
--prompt="What is the capital of France?"
🌟 Key Features
- 📱 Cross-Platform Support: Android, iOS, Web, Desktop, and IoT (e.g.
- 🚀 Hardware Acceleration: Peak performance via GPU and NPU accelerators.
- 👁️ Multi-Modality: Support for vision and audio inputs.
- 🔧 Tool Use: Function calling support for agentic workflows.
- 📚 Broad Model Support: Gemma, Llama, Phi-4, Qwen, and more.
--------------------------------------------------------------------------------
🚀 Production-Ready for Google's Products
LiteRT-LM powers on-device GenAI experiences in Chrome, Chromebook Plus, Pixel Watch, and more.
You can also try the Google AI Edge Gallery app to run models immediately on your device.
Install the app today from Google Play | Install the app today from App Store
:----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: | :--------------------------------------:
|
📰 Blogs & Announcements
Link | Description :-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | :---------- Bringing Gemma 4 12B to your Laptop: Unlocking Local, Agentic Workflows with Google AI Edge | Bring agentic, multimodal AI capabilities to everyday laptops, enabling local data processing and visual insight generation. Blazing-fast on-device GenAI with LiteRT-LM | Unlock Gemma 4's full potential with blazing speed and incredible efficiency using newly added Swift, JavaScript, and Flutter APIs. Accelerating Gemma 4: faster inference with multi-token prediction drafters | An overview of how Multi-Token Prediction (MTP) drafters are making Gemma 4 models up to 3x faster at inference. Bring state-of-the-art agentic skills to the edge with Gemma 4 | Deploy Gemma 4 in-app and across a broader range of devices with stellar performance and broad reach using LiteRT-LM. On-device GenAI in Chrome, Chromebook Plus and Pixel Watch | Deploy language models on wearables and browser-based platforms using LiteRT-LM at scale. On-device Function Calling in Google AI Edge Gallery | Explore how to fine-tune FunctionGemma and enable function calling capabilities powered by LiteRT-LM Tool Use APIs. Google AI Edge small language models, multimodality, and function calling | Latest insights on RAG, multimodality, and function calling for edge language models.
--------------------------------------------------------------------------------
🏃 Quick Start
🔗 Key Links
including performance benchmarks, model support, and more.- 👉 LiteRT-LM CLI Guide including
⚡ Quick Try (No Code)
Try LiteRT-LM immediately from your terminal without writing a single line of
code using uv:
uv tool install litert-lm
litert-lm run \
--from-huggingface-repo=google/gemma-3n-E2B-it-litert-lm \
gemma-3n-E2B-it-int4 \
--prompt="What is the capital of France?"
--------------------------------------------------------------------------------
📚 Supported Language APIs
Ready to get started? Explore our language-specific guides and setup instructions.
Language | Status | Best For... | Documentation :------------------- | :-------------- | :---------------------- | :------------ Python | ✅ Stable | Prototyping & Scripting | Python Guide Kotlin | ✅ Stable | Android apps & JVM | Kotlin Guide Swift | 🚀 Early Preview | Native iOS & macOS | Swift Guide JavaScript (web) | 🚀 Early Preview | Browser environments | JavaScript Guide Flutter | 🚀 Community | Cross-platform mobile | Flutter Guide C++ | ✅ Stable | High-performance native | C++ Guide
🏗️ Build From Source
This guide shows how you can compile
LiteRT-LM from source. If you want to build the program from source, you should
checkout the stable
tag.
--------------------------------------------------------------------------------