mybigday/llama.rn

★ 1,041⑂ 122

React Native binding of llama.cpp

About mybigday/llama.rn

mybigday/llama.rn is an open-source project on GitHub, mainly written in C. React Native binding of llama.cpp It currently holds 1,041 stars and 122 forks with 0 open issues, and was last pushed on an unknown date (repository created unknown).

Project Overview

AI Homed tracks it on the Local & On-Device AI board.

GitHub Repository Details

Repository mybigday/llama.rn · default branch - · size 0 KB · watchers 0 · source: GitHub REST API and repository README

README

llama.rn

Actions Status License: MIT npm

React Native binding of llama.cpp - LLM inference in C/C++

Key Features:

[!IMPORTANT]
Starting with v0.10, llama.rn requires React Native's New Architecture.
> For Old Architecture support or documentation for v0.9.x, please refer to the v0.9 branch.

Installation

npm install llama.rn

llama.rn downloads the pre-built ios/rnllama.xcframework and android/src/main/jniLibs from the matching GitHub release during postinstall. Existing downloads are reused, and each archive is verified with SHA-256 before extraction.

Bun

Bun does not run dependency lifecycle scripts unless the package is trusted. If you install llama.rn with Bun and want the native downloads to run automatically, add it to trustedDependencies and re-run bun install:

{
  "trustedDependencies": ["llama.rn"]
}

If you prefer not to trust dependency lifecycle scripts, run the downloader manually before npx pod-install or your Android build:

node ./node_modules/llama.rn/install/download-native-artifacts.js

iOS

Please re-run npx pod-install again.

By default, llama.rn will use pre-built rnllama.xcframework for iOS. If you want to build from source, please set RNLLAMA_BUILD_FROM_SOURCE to 1 in your Podfile.

Android

Add proguard rule if it's enabled in project (android/app/proguard-rules.pro):

# llama.rn
-keep class com.rnllama.** { *; }

By default, llama.rn will use pre-built libraries for Android. If you want to build from source, please set rnllamaBuildFromSource to true in android/gradle.properties.

OpenCL (GPU acceleration)
Hexagon (NPU acceleration) (Experimental)

Expo

For use with the Expo framework and CNG builds, you will need expo-build-properties to utilize iOS and OpenCL features. Simply add the following to your app.json/app.config.js file:

module.exports = {
  expo: {
    // ...
    plugins: [
      // ...
      [
        'llama.rn',
        // optional fields, below are the default values
        {
          enableEntitlements: true,
          entitlementsProfile: 'production',
          forceCxx20: true,
          enableOpenCL: true,
        },
      ],
    ],
  },
}

Obtain the model

You can search HuggingFace for available models (Keyword: GGUF).

For get a GGUF model or quantize manually, see quantize documentation in llama.cpp.

Usage

💡 You can find complete examples in the example project.

Load model info only:

import { loadLlamaModelInfo } from 'llama.rn'

const modelPath = 'file://' console.log('Model Info:', await loadLlamaModelInfo(modelPath))

Initialize a Llama context & do completion:

import { initLlama } from 'llama.rn'

// Initial a Llama context with the model (may take a while) const context = await initLlama({ model: modelPath, use_mlock: true, n_ctx: 2048, n_gpu_layers: 99, // number of layers to store in GPU memory (Metal/OpenCL) // embedding: true, // use embedding })

const stopWords = ['', '<|end|>', '<|eot_id|>', '<|end_of_text|>', '<|im_end|>', '<|EOT|>', '<|END_OF_TURN_TOKEN|>', '<|end_of_turn|>', '<|endoftext|>']

// Do chat completion const msgResult = await context.completion( { messages: [ { role: 'system', content: 'This is a conversation between user and assistant, a friendly chatbot.', }, { role: 'user', content: 'Hello!', }, ], n_predict: 100, stop: stopWords, // ...other params }, (data) => { // This is a partial completion callback const { token } = data }, ) console.log('Result:', msgResult.text) console.log('Timings:', msgResult.timings)

// Or do text completion const textResult = await context.completion( { prompt: 'This is a conversation between user and llama, a friendly chatbot. respond in simple markdown.\n\nUser: Hello!\nLlama:', n_predict: 100, stop: [...stopWords, 'Llama:', 'User:'], // ...other params }, (data) => { // This is a partial completion callback const { token } = data }, ) console.log('Result:', textResult.text) console.log('Timings:', textResult.timings)

The binding's design inspired by server.cpp example in llama.cpp:

Please visit the Documentation for more details.

You can also visit the example to see how to use it.

MTP Speculative Decoding

MTP speculative decoding can be enabled for GGUF models that contain MTP/NextN layers:

const context = await initLlama({
  model: modelPath,
  n_ctx: 4096,
  n_batch: 1024,
  n_ubatch: 512,
  n_gpu_layers: 99,
  flash_attn_type: 'auto',
  cache_type_k: 'q8_0',
  cache_type_v: 'q8_0',
  speculative: {
    type: 'draft-mtp',
    n_max: 3,
  },
})

const result = await context.completion({ messages: [ { role: 'user', content: 'Write a concise TypeScript function that groups an array of objects by a key.', }, ], chat_template_kwargs: { preserve_thinking: true, }, n_predict: 128, temperature: 0.6, top_k: 20, top_p: 0.95, speculative: { type: 'draft-mtp', n_max: 3, }, })

console.log(result.text) console.log(result.draft_tokens, result.draft_tokens_accepted)

Use speculative: false on a completion call to disable MTP for that request. For recurrent or hybrid models, enable MTP at initLlama time with a positive spec_draft_n_max or speculative.draft.n_max so llama.cpp can allocate rollback state. Current MTP support is text-only, including queued parallel completions.

Multimodal (Vision & Audio)

llama.rn supports multimodal capabilities including vision (images) and audio processing. This allows you to interact with models that can understand both text and media content.

Supported Media Formats

Images (Vision):

Audio:

Setup

First, you need a multimodal model and its corresponding multimodal projector (mmproj) file, see how to obtain mmproj for more details.

Initialize Multimodal Support

import { initLlama } from 'llama.rn'

// First initialize the model context const context = await initLlama({ model: 'path/to/your/multimodal-model.gguf', n_ctx: 4096, n_gpu_layers: 99, // Recommended for multimodal models // Important: Disable context shifting for multimodal ctx_shift: false, })

// Initialize multimodal support with mmproj file const success = await context.initMultimodal({ path: 'path/to/your/mmproj-model.gguf', use_gpu: true, // Recommended for better performance })

// Check if multimodal is enabled console.log('Multimodal enabled:', await context.isMultimodalEnabled())

if (success) { console.log('Multimodal support initialized!')

// Check what modalities are supported const support = await context.getMultimodalSupport() console.log('Vision support:', support.vision) console.log('Audio support:', support.audio) } else { console.log('Failed to initialize multimodal support') }

// Release multimodal context await context.releaseMultimodal()

Usage Examples

Vision (Image Processing)

const result = await context.completion({
  messages: [
    {
      role: 'user',
      content: [
        {
          type: 'text',
          text: 'What do you see in this image?',
        },
        {
          type: 'image_url',
          image_url: {
            url: 'file:///path/to/image.jpg',
            // or base64: 'data:image/jpeg;base64,/9j/4AAQSkZJRgABAQEAYABgAAD...'
          },
        },
      ],
    },
  ],
  n_predict: 100,
  temperature: 0.1,
})

console.log('AI Response:', result.text)

Audio Processing

// Method 1: Using structured message content (Recommended)
const result = await context.completion({
  messages: [
    {
      role: 'user',
      content: [
        {
          type: 'text',
          text: 'Transcribe or describe this audio:',
        },
        {
          type: 'input_audio',
          input_audio: {
            data: 'data:audio/wav;base64,UklGRiQAAABXQVZFZm10...',
            // or url: 'file:///path/to/audio.wav',
            format: 'wav', // or 'mp3'
          },
        },
      ],
    },
  ],
  n_predict: 200,
})

console.log('Transcription:', result.text)

Tokenization with Media

// Tokenize text with media
const tokenizeResult = await context.tokenize(
  'Describe this image: <__media__>',
  {
    media_paths: ['file:///path/to/image.jpg']
  }
)

console.log('Tokens:', tokenizeResult.tokens) console.log('Has media:', tokenizeResult.has_media) console.log('Media positions:', tokenizeResult.chunk_pos_media)

Notes

Tool Calling

llama.rn has universal tool call support by using minja (as Jinja template parser) and chat.cpp in llama.cpp.

Example:

import { initLlama } from 'llama.rn'

const context = await initLlama({ // ...params })

const { text, tool_calls } = await context.completion({ // ...params tool_choice: 'auto', tools: [ { type: 'function', function: { name: 'ipython', description: 'Runs code in an ipython interpreter and returns the result of the execution after 60 seconds.', parameters: { type: 'object', properties: { code: { type: 'string', description: 'The code to run in the ipython interpreter.', }, }, required: ['code'], }, }, }, ], messages: [ { role: 'system', content: 'You are a helpful assistant that can answer questions and help with tasks.', }, { role: 'user', content: 'Test', }, ], }) console.log('Result:', text) // If tool_calls is not empty, it means the model has called the tool if (tool_calls) console.log('Tool Calls:', tool_calls)

You can check chat.cpp for models has native tool calling support, or it will fallback to GENERIC type tool call.

The generic tool call will be always JSON object as output, the output will be like {"response": "..."} when it not decided to use tool call.

Grammar Sampling

GBNF (GGML BNF) is a format for defining formal grammars to constrain model outputs in llama.cpp. For example, you can use it to force the model to generate valid JSON, or speak only in emojis.

You can see GBNF Guide for more details.

llama.rn provided a built-in function to convert JSON Schema to GBNF:

Example gbnf grammar:

root   ::= object
value  ::= object | array | string | number | ("true" | "false" | "null") ws

object ::= "{" ws ( string ":" ws value ("," ws string ":" ws value)* )? "}" ws

array ::= "[" ws ( value ("," ws value)* )? "]" ws

string ::= "\"" ( [^"\\\x7F\x00-\x1F] | "\\" (["\\bfnrt] | "u" [0-9a-fA-F]{4}) # escapes )* "\"" ws

number ::= ("-"? ([0-9] | [1-9] [0-9]{0,15})) ("." [0-9]+)? ([eE] [-+]? [0-9] [1-9]{0,15})? ws

Optional space: by convention, applied in this grammar after literal chars when allowed

ws ::= | " " | "\n" [ \t]{0,20}
import { initLlama } from 'llama.rn'

const gbnf = '...'

const context = await initLlama({ // ...params grammar: gbnf, })

const { text } = await context.completion({ // ...params messages: [ { role: 'system', content: 'You are a helpful assistant that can answer questions and help with tasks.', }, { role: 'user', content: 'Test', }, ], }) console.log('Result:', text)

Also, this is how json_schema works in response_format during completion, it converts the json_schema to gbnf grammar.

Parallel Decoding

llama.rn supports slot-based parallel request processing for concurrent completion requests, enabling multiple prompts to be processed simultaneously with automatic queue management. It is similar to the llama.cpp server.

Usage

import { initLlama } from 'llama.rn'

const context = await initLlama({ model: modelPath, n_ctx: 8192, n_gpu_layers: 99, n_parallel: 4, // Max number of parallel slots supported })

// Enable parallel mode with 4 slots await context.parallel.enable({ n_parallel: 4, // new_n_ctx (2048) = n_ctx / n_parallel n_batch: 512, })

// Queue multiple completion requests const request1 = await context.parallel.completion( { messages: [{ role: 'user', content: 'What is AI?' }], n_predict: 100, }, (requestId, data) => { console.log(Request ${requestId}:, data.token) } )

const request2 = await context.parallel.completion( { messages: [{ role: 'user', content: 'Explain quantum computing' }], n_predict: 100, }, (requestId, data) => { console.log(Request ${requestId}:, data.token) } )

// Cancel a request if needed await request1.stop()

// Wait for completion const result = await request2.promise console.log('Result:', result.text)

// Disable parallel mode when done await context.parallel.disable()

API

context.parallel.enable(config?):

context.parallel.disable(): context.parallel.configure(config): context.parallel.completion(params, onToken?): context.parallel.embedding(text, params?): context.parallel.rerank(query, documents, params?):

Notes

Session (State)

The session file is a binary file that contains the state of the context, it can saves time of prompt processing.

const context = await initLlama({ ...params })

// After prompt processing or completion ...

// Save the session await context.saveSession('')

// Load the session await context.loadSession('')

Notes

Embedding

The embedding API is used to get the embedding of a text.

const context = await initLlama({
  ...params,
  embedding: true,
})

const { embedding } = await context.embedding('Hello, world!')

Rerank

The rerank API is used to rank documents based on their relevance to a query. This is particularly useful for improving search results and implementing retrieval-augmented generation (RAG) systems.

const context = await initLlama({
  ...params,
  embedding: true, // Required for reranking
  pooling_type: 'rank', // Use rank pooling for rerank models
})

// Rerank documents based on relevance to query const results = await context.rerank( 'What is artificial intelligence?', // query [ 'AI is a branch of computer science.', 'The weather is nice today.', 'Machine learning is a subset of AI.', 'I like pizza.', ], // documents to rank { normalize: 1, // Optional: normalize scores (default: from model config) } )

// Results are automatically sorted by score (highest first) results.forEach((result, index) => { console.log(Rank ${index + 1}:, { score: result.score, document: result.document, originalIndex: result.index, }) })

Notes

Recommended Models

Text-to-Speech (TTS)

[!WARNING]
Experimental. The TTS / codec integration is not production-ready. The API surface may change without a major version bump, and several model families do not yet produce correct speech on all backends — see Tested models before picking one.

llama.rn runs on-device neural text-to-speech through codec.cpp as the audio-codec / vocoder backend. Load a TTS backbone as usual, attach its codec (vocoder) GGUF, then drive the synthesis flow the library detects for the model.

Supported models

The model family is auto-detected natively; getTTSCapabilities() reports it and getFormattedAudioCompletion() returns the right synthesis flow.

| Family | Variants | Flow | |---|---|---| | OuteTTS | v0.1 · v0.2 · v0.3 · v1.0 | tokens | | Soprano | 1.1 (80M) | tokens | | NeuTTS | Nano · Air (needs a phonemizer) | tokens | | CSM | 1B | tokens | | Qwen3-TTS | 0.6B | tokens | | MOSS-TTSD | v0.5 | tokens | | MOSS-TTS-Realtime | streaming interleave | tokens | | Chatterbox (T3) | English · 23-language multilingual | tokens | | BlueMagpie-TTS | Barbet backbone + AudioVAE | continuous_embd |

"Supported" means the family is detected and wired end-to-end. It does not mean every family is verified on every backend — see Tested models.

There are two synthesis flows:

Codec-LM AR and continuous-latent models must be initialized with embedding: true.

Usage

```ts import { initLlama } from 'llama.rn'

const ctx = await initLlama({ model: backbonePath,

GitHub Stars & Activity

1,041Stars
122Forks
0Open issues
CLanguage

GitHub Popularity

GitHub stars1,041
Forks122
Open issues0
Primary languageC
License-
Stars gained today0
Created-
Last pushed-

Trending History

Trending statusnot on today's boards

Related AI Projects

1

QwenAudio / Fun-ASR

C★ 1,549⑂ 152
2

cosmo-wander-ai / cosmo-edge

C★ 1,218⑂ 222
3

ollama / ollama

Go★ 181,337⑂ 17,947
4

open-webui / open-webui

Python★ 152,659⑂ 22,336
5

HKUDS / nanobot

Python★ 48,435⑂ 8,555
6

chatchat-space / Langchain-Chatchat

Python★ 38,654⑂ 6,265
7

Zackriya-Solutions / meetily

Rust★ 30,978⑂ 3,365
8

mozilla-ai / llamafile

C++★ 26,010⑂ 1,600

More AI Rankings