spiceai/spiceai
Add a real-time analytics node to your operational database. Spice is a portable, accelerated SQL query, search, and LLM-inference engine in Rust for data-grounded AI apps and agents.
About spiceai/spiceai
spiceai/spiceai is an open-source project on GitHub, mainly written in Rust. Add a real-time analytics node to your operational database. Spice is a portable, accelerated SQL query, search It currently holds 3,086 stars and 0 forks with 0 open issues, and was last pushed on an unknown date (repository created unknown).
Project Overview
AI Homed tracks it on the AI Models & LLM Tools board.
GitHub Repository Details
README
📄 Docs | ⚡️ Quickstart | 🧑🍳 Cookbook | 🤖 AI Skills | 📰 Blog
Spice is a portable, accelerated SQL query, search, and LLM-inference engine, written in Rust, for data-grounded AI apps and agents. Run it as a sidecar next to your application — or scale to a multi-node distributed cluster — to get millisecond data and AI on localhost, backed by your existing data sources.
🎯 Goal: Build data-grounded AI apps and agents in minutes, not months. No pipelines. No glue. Just SQL, search, and inference — federated across your data, accelerated locally, served on localhost.
🆕 New in Spice 2.0 — add a real-time analytics node to your operational database. Point Spice at PostgreSQL, MySQL, or MongoDB and it maintains a sandboxed, analytics-ready replica with high-throughput CDC replication — sub-second queries, ~2-second freshness, and zero analytical load on production. No ETL, no Debezium, no Kafka required. Read the Spice 2.0 launch →
Why Spice?
- ⚡ Real-time analytics node for your operational database — Add a sandboxed analytics replica to PostgreSQL, MySQL, and MongoDB via native CDC (WAL, binlog, change streams) plus DynamoDB Streams — ~2-second freshness, zero load on production, no ETL, no Debezium or Kafka required.
- 🚀 Localhost latency at any scale — Millisecond queries against a sandboxed working set on each pod, transparently delegated to a distributed cluster for the long tail.
- 🦀 Built in Rust on industry-leading open foundations: Apache DataFusion, Apache Ballista, Apache Arrow, Apache Iceberg, Vortex, DuckDB, and SQLite.
- ⚡ Distributed query without the operational tax — Apache Ballista with multi-active schedulers coordinated through object storage. 2.9x faster than single-node DataFusion on TPC-H SF100, 8x less RAM than Spark.
- 💎 Spice Cayenne data accelerator on Vortex (GA) — 1.5x faster than DuckDB with 3x less memory on TPC-H SF100, 26x faster than Spice 1.x on TPC-DS SF100, 100x faster random access vs. Parquet.
- 🔍 Petabyte-scale hybrid search — Native Amazon S3 Vectors, Tantivy BM25, DuckDB HNSW, and Elasticsearch kNN, with reciprocal rank fusion (RRF) and reranker UDTFs — all in a single SQL query.
- 🤖 AI-native runtime — OpenAI-compatible APIs, MCP server + gateway, LLM memory, NSQL text-to-SQL, multi-vector ColBERT-style embeddings, provider-aware prompt caching.
- 🔗 30+ data connectors with advanced query push-down — federate Postgres, MySQL, Snowflake, Databricks, Iceberg, Delta Lake, S3, Spark, MSSQL, DynamoDB, MongoDB, GitHub, SharePoint, Kafka, and more.
- 📝 Open table formats, first-class — Query, accelerate, and write to Apache Iceberg with ACID guarantees via standard SQL
INSERT INTO. No Spark required. - 🛡️ Enterprise-ready — HashiCorp Vault and Azure Key Vault secret stores, mTLS, read-only API keys, observability via OpenTelemetry, and an extensibility model used in production at companies like Twilio and Barracuda.
What you get
Spice provides five APIs and interfaces in a lightweight, portable runtime (single binary or container):
1. SQL Query & Search: HTTP, Arrow Flight, Arrow Flight SQL, ODBC, JDBC, and ADBC APIs; vector_search, text_search, rrf, and rerank UDTFs.
2. Text-to-SQL (NSQL): Natural-language SQL generation grounded in your federated schema with built-in sampling tools — usable from the HTTP API, the SQL REPL, or directly inside agent tool calls.
3. OpenAI-Compatible APIs: Hosted LLM gateway (OpenAI, Anthropic, xAI, Bedrock) and local model serving (CUDA/Metal accelerated). Includes the OpenAI Responses API, web search, and tool calls.
4. Iceberg Catalog REST APIs: A unified Iceberg REST Catalog API for query and write.
5. MCP HTTP+SSE APIs: Model Context Protocol server and gateway with Streamable HTTP transport. Dual-era: serves 2026-07-28 (server/discover, sessionless) and still answers legacy initialize.
🎥 Watch & Learn
- 🎓 CMU Databases: Accelerating Data and AI with Spice.ai Open-Source — Luke Kim at the Carnegie Mellon Database Group
- ☁️ AWS re:Invent 2025 (STG364): How Spice AI operationalizes data lakes for AI using Amazon S3
- 🔍 How to search with Amazon S3 Vectors
- 💎 Introducing the Spice Cayenne Data Accelerator
- 🧊 Writing to Apache Iceberg Tables with Spice.ai
- 🔌 Using Spice as an MCP Server and Gateway
- 🛠️ How to Query Data using Spice, OpenAI, and MCP
What's New
Analytics node for operational databases — real-time CDC, no ETL
Add a sandboxed, analytics-ready replica alongside PostgreSQL, MySQL, and MongoDB in minutes — ~2-second end-to-end freshness, zero analytical load on production, and no ETL. Spice replicates committed inserts, updates, and deletes directly from the native change log at up to ~170x the ingest throughput of Spice 1.x, so production never runs a single analytical query. It's incrementally adoptable: start with 1 table and be querying operational data in minutes, then join across replicated sources in a single SQL query. In the CH-BenCHmark HTAP benchmark, 1 Spice node served 1,046 analytical queries/hour at SF1000 (1,000 warehouses, 300M+ rows) while the source sustained a 266,000+ tpmC live transactional load. Read the Spice 2.0 launch →
- PostgreSQL (WAL), MySQL (binlog), and MongoDB (change streams) — native replication with auto-managed replication state (slots, binlog positions, resume tokens) and bootstrapped initial snapshots. No Debezium or Kafka required.
- DynamoDB Streams — two-tier acceleration that fans out from a central Spice layer to thousands of edge sidecars with sub-second propagation. Used in production for global control-plane sync. Read the pattern →
- Debezium — Kafka consumer (
from: debezium:…) or push ingest without Kafka (from: cdc:…+POST /v1/datasets/{name}/cdc, JSON/Avro).
Cluster-Sidecar Architecture: localhost latency, cluster scale
Each application gets a complete data plane on localhost. A lightweight Spice sidecar runs in the application pod, serves SQL/search/LLM-inference from a scoped working set, and transparently delegates the long tail to a central Spice cluster (Ballista distributed query, Cayenne acceleration, hybrid search indexing) over Arrow Flight. Three latency tiers: results cache (microseconds) → local working set (single-digit milliseconds) → cluster delegation. The application never holds credentials to Postgres, S3, Snowflake, or Iceberg — only a token to its sidecar. Read the architecture deep dive →
Apache Ballista distributed query
Spice extends Apache Ballista with multi-active scheduler HA coordinated through object storage (no etcd, ZooKeeper, or Redis required), bidirectional gRPC control streams, mandatory mTLS, multiple shuffle backends (local, in-memory, S3/Azure/GCS), Vortex-encoded shuffle data, and distributed embeddings inside SQL. TPC-H SF100: 2.9x faster on 3 executors than 1 node. 8x less RAM than Apache Spark with 2–8x better query performance — now generally available. Read the engineering deep dive →
Spice Cayenne — next-gen data acceleration on Vortex
Cayenne pairs the Vortex columnar format with SQLite metadata to deliver multi-file acceleration without DuckDB's single-file ceiling or memory overhead. Now GA with atomic WAL-staged writes, high-throughput CDC ingestion, MERGE INTO, and SQL-defined partitioning. TPC-H SF100: 1.5x faster than DuckDB with 3x less memory. TPC-DS SF100: 26x faster than Spice 1.x. ClickBench: 14% faster, 3.4x less memory. Vortex itself is 100x faster on random access, 10–20x faster on full scans, and 5x faster writes than Parquet — compute kernels run directly on encoded data, skipping decompression entirely for many operations. Read the Vortex deep dive →
Apache Iceberg: query, accelerate, and write
Connect to any Iceberg catalog (REST, AWS Glue, Hadoop), query tables with full SQL semantics, selectively accelerate hot datasets for sub-10ms reads (down from 500ms–5s on S3), and write back with ACID guarantees via Iceberg's optimistic concurrency protocol — using standard SQL INSERT INTO. No Spark required. Read the Iceberg deep dive →
Petabyte-scale hybrid search
Native Amazon S3 Vectors (Day 1 launch partner) for billions of vectors at up to 90% lower cost than traditional vector DBs. Plus DuckDB HNSW and Elasticsearch kNN as .vectors.engine backends. Spice manages the full lifecycle — ingestion → embedding (AWS Bedrock, HuggingFace, OpenAI, Model2Vec for 500x faster static embeddings, multi-vector ColBERT-style late interaction with MaxSim) → indexing → query. SQL-integrated via vector_search, text_search, rrf (reciprocal rank fusion), and rerank UDTFs.
SELECT * FROM rerank(
rrf(
vector_search('docs', 'how does Spice accelerate Iceberg?'),
text_search('docs', 'how does Spice accelerate Iceberg?')
),
document => content
) LIMIT 10;
Multi-tenancy for AI agents — without per-tenant pipelines
Spin up one Spice runtime per tenant or agent — each with its own sandboxed datasets, accelerators, secrets, and policies. Or share a runtime with config-level tenant isolation. Or do both with a hybrid model. The lightweight ~140MB runtime makes "one Spicepod per tenant" actually viable — even at thousands of tenants. Read the patterns →
Spice Skills for AI coding agents
Drop-in skills for Claude Code, Cursor, and any agent that supports the open Agent Skills format. Skills auto-activate to set up datasets, connect data sources, configure acceleration, run federated queries, and wire models — without you re-explaining Spice's configuration model.
In Claude Code:
/plugin marketplace add spiceai/skills
github.com/spiceai/skills | Read the announcement →
Acceleration Snapshots
Bootstrap accelerated datasets from S3 in seconds, not minutes. Cold-start ephemeral pods with pre-built Vortex/DuckDB/SQLite files. Recover from federated source outages by serving from the last known good snapshot. Critical for sidecar deployments and serverless environments.
Enterprise hardening (latest)
- HashiCorp Vault and Azure Key Vault secret stores
- Read-only API keys enforced on Flight DoGet and async query paths
- Provider-aware LLM prompt caching for cost reduction
- mTLS for all internal cluster communication; OpenTelemetry metric export with delta temporality
- Streamable HTTP MCP transport (
2026-07-28+ legacyinitialize), MCP gateway, MCP server - 30+ data connectors with shared HTTP rate control, dynamic headers, schema decomposition
How is Spice different?
1. Cluster-sidecar architecture — Each application gets its own Spice sidecar serving SQL, search, and LLM inference on localhost, transparently delegating the long tail to a central Spice cluster (Ballista distributed query, Cayenne acceleration, hybrid search indexing) over Arrow Flight. You get three latency tiers in one engine: results cache (microseconds) → local working set (single-digit milliseconds) → cluster delegation (distributed). No other open-source runtime gives you all three behind one connection. Read the architecture →
2. Structural data sandboxing — Datasets a sidecar doesn't declare in its spicepod.yaml are physically absent from the catalog, not filtered at query time. The application never holds credentials to Postgres, S3, Snowflake, or Iceberg — only a token to its sidecar. A compromised pod gets a loopback scoped to that tenant's working set, not database credentials.
3. Ingest once, serve everywhere — The cluster ingests each source dataset once and produces one authoritative materialization that every sidecar pulls. Source systems see one stable connection pool, not one per pod. Pull-based refresh + acceleration snapshots in S3 mean cold starts in seconds and graceful degradation when the cluster is unreachable.
4. AI-Native Runtime — Data query and AI inference live in one engine, so retrieval, ranking, and generation happen in one query plan, in one process — vector_search, text_search, rrf, rerank, NSQL, and tool calls are all SQL primitives.
5. Dual-engine acceleration — Per-dataset choice of OLAP (Cayenne/Vortex, Arrow, DuckDB) and OLTP (SQLite, PostgreSQL) engines, so you can match workload to engine instead of forcing everything into one shape.
6. Edge to cloud, single binary — Runs on a laptop, as a Kubernetes sidecar, as a microservice, or as a multi-node Ballista cluster across edge, on-prem, and public clouds. Self-hosted OSS, Spice Cloud (managed cluster), and Spice.ai Enterprise (on-prem full stack) all use identical spicepod.yaml manifests — no app changes to migrate.
If you build with DataFusion, DuckDB, Vortex, Iceberg, or Ballista, Spice gives you a flexible, production-ready engine you can just use — instead of stitching them together yourself.
Example Use-Cases
Real-time Analytics on Operational Data (no ETL)
- Analytics node for PostgreSQL, MySQL, and MongoDB: Point Spice at a live operational database and it maintains a continuously updated, sandboxed analytics replica via native CDC — sub-second queries, ~2-second freshness, and zero analytical queries against production. Start with one table, then join across replicated sources in one SQL query. CDC Docs
- HTAP at scale: Sustain analytics and transactions on the same data — 1,046 analytical QPH at SF1000 under a 266,000+ tpmC transactional load in CH-BenCHmark, all served from the replica. Spice 2.0 launch →
- Bring your own BI tools: Query the replica from Power BI, Tableau, Looker, and Apache Superset over Arrow Flight SQL, ODBC, and JDBC — or from Python and the Go, Rust, Java, and JavaScript SDKs.
Data-grounded Agentic AI Applications
- OpenAI-compatible AI Gateway: Hosted (OpenAI, Anthropic, xAI, Bedrock) or local models (Llama, NVIDIA NIM) with Responses API, streaming tool calls, web search, and provider-aware prompt caching. AI Gateway Recipe
- Federated Data Access: SQL and NSQL (text-to-SQL) across 30+ sources with advanced push-down, scaling to multi-node Ballista. Federated SQL Query Recipe
- Search and RAG: Petabyte-scale vector search via Amazon S3 Vectors, BM25 full-text via Tantivy, ColBERT-style multi-vector embeddings with MaxSim, hybrid search with RRF, rerank UDTF. Amazon S3 Vectors Recipe
- LLM Memory and Observability: Persistent agent memory + deep visibility into data flows, model performance, and traces. LLM Memory Recipe | Observability Docs
Database CDN and Query Mesh
- Co-located acceleration: Materialize working sets as Cayenne (Vortex), Arrow, SQLite, DuckDB, or Postgres alongside your app for sub-second query. Bootstrap from S3 snapshots. DuckDB Accelerator Recipe
- Resiliency: Maintain availability with local replicas of critical datasets; recover from source outages from snapshots. Local Dataset Replication Recipe
- Responsive dashboards: Sub-second BI with configurable refresh and CDC. Sales BI Demo
- Legacy modernization: One endpoint that federates legacy systems with modern infrastructure. Federation Recipe
Multi-Tenant AI Agents
- One Spicepod per tenant or per agent — sandboxed datasets, sources, secrets, and policies per agent. The runtime is light enough to make this actually viable. Patterns →
Retrieval-Augmented Generation (RAG)
- Hybrid search in SQL: Combine vector + BM25 with RRF and rerank, in one query plan, against your own data — accelerated.
- Semantic Knowledge Layer: Define a semantic context model so agents understand the shape and meaning of your data. Semantic Model Docs
- Text-to-SQL: Built-in NSQL with sampling tools for grounded SQL generation. Text-to-SQL Recipe
FAQ
- Is Spice a cache? Not exactly — think of Spice acceleration as an active cache: a materialization or data prefetcher. A cache fetches on miss; Spice prefetches and materializes filtered data on an interval, trigger, or via CDC. Spice also supports results caching.
- Is Spice a CDN for databases? Yes — a common use-case is shipping a working set of a database, data lake, or data warehouse to where it's most frequently accessed: data-intensive applications and AI context.
- Can I use Spice without Spice Cloud? Yes, the entire runtime is open-source under Apache 2.0. Spice Cloud is an optional managed cluster.
Watch a 30-sec BI dashboard acceleration demo
See more demos on YouTube.
Supported Data Connectors
| Name | Description | Status | Protocol/Format |
| ---------------------------------- | ------------------------------------- | ----------------- | ---------------------------- |
| databricks (mode: delta_lake) | [Databricks][databricks] | Stable | S3/Delta Lake |
| delta_lake | Delta Lake | Stable | Delta Lake |
| dremio | [Dremio][dremio] | Stable | Arrow Flight |
| duckdb | DuckDB | Stable | Embedded |
| file | File | Stable | Parquet, CSV |
| github | GitHub | Stable | GitHub API |
| postgres | PostgreSQL (with native WAL CDC) | Stable | |
| s3 | [S3][s3] | Stable | Parquet, CSV |
| mysql | MySQL (with native binlog CDC) | Stable | |
| spice.ai | [Spice.ai][spiceai] | Stable | Arrow Flight |
| dynamodb | Amaz