llm-d/llm-d

▲ 525 stars today★ 4,550⑂ 770

Achieve state of the art inference performance with modern accelerators on Kubernetes

About llm-d/llm-d

llm-d/llm-d is an open-source project on GitHub, mainly written in Shell. Achieve state of the art inference performance with modern accelerators on Kubernetes It currently holds 4,550 stars and 770 forks with 0 open issues, and was last pushed on an unknown date (repository created unknown).

Project Overview

AI Homed tracks it on the Today's Trending board.

GitHub Repository Details

Repository llm-d/llm-d · default branch - · size 0 KB · watchers 0 · source: GitHub REST API and repository README

README

https://github.com/llm-d/llm-d/blob/HEAD/llm-d Logo

Achieve SOTA Inference Performance On Any Accelerator

Documentation FOSSA Status Release Status License Join Slack

llm-d is a high-performance distributed inference serving stack optimized for production deployments on Kubernetes. We help you achieve the fastest "time to state-of-the-art (SOTA) performance" for key OSS large language models across most hardware accelerators and infrastructure providers with well-tested guides and real-world benchmarks.

llm-d is a Cloud Native Computing Foundation (CNCF) sandbox project, founded by Red Hat, Google Cloud, IBM Research, CoreWeave, and NVIDIA.

What does llm-d offer to production inference?

Model servers like vLLM and SGLang handle efficiently running large language models on accelerators. llm-d provides state-of-the-art orchestration and optimizations above model servers to serve high-scale real-world traffic efficiently and reliably. Our offerings are organized into the following themes:

For a complete list of tested recipes and architectural patterns, see our well-lit path guides. These guides provide benchmarked recipes and Helm charts to start serving quickly with best practices common to production deployments. Our intent is to eliminate the heavy lifting common in tuning and deploying generative AI inference on modern accelerators.

Performance Highlights

Validated performance gains from production deployments and partner benchmarks:

Explore detailed, reproducible benchmarks on Prism.

Get Started Now

Ready to achieve SOTA performance? Follow our Quickstart Guide to deploy your first optimized inference service on Kubernetes. You'll learn how to set up the llm-d stack, configure the intelligent router, and validate performance with production-ready benchmarks.

[!TIP]
Most users begin with our Optimized Baseline, which provides a high-performance foundation for a wide range of LLM serving use cases.

Latest News 🔥

🧱 Architecture

llm-d accelerates distributed inference by integrating industry-standard open technologies like vLLM and Kubernetes. For more details, see our full Architecture Documentation.

https://github.com/llm-d/llm-d/blob/HEAD/llm-d Arch

📦 Releases

Our guides are living docs and kept current. For details about the Helm charts and component releases, visit our GitHub Releases page to review release notes.

See the accelerator docs for points of contact and more details about the accelerators, networks, and configurations tested.

Contribute

We adhere to the CNCF Code of Conduct.

License

This project is licensed under Apache License 2.0. See the LICENSE file for details.

FOSSA Status

GitHub Stars & Activity

4,550Stars
770Forks
0Open issues
ShellLanguage

GitHub Popularity

GitHub stars4,550
Forks770
Open issues0
Primary languageShell
License-
Stars gained today525
Created-
Last pushed-

Trending History

Monthly boardrank #66 · ▲ 525 stars

Related AI Projects

1

obra / superpowers

Shell★ 287,295⑂ 25,690▲ 522 stars
2

mattpocock / skills

Shell★ 263,045⑂ 22,188▲ 820 stars
3

a2aproject / A2A

Shell★ 25,792⑂ 2,610▲ 9 stars
4

Donchitos / Claude-Code-Game-Studios

Shell★ 25,131⑂ 3,588▲ 36 stars
5

kunchenguid / firstmate

Shell★ 6,086⑂ 1,887▲ 141 stars
6

ollama / ollama

Go★ 181,102⑂ 17,896▲ 140 stars
7

ggml-org / llama.cpp

C++★ 128,384⑂ 23,249▲ 128 stars
8

openai / codex

Rust★ 124,542⑂ 19,253▲ 314 stars

More AI Rankings