llm-d/llm-d-router

★ 374⑂ 419

llm-d Router: The intelligent entry point for inference requests

About llm-d/llm-d-router

llm-d/llm-d-router is an open-source project on GitHub, mainly written in Go. llm-d Router: The intelligent entry point for inference requests It currently holds 374 stars and 419 forks with 0 open issues, and was last pushed on an unknown date (repository created unknown).

Project Overview

AI Homed tracks it on the Today's Trending board, currently at rank #69 with 0 new stars today.

GitHub Repository Details

Repository llm-d/llm-d-router · default branch - · size 0 KB · watchers 0 · source: GitHub REST API and repository README

README

Go Report Card Go Reference License Join Slack FOSSA Status

llm-d Router

[!IMPORTANT]
Terminology Change: The Inference Scheduler has been renamed to llm-d Router; see Terminology.
[!IMPORTANT]
API & Code Consolidation: Core Endpoint Picker (EPP) code and the InferenceObjective and InferenceModelRewrite APIs have been merged into this repository from [Gateway API Inference Extension (GIE)]. The GIE repository now exclusively hosts the InferencePool API—an extension of the [Kubernetes Gateway API]—and defines the Endpoint Picker Protocol.

The llm-d Router is the intelligent entry point for inference traffic, delivering LLM load and prefix-cache aware routing, request prioritization, and advanced flow control across diverse request formats to fulfill complex serving objectives. It supports a flexible deployment model: it can run in Standalone Mode, where a self-managed Envoy proxy runs either alongside the EPP or as a separate service, or integrate with L7 load balancers—including self-managed instances (e.g., Istio, AgentGateway) and cloud-managed services (e.g., Google Cloud's Application Load Balancer)—via the Kubernetes Gateway API.

The router achieves its intelligence through an Endpoint Picker (EPP) that integrates with production-grade proxies (such as [Envoy]) via the [ext-proc] protocol, injecting real-time signals into the data plane to optimize request placement.

https://github.com/llm-d/llm-d-router/blob/HEAD/llm-d Router Architecture

Core Components and APIs

This repository hosts the following core components:

Modes of Operation

The llm-d Router supports two primary deployment modes as specified in the [Kubernetes Gateway API Inference Extensions]:

1. Standalone Mode

A deployment with a self-managed Envoy proxy that does not require Gateway API infrastructure. The standalone Helm chart supports two Envoy proxy topologies: See the [Helm chart documentation] for configuration examples.

2. Gateway Mode (Inference Gateway)

The recommended mode for production environments, leveraging the official [Gateway API]. In this mode, the EPP acts as a backend for an InferencePool, which is referenced by an HTTPRoute on a shared Gateway. This enables advanced traffic management, multi-cluster load balancing, and shared infrastructure for both inference and traditional workloads.

For more details on the router architecture, routing logic, and different plugins (filters and scorers), see the [Architecture Documentation]. For resource provisioning and container sizing recommendations under heavy or long-context workloads, see the [EPP Container Sizing Guide]. The [OpenTelemetry JSON stdout logs] document describes the log record format used by the EPP and routing sidecar. The [TLS Documentation] covers the serving and metrics TLS flags of the EPP, the routing sidecar and the coordinator.

---

[!NOTE]
The project provides tools for automatic Envoy installation. However, if you install or
configure it yourself, please note that the only supported request_body_mode and response_body_mode
is FULL_DUPLEX_STREAMED

Terminology

To ensure clarity across the project, we use the following standard terminology:

[Kubernetes]:https://kubernetes.io [Kubernetes Gateway API]:https://gateway-api.sigs.k8s.io/ [Architecture Documentation]:docs/architecture.md [Disaggregation Documentation]:docs/disaggregation.md [EPP Container Sizing Guide]:docs/operations.md [TLS Documentation]:docs/tls.md [OpenTelemetry JSON stdout logs]:docs/otel-json-stdout.md [InferencePool]:https://github.com/kubernetes-sigs/gateway-api-inference-extension [Gateway API Inference Extension (GIE)]:https://github.com/kubernetes-sigs/gateway-api-inference-extension [Kubernetes Gateway API Inference Extensions]:https://github.com/kubernetes-sigs/gateway-api-inference-extension [Gateway API]:https://github.com/kubernetes-sigs/gateway-api [Envoy]:https://github.com/envoyproxy/envoy [ext-proc]:https://www.envoyproxy.io/docs/envoy/latest/configuration/http/http_filters/ext_proc_filter [Helm chart documentation]:config/charts/README.md

Contributing

Start with the [llm-d organization contributing guide][org-contributing] for project-wide guidelines, code of conduct, and community resources, then see CONTRIBUTING.md for what is specific to this repository, including how to claim an issue.

Our community meeting is bi-weekly at Wednesday 10AM PDT ([Google Meet], [Meeting Notes]).

We currently utilize the [#sig-router] channel in llm-d Slack workspace for communications.

For large changes please [create an issue] first describing the change so the maintainers can do an assessment, and work on the details with you. See DEVELOPMENT.md for details on how to work with the codebase.

Contributions are welcome!

[org-contributing]:https://github.com/llm-d/llm-d/blob/main/CONTRIBUTING.md [create an issue]:https://github.com/llm-d/llm-d-router/issues/new [discussion]:https://github.com/llm-d/llm-d-router/discussions/new?category=q-a [Slack]:https://llm-d.slack.com/ [Google Meet]:https://meet.google.com/ozx-goao-cxh [Meeting Notes]:https://docs.google.com/document/d/1Pf3x7ZM8nNpU56nt6CzePAOmFZ24NXDeXyaYb565Wq4 [#sig-router]:https://llm-d.slack.com/?redir=%2Fmessages%2Fsig-router

Security

See SECURITY.md for vulnerability reporting. Published container images carry a signed provenance attestation and an SBOM (software bill of materials). See Verifying Published Artifacts for how to check them.

License

FOSSA Status

GitHub Stars & Activity

374Stars
419Forks
0Open issues
GoLanguage

GitHub Popularity

GitHub stars374
Forks419
Open issues0
Primary languageGo
License-
Stars gained today0
Created-
Last pushed-

Trending History

Daily boardrank #69 · ▲ 0 stars

Related AI Projects

1

ollama / ollama

Go★ 182,379⑂ 18,117▲ 137 stars
→
2

JuliusBrussee / caveman

Go★ 110,188⑂ 6,379▲ 211 stars
→
3

livekit / livekit

Go★ 21,307⑂ 2,420▲ 16 stars
→
4

slavakurilyak / awesome-ai-agents

Go★ 2,360⑂ 604▲ 14 stars
→
5

Autumn-27 / ARTEX

Go★ 1,689⑂ 385▲ 111 stars
→
6

ys-ll / uniterm

Go★ 743⑂ 114▲ 30 stars
→
7
→
8

jiwoochris / artex-ko

Go★ 67⑂ 29
→

More AI Rankings