Slide 1 of 8
llm-d llm-d Open Source Framework

Year in Review

From 45 tok/s to 50,000 tok/s in 13 months. The story of how an open-source project rewrote the rules of distributed LLM inference on Kubernetes.

1,000x
Throughput Gain
CNCF
Sandbox Project
13+
Contributing Partners
6
Releases Shipped

Press or click to navigate

Slide 2 of 8

The Problem That Started It All

LLM inference was broken. Traditional load balancers treated every GPU pod as interchangeable, randomly spraying requests across the fleet.

🎯
Random Routing
Requests scattered across pods with no awareness of which GPU already held relevant KV cache data.
🔥
Wasted GPU Cycles
Cold starts on every request. Expensive prefill recomputed again and again across the cluster.
🔒
Vendor Lock-in
Proprietary inference platforms offered optimization but at the cost of flexibility and portability.
⚙️
No K8s Answer
The cloud-native ecosystem had no solution for intelligent, GPU-aware LLM inference scheduling.
Slide 3 of 8

A Year of Milestones

Seven inflection points in thirteen months. Each milestone built on the last, accelerating from research prototype to CNCF Sandbox.

Launch
May 2025
Red Hat Summit Launch
Founded by Red Hat, Google Cloud, IBM Research, CoreWeave, and NVIDIA.
Release
May 2025
v0.1 -- First Release
Baseline: ~45 tok/s. The foundation is laid.
Release
July 2025
v0.2 -- Well-Lit Paths
Prefix-cache-aware scheduling and prefill/decode disaggregation.
Release
October 2025
v0.3 -- Multi-Hardware
Intel XPU, Google TPU, NIXL integration. 2,200 tok/s per GPU.
Release
February 2026
v0.5 -- 50,000 tok/s
16x16 B200 topology, HA failover, scale-to-zero autoscaling.
CNCF
March 2026
CNCF Sandbox Accepted
Accepted at KubeCon Europe, Amsterdam. Donated by IBM, Red Hat, Google Cloud.
Release
June 2026
v0.7 -- Standalone Maturity
Standalone mode, expanded integrations, production stability.
Slide 4 of 8

Relentless Acceleration

From 45 tok/s on day one to 50,000 tok/s on a 16x16 B200 topology. Every release pushed the boundary further.

1,000x improvement in 13 months
Slide 5 of 8

The All-Star Roster

13+ organizations collaborating across core development, hardware, cloud, research, and AI models.

Core Development
Red Hat Google Cloud
Hardware
NVIDIA AMD Intel
Cloud
CoreWeave Cisco Lambda
Research & Models
IBM Research UC Berkeley (vLLM) Univ of Chicago Hugging Face Mistral AI
13+
Partners
85+
Contributors
340+
PRs Merged
1,200+
GitHub Stars
Slide 6 of 8

Production-Grade Features

What makes llm-d different: intelligent routing, disaggregated inference, and enterprise-ready resilience.

🧠
Intelligent Routing
The Endpoint Picker (EPP) scores every pod by cache affinity, queue depth, and active load. Requests land on the GPU most likely to serve them instantly, achieving 80%+ cache hit rates.
Prefill/Decode Split
Separate GPU pools for prompt processing and token generation. Each phase scales independently, connected by high-speed NIXL interconnect for KV cache transfer.
🚀
Scale-to-Zero & HA
Active-active high availability with no single point of failure. Scale-to-zero autoscaling releases idle GPUs. Hierarchical KV cache offloading preserves context across scale events.
Slide 7 of 8

CNCF Sandbox Acceptance

March 24, 2026 -- KubeCon Europe, Amsterdam. Just 10 months from public announcement to CNCF acceptance.

Accepted into the Cloud Native Computing Foundation Sandbox. A testament to the velocity of the community and the urgency of the problem.

Donated By

IBM Research
Red Hat
Google Cloud
10
Months to Acceptance
13+
Contributing Partners
3
Donating Organizations
Slide 8 of 8

The Road Ahead

From sandbox to production-grade. The next chapter of llm-d is being written right now.

v1.0 Stable API
GA-quality APIs, backwards compatibility guarantees, and full documentation for production adopters.
75%
Multi-Cluster Federation
Route inference requests across multiple Kubernetes clusters with unified scheduling.
40%
Agentic Inference
First-class support for multi-turn agentic workloads with session affinity and tool-call routing.
30%
Cost-Aware Routing
Routing decisions that factor in GPU cost, spot vs. on-demand pricing, and multi-cloud economics.
15%