From 45 tok/s to 50,000 tok/s in 13 months. The story of how an open-source project rewrote the rules of distributed LLM inference on Kubernetes.
Press → or click to navigate
LLM inference was broken. Traditional load balancers treated every GPU pod as interchangeable, randomly spraying requests across the fleet.
Seven inflection points in thirteen months. Each milestone built on the last, accelerating from research prototype to CNCF Sandbox.
From 45 tok/s on day one to 50,000 tok/s on a 16x16 B200 topology. Every release pushed the boundary further.
13+ organizations collaborating across core development, hardware, cloud, research, and AI models.
What makes llm-d different: intelligent routing, disaggregated inference, and enterprise-ready resilience.
March 24, 2026 -- KubeCon Europe, Amsterdam. Just 10 months from public announcement to CNCF acceptance.
Accepted into the Cloud Native Computing Foundation Sandbox. A testament to the velocity of the community and the urgency of the problem.
Donated By
From sandbox to production-grade. The next chapter of llm-d is being written right now.