Plan, compare, and build production-ready LLM serving architectures with llm-d. Answer five questions and get sizing, routing, and Helm values in minutes, not weeks.
Model parameters determine GPU memory, tensor parallelism, routing strategy, and whether disaggregated serving is beneficial. Five decisions shape everything.
Every request flows through four layers, each one optimized for production inference at scale.
Your model size and GPU budget determine the architecture. The decision tree maps every combination to a recommended stack.
Answer five questions and get a complete llm-d architecture: GPU sizing, EPP routing strategy, Helm values, and deployment guidance.
From MVP to hyperscale, each stack is battle-tested and production-ready.
Where each llm-d component sits on the adoption curve. From proven production staples to emerging capabilities.
Wherever you are today, there is a tested path forward. Effort estimates and rollback strategies for every source framework.
llm-d is a CNCF Sandbox project. 100% open source, zero vendor lock-in, Kubernetes-native from day one. Your inference stack, your way.