Content 1: Deploy llm-d in 5 Minutes
Blog submission for deploying llm-d on Kubernetes. Quick-start guide walkthrough with step-by-step instructions.
Content 2: llm-d Benchmark
47.5x faster TTFT with cache-aware routing on H200 GPUs. Interactive web app with 8 chart visualizations and blog submission.
Content 3: Prefix-Cache-Aware Routing
Deep dive into llm-d's routing algorithm. Animated diagram showing how the EPP tracks KV cache state and scores routing decisions. Blog submission included.
Content 4: Monitoring & Observability
Interactive Grafana-themed monitoring dashboard simulator. Three metric layers (EPP, vLLM, DCGM), five scenarios, setup guide, PromQL reference, and alert rules.
Content 5: Inference Routing Landscape
Competitive landscape comparison of inference routing approaches with business view toggle. Interactive comparison web app scoring on vendor lock-in, hardware flexibility, and partner ecosystem. Blog submission included.
Content 6: Deploying llm-d with Disaggregated Prefill & Decode
How to separate prefill and decode phases across GPU pools for maximum throughput. Step-by-step deployment guide with performance results.
Content 7: Serving Multiple LoRA Adapters with llm-d
Deploy dozens of fine-tuned models on shared GPU infrastructure. How llm-d routes requests to the right adapter with minimal memory overhead.
Content 8: llm-d for Multi-Tenant AI on OpenShift
Isolation, fair scheduling, and cost allocation across teams sharing GPU infrastructure. The multi-tenant inference platform story.
Content 9: llm-d at Scale: Lessons from Production
What we learned running llm-d in production. Scaling patterns, failure modes, and the operational playbook for distributed inference.
Content 10: llm-d + OpenShift AI: Enterprise LLM Serving
How llm-d fits into Red Hat’s AI platform. From open-source project to enterprise-grade inference with OpenShift AI integration.
Content 11: 2026 Year in Review: The State of Open Source LLM Inference
Community growth, benchmark milestones, partner ecosystem expansion, and the roadmap ahead. The annual state-of-the-project piece.
Content 12: The GPU Cost Cliff
Interactive cost visualization: drag from 10 to 1,000 concurrent users. Real GPU pricing. Terrain-style viz. The CFO conversation starter.
Content 13: “What Breaks at Scale” -- POC-to-Production Walkthrough
Interactive walkthrough of the 5 things that break when you go from POC to production inference. Walk through each failure mode and see how llm-d solves it.
Content 14: “Who’s Building llm-d” -- The Community Story
CoreWeave, Google, IBM, NVIDIA, AMD, Microsoft, Oracle -- why each showed up and what it signals. The ecosystem momentum piece.
Content 15: “Choose Your Own Inference Stack” -- Config Generator
Answer 5 questions about your workload, get a visual architecture diagram and shareable deployment config. The tool someone bookmarks and sends to their team.
Content 16: Agentic Demand Visualizer
Visualize agent inference patterns -- bursty, multi-turn, tool-calling. Compare naive routing vs. llm-d sticky sessions and prefix hits. Air traffic control aesthetic.
Content 17: “Inference Under the Hood” -- Transparent Cluster
Visual cluster sandbox: see pods, send requests, watch the EPP route them in real time, see KV cache update, watch responses stream back.
Content 19: llm-d Routing Deep Dive -- Interactive Technical Paper
Distill.pub-style long-form explainer with interactive inline figures. Scroll and diagrams animate, charts respond to input.
Content 20: From llm-d to Red Hat AI Inference Server
How the open-source llm-d project becomes the Red Hat AI Inference Server inside Red Hat AI 3. Architecture, packaging, support, and what it means for enterprises.
“llm-d vs. the Field”
Where TensorRT-LLM wins. Where SGLang shines. Where llm-d leads. What we’re still working on. Radical honesty as developer advocacy.
Content 21: “llm-d Explained Simply”
Explain llm-d using airport gate routing, restaurant kitchen management, or highway traffic flow. Zero jargon. The piece a VP sends to their VP.
Content 22: vLLM Office Hours -- llm-d Integration
Alternate existing vLLM office hours with llm-d sessions. Find speakers from engineering, community, and production users. Phase 2: spin into dedicated llm-d office hours if traction builds.