Web Apps

Quick access to all projects

22 items
Content Library
All llm-d content -- published and planned. Blog submissions, interactive web apps, cost visualizations, community stories, and supporting material.
9 items
Conference & Field Material
Demos, tools, and resources for conferences, booth presentations, and field engineering conversations. Interactive demos, ROI calculator, and comparison tools.
3 items
Platform/Scale Content
Platform & Scale pillar resources: interactive architecture diagram, content plan (July-Dec 2026), and strategic hub. Full-stack AI infrastructure from API gateway to GPU metal.
Private
Downstream llm-d Content Plan
Net-new llm-d content roadmap tied to Red Hat OpenShift. 20-piece timeline with business value alignment, conference strategy, and social media content kit.
Live
CFP Tracker
Beautiful, animated web app for tracking conference submissions. AI-powered assistance for talk titles and abstracts via ask.llm-d integration.
llm-d
Taken Down
llm-d Learning Center
Interactive platform with deployment configurator, routing visualizer, capacity planner, and learning resources for llm-d. Site taken down pending community buy-in.
Awaiting review by llm-d leadership. Code preserved in repo.
llm-d
Taken Down
Show and Tell
Engineers share rough demos and cool work, we clean it up and amplify it. Site taken down pending community buy-in.
Awaiting proposal review by llm-d leadership. Code preserved in repo.

Content Library

llm-d on OpenShift content series -- blog submissions, interactive web apps, and supporting material

llm-d
1 item
Content 1: Deploy llm-d in 5 Minutes
Blog submission for deploying llm-d on Kubernetes. Quick-start guide walkthrough with step-by-step instructions.
llm-d
2 items
Content 2: llm-d Benchmark
47.5x faster TTFT with cache-aware routing on H200 GPUs. Interactive web app with 8 chart visualizations and blog submission.
llm-d
2 items
Content 3: Prefix-Cache-Aware Routing
Deep dive into llm-d's routing algorithm. Animated diagram showing how the EPP tracks KV cache state and scores routing decisions. Blog submission included.
llm-d
2 items
Content 4: Monitoring & Observability
Interactive Grafana-themed monitoring dashboard simulator. Three metric layers (EPP, vLLM, DCGM), five scenarios, setup guide, PromQL reference, and alert rules.
llm-d
2 items
Content 5: Inference Routing Landscape
Competitive landscape comparison of inference routing approaches with business view toggle. Interactive comparison web app scoring on vendor lock-in, hardware flexibility, and partner ecosystem. Blog submission included.
llm-d
Live
Content 6: Deploying llm-d with Disaggregated Prefill & Decode
How to separate prefill and decode phases across GPU pools for maximum throughput. Step-by-step deployment guide with performance results.
llm-d
Live
Content 7: Serving Multiple LoRA Adapters with llm-d
Deploy dozens of fine-tuned models on shared GPU infrastructure. How llm-d routes requests to the right adapter with minimal memory overhead.
llm-d
Live
Content 8: llm-d for Multi-Tenant AI on OpenShift
Isolation, fair scheduling, and cost allocation across teams sharing GPU infrastructure. The multi-tenant inference platform story.
llm-d
Live
Content 9: llm-d at Scale: Lessons from Production
What we learned running llm-d in production. Scaling patterns, failure modes, and the operational playbook for distributed inference.
llm-d
Live
Content 10: llm-d + OpenShift AI: Enterprise LLM Serving
How llm-d fits into Red Hat’s AI platform. From open-source project to enterprise-grade inference with OpenShift AI integration.
llm-d
Live
Content 11: 2026 Year in Review: The State of Open Source LLM Inference
Community growth, benchmark milestones, partner ecosystem expansion, and the roadmap ahead. The annual state-of-the-project piece.
llm-d
Live
Content 12: The GPU Cost Cliff
Interactive cost visualization: drag from 10 to 1,000 concurrent users. Real GPU pricing. Terrain-style viz. The CFO conversation starter.
llm-d
Live
Content 13: “What Breaks at Scale” -- POC-to-Production Walkthrough
Interactive walkthrough of the 5 things that break when you go from POC to production inference. Walk through each failure mode and see how llm-d solves it.
llm-d
Live
Content 14: “Who’s Building llm-d” -- The Community Story
CoreWeave, Google, IBM, NVIDIA, AMD, Microsoft, Oracle -- why each showed up and what it signals. The ecosystem momentum piece.
llm-d
Live
Content 15: “Choose Your Own Inference Stack” -- Config Generator
Answer 5 questions about your workload, get a visual architecture diagram and shareable deployment config. The tool someone bookmarks and sends to their team.
llm-d
Live
Content 16: Agentic Demand Visualizer
Visualize agent inference patterns -- bursty, multi-turn, tool-calling. Compare naive routing vs. llm-d sticky sessions and prefix hits. Air traffic control aesthetic.
llm-d
Live
Content 17: “Inference Under the Hood” -- Transparent Cluster
Visual cluster sandbox: see pods, send requests, watch the EPP route them in real time, see KV cache update, watch responses stream back.
llm-d
Live
Content 19: llm-d Routing Deep Dive -- Interactive Technical Paper
Distill.pub-style long-form explainer with interactive inline figures. Scroll and diagrams animate, charts respond to input.
llm-d
Live
Content 20: From llm-d to Red Hat AI Inference Server
How the open-source llm-d project becomes the Red Hat AI Inference Server inside Red Hat AI 3. Architecture, packaging, support, and what it means for enterprises.
llm-d
Recurring
“llm-d vs. the Field”
Where TensorRT-LLM wins. Where SGLang shines. Where llm-d leads. What we’re still working on. Radical honesty as developer advocacy.
llm-d
Planned
Content 21: “llm-d Explained Simply”
Explain llm-d using airport gate routing, restaurant kitchen management, or highway traffic flow. Zero jargon. The piece a VP sends to their VP.
llm-d
Planned
Content 22: vLLM Office Hours -- llm-d Integration
Alternate existing vLLM office hours with llm-d sessions. Find speakers from engineering, community, and production users. Phase 2: spin into dedicated llm-d office hours if traction builds.

Content 1: Deploy llm-d in 5 Minutes

Blog submission for the llm-d quick-start deployment guide

Content 2: llm-d Benchmark

Interactive benchmark web app and blog submission for cache-aware routing results

Content 3: Prefix-Cache-Aware Routing

Animated diagram web app and blog submission explaining the llm-d routing algorithm

Content 4: Monitoring & Observability

Interactive Grafana-themed dashboard simulator with setup guides for monitoring llm-d on OpenShift

Content 5: Inference Routing Landscape

Competitive comparison of inference routing approaches and blog submission

Content 6: Deploying llm-d with Disaggregated Prefill & Decode

Interactive web app and blog submission for disaggregated prefill and decode deployment

Content 7: Serving Multiple LoRA Adapters with llm-d

Interactive web app and blog submission for LoRA adapter routing

Content 8: llm-d for Multi-Tenant AI on OpenShift

Interactive web app and blog submission for multi-tenant inference on OpenShift

Content 9: llm-d at Scale: Lessons from Production

Interactive web app and blog submission for production lessons at scale

Content 10: llm-d + OpenShift AI: Enterprise LLM Serving

Interactive web app and blog submission for OpenShift AI integration

Content 11: 2026 Year in Review: The State of Open Source LLM Inference

Interactive web app and blog submission for the 2026 year in review

Content 12: The GPU Cost Cliff

Interactive cost visualization web app and blog submission

Content 13: “What Breaks at Scale” -- POC-to-Production Walkthrough

Interactive walkthrough web app and blog submission

Content 14: “Who’s Building llm-d” -- The Community Story

Interactive web app and blog submission for the community story

Content 15: “Choose Your Own Inference Stack” -- Config Generator

Interactive config generator web app and blog submission

Content 16: Agentic Demand Visualizer

Interactive agentic demand visualization web app and blog submission

Content 17: “Inference Under the Hood” -- Transparent Cluster

Interactive transparent cluster web app and blog submission

Content 19: llm-d Routing Deep Dive -- Interactive Technical Paper

Distill-style interactive web app and blog submission

Content 20: From llm-d to Red Hat AI Inference Server

Interactive web app and blog submission for the llm-d to RHAIS journey

“llm-d vs. the Field”

Recurring series -- Blog (Recurring)

llm-d
Recurring
“llm-d vs. the Field”
Where TensorRT-LLM wins. Where SGLang shines. Where llm-d leads. What we’re still working on. Radical honesty as developer advocacy.

Content 21: “llm-d Explained Simply”

Planned -- Blog / Video

llm-d
Planned
“llm-d Explained Simply”
Explain llm-d using airport gate routing, restaurant kitchen management, or highway traffic flow. Zero jargon. The piece a VP sends to their VP.

Content 22: vLLM Office Hours -- llm-d Integration

Planned -- Community · Live Sessions

llm-d
Phase 1
Every Other vLLM Office Hours = llm-d
Alternate existing vLLM office hours with llm-d sessions starting July 2026. Find speakers from engineering, community, and production users. Uses existing timeslot and audience -- no cold start.
llm-d
Phase 2
Dedicated llm-d Office Hours
If alternating sessions gain traction and there’s enough speaker pipeline, spin llm-d into its own standalone office hours series. Community-driven, not marketing-driven.

Conference & Field Material

Demos, tools, and resources for conferences, booth presentations, and field engineering conversations

llm-d
Live
llm-d Interactive Demo
Interactive storytelling demo for llm-d. Walk someone through what cache-aware routing does in 2 minutes at a booth.
llm-d
Live
Ask llm-d
Domain-specific AI agent for llm-d. Explain architecture, generate deployment configs, and simulate routing performance -- tailored to 17 industries. Powered by Llama 3.3 70B via Groq.
Live
ROI Calculator
GPU savings calculator for cache-aware routing. Enter model size, GPU type, and request volume -- see how many fewer GPUs you need with llm-d. Includes before/after comparison (single-node vLLM vs. llm-d distributed), cost cliff visualization, and industry presets for banking, telco, and healthcare.
llm-d
Live
Inference Routing Landscape
Interactive comparison of inference routing approaches: round-robin, KServe, standalone vLLM, SGLang, and llm-d. Includes a business view toggle that scores competitors on vendor lock-in risk, hardware flexibility, enterprise support, and partner ecosystem -- not just routing features.
Live
Live Demo Environment
Simulated inference endpoint running llm-d. Send a prompt, see which pod it routes to, watch the cache hit on the second request. Proof over storytelling for the skeptic at the booth.
llm-d
Live
“What Breaks at Scale” -- Interactive Walkthrough
Start with working single-node inference, crank up load, watch things break visually. Then install llm-d components one by one and watch each problem resolve. The POC-to-production story told interactively.
llm-d
Live
“Inference Under the Hood” -- Transparent Cluster
Visual cluster sandbox: see pods, send requests, watch the EPP route them in real time, see KV cache update, watch responses stream back. The booth crowd magnet.
llm-d
Live
“Choose Your Own Inference Stack” -- Config Generator
Answer 5 questions about your workload, get a visual architecture diagram and shareable deployment config. The tool someone bookmarks and sends to their team.
llm-d
Live
Agentic Demand Visualizer
Visualize agent inference patterns -- bursty, multi-turn, tool-calling. Compare naive routing (cache thrashing) vs. llm-d (sticky sessions, prefix hits). Air traffic control aesthetic.

Platform/Scale Content

Full-stack AI infrastructure pillar: architecture, content plan, and strategic resources