Slide 1 of 7
llm-d llm-d Production Lessons

12 Incidents. 0 Survivors Who Didn't Learn Something.

Real production failures from running llm-d at scale. The kind that page you at 2 AM, crash your SLOs, and teach you everything the docs could not.

4
P0 Incidents
14h 23m
Total Downtime
12
Lessons Learned
Slide 2 of 7

How Production Fails

Failures cascade. A single misconfiguration can trigger a chain reaction that takes down your entire serving cluster in under 90 seconds.

Slide 3 of 7

The 12 Incidents

Each one paged someone at the worst possible time. Each one taught a lesson no documentation could.

P0
KV Cache Eviction Storm
Cache hit rate cliff-dived from 85% to 3% in 90 seconds. Pure LRU evicted hot system prompts first.
P0
Model Download Thundering Herd
24 pods pulling 140GB simultaneously. 38-minute total outage from saturated network.
P0
RDMA Network Partition
Switch failure. Pods look healthy. Requests hang forever. Every dashboard showed green.
P0
Prefill Pod Starvation
50 concurrent 128K requests. Short-context TTFT jumped from 200ms to 45 seconds.
P1
Autoscaler Oscillation
HPA thrashing 8 to 20 pods every 5 minutes. Pods terminated mid-inference.
P1
GPU Memory Fragmentation
OOM kills with 40% free memory. nvidia-smi was lying. Memory scattered in fragments.
P1
NUMA Node Misplacement
4 pods at 38ms/tok instead of 18ms/tok. GPU was fine. Memory path was not.
P1
Health Check False Positives
Pods say "ready." Model not loaded. 5,000 requests timed out during rolling update.
P1
Prefix Cache Collision
"The model is hallucinating." It was not. Hash collision caused cross-tenant cache hits.
P1
ConfigMap Update Race
30 seconds of split-brain routing. Cache coherence destroyed by eventual consistency.
P2
Token Budget Exhaustion
"The model got dumber." 340K silently truncated responses. Zero alerts fired for 3 days.
P2
Decode Latency from Defrag
Every 6 hours, defrag fired and decode latency jumped 5x. The fix needed a fix.
Slide 4 of 7

Your First Week Checklist

Every item addresses a specific failure mode from the incidents. Complete these before your first real traffic.

Deploy model cache DaemonSet to node-local NVMe
Configure compound readiness probes (/health + /v1/models + warmup)
Set HPA stabilization to 15+ minutes
Set Topology Manager to single-numa-node
Add GPU fragmentation monitoring (alert at 0.3)
Monitor finish_reason distribution (alert at 5% length)
Configure PodDisruptionBudgets (minAvailable 75%)
Set KV cache eviction to frequency-weighted
Add RDMA health checks to readiness probes
Deploy priority queues by context length
Slide 5 of 7

The Patterns We Learned

Every incident leaves behind a story. These are the lessons we tell new engineers on their first day.

Cache is Your SLO
Your cache eviction policy IS your SLO policy under load. LRU evicts your hottest system prompts first. Use frequency-weighted eviction and set admission control thresholds 15% below your eviction trigger.
Models are Infrastructure
Model weights are infrastructure, not application artifacts. If your model is not on local NVMe before the pod starts, you have already lost. Treat model loading as a first-class infrastructure concern.
Two Networks, One Blind Spot
In disaggregated serving, you have TWO networks. Kubernetes only monitors one. RDMA will fail silently at 4 AM and every dashboard will show green. Monitor both independently.
Silent Truncation Kills Quality
Silent truncation is worse than a loud error. Always monitor finish_reason distribution. We spent 3 days investigating "model quality" before finding a one-line config error.
GPU Memory Lies to You
nvidia-smi shows 40% free. It does not show you that the free memory is scattered across 10,000 fragments too small to use. Monitor fragmentation ratio, not just total free.
Not All Tokens Are Equal
A 128K prefill consumes the same compute as 256 short requests. Without priority queues, your shortest requests get hurt the most. Your scheduler needs context-length awareness.
Slide 6 of 7

Where Things Break

A production llm-d cluster has dozens of failure points. This diagram shows the most critical ones, with live data flow and failure indicators.

Slide 7 of 7

Start Here, Before 3 AM Finds You

Every lesson on these slides came from a real pager alert. The checklist exists so yours does not have to.

Reading about incidents is one thing. Diagnosing them under pressure is another. The cluster is live. Something is about to break. Can you find the fix before the SLO burns?
Built for SREs who got paged at 3 AM