Real production failures from running llm-d at scale. The kind that page you at 2 AM, crash your SLOs, and teach you everything the docs could not.
Failures cascade. A single misconfiguration can trigger a chain reaction that takes down your entire serving cluster in under 90 seconds.
Each one paged someone at the worst possible time. Each one taught a lesson no documentation could.
Every item addresses a specific failure mode from the incidents. Complete these before your first real traffic.
Every incident leaves behind a story. These are the lessons we tell new engineers on their first day.
A production llm-d cluster has dozens of failure points. This diagram shows the most critical ones, with live data flow and failure indicators.
Every lesson on these slides came from a real pager alert. The checklist exists so yours does not have to.