Send an Inference Request
Type a prompt or choose a preset, then hit Send. Watch how llm-d routes the request.
You are a helpful AI assistant specializing in cloud-native infrastructure. Answer concisely and technically.
Quick Prompts
Request History
No requests sent yet. Hit Send to begin.
Cluster State
3-pod vLLM cluster on H200 GPUs · Llama 3.1 70B (FP8)
Cache-Aware Routing (llm-d EPP)
Idle — waiting for request
Gateway API
llm-d EPP
0
Total Requests
0
Cache Hits
Hit Rate
Avg TTFT
Send a request to see the inference response and timing breakdown.