https://<your-organization-host>/grafana.
Dashboards
The bundled dashboards are in the Warpscale folder.Warpscale · vLLM
Serving behaviour for one vLLM instance, picked from the dropdown. Two rows.- vLLM Instance — peak KV cache use, peak running and waiting requests, and the hourly medians for queue wait, prefill throughput, and time per output token
- vLLM Tenant — the same instance broken down by tenant: prefill and decode tokens, token share, KV residency and occupancy, queue wait, decode excess, contention, contention per request, finished requests
_untagged.