Before you start
- A GPU with the NVIDIA driver installed. For the Docker path, also the NVIDIA Container Toolkit.
- The probe, running on the same node
- An organization API key. Create one under API keys in the dashboard.
127.0.0.1:8000. Adjust the URLs if yours differs.
1. Start the probe with vLLM settings
The plugin reports what happens inside the engine. These two settings add the server’s own metrics and traces.--vllm-metrics-url points at vLLM’s Prometheus endpoint.
--enable-otlp-receiver starts a receiver on 127.0.0.1:4317. vLLM exports per-request traces to it in the next step.
Start the probe first — it polls on --vllm-poll-interval until vLLM comes up.
See the full flag list.
2. Start vLLM
The plugin runs inside vLLM and the probe runs beside it, so it needs to know where to send what it collects. The plugin has to live in the same Python environment as vLLM.- Docker
- Virtualenv
vLLM’s official image already carries a matched torch, CUDA, and NCCL. Add the plugin on start.
--network host is required, because the probe runs on the host and both connections between the two use the host’s loopback. The probe scrapes vLLM at 127.0.0.1:8000, and vLLM sends traces to 127.0.0.1:4317. A container with its own network reaches neither.For anything long-lived, bake the plugin into an image instead of installing it on every start.Dockerfile
WarpscaleMpExecutor instead.
Tenant attribution
A single vLLM instance often serves requests from more than one tenant. The--middleware warpscale_vllm.middleware.TenantCapture flag in the commands above tags each request with the tenant that sent it, and the Warpscale · vLLM dashboard breaks tokens, KV cache occupancy, queue wait, and contention down per tenant.
Your gateway names the tenant on every request it forwards:
WS_VLLM_TENANT_HEADER if your gateway uses a different header name.
The value is lowercased, then matched against ^[a-z0-9][a-z0-9_-]{0,63}$. A request that fails the check, or arrives without the header, is served as normal and counted under _untagged. Serving one tenant is that same case, so leave the flag in.
Plugin settings
The plugin reads these env vars from the vLLM process.Trace settings
Traces are vLLM’s own OpenTelemetry export, so these env vars are read by the OTel SDK inside vLLM rather than by the plugin. The receiver is a local TCP port with no certificate, which is what both settings are about.3. Send some traffic
Fire the requests concurrently. A single request tells you the plugin works, but batching, queueing, and KV cache pressure only show up under load.4. See the data
Your server appears under Inference in the console, athttps://<your-organization-host>/inference. Open the instance to see every request it has served.
For metrics over time, open the Warpscale · vLLM dashboard and pick your instance. KV cache use, running and queued requests, queue wait, prefill throughput, and time per output token appear within a poll interval of the requests completing. The per-tenant row fills in alongside them once your gateway starts sending the tenant header.
Troubleshooting
vLLM fails to start after adding the executor backend
vLLM fails to start after adding the executor backend
The plugin is not in the environment vLLM is running from. Confirm with
python -c "import warpscale_vllm" using that interpreter, or docker run --rm --entrypoint python3 <your-image> -c "import warpscale_vllm" for the image.The server appears but has no request statistics
The server appears but has no request statistics
The probe cannot scrape vLLM. Confirm
--vllm-metrics-url is set on the probe, then fetch that same URL from the host with curl. A server on a non-default port needs the URL updated to match; a container without --network host is not reachable there at all.No engine-internal data, but metrics are arriving
No engine-internal data, but metrics are arriving
WS_VLLM_SINK is unset or points somewhere the probe is not listening. The probe accepts records on the path given by --user-event-socket, which defaults to /var/run/warpscale/warpscaled.sock.Every request lands under _untagged
Every request lands under _untagged
The header is not reaching the plugin, or its value was rejected. Confirm the gateway sets
X-Tenant-Id — or whatever WS_VLLM_TENANT_HEADER names — on the request it forwards to vLLM rather than only on the one it received. Then check the value itself: uppercase is fine and gets lowercased, but a space, a dot, a leading -, or anything past 64 characters is dropped._untagged is also what you get when --middleware warpscale_vllm.middleware.TenantCapture is missing from vllm serve.No traces, but metrics are arriving
No traces, but metrics are arriving
Two causes. The receiver accepts gRPC only, so confirm
OTEL_EXPORTER_OTLP_TRACES_PROTOCOL is unset. Then check the container is on --network host, since the receiver listens on the host’s loopback at 127.0.0.1:4317 and nothing outside that network namespace can reach it.The probe logs first engine span batch received once, when traces start arriving. No such line means nothing reached it, and the fault is between vLLM and the probe rather than after it.