Metrics by component
Scrape each producer pod. Retain namespace, pod UID and engine identity. These are reference names—match the deployed version.
Endpoint Picker
Scrape each Endpoint Picker pod’s verified metrics URL. Connect to the pod · Save a scrape.
| Metric / field | Plain English | Why needed | What it shows |
|---|---|---|---|
llm_d_epp_info | Build identity | Match behavior to the running build. | Commit/build reference; confirm image digest if unknown. |
llm_d_epp_request_total | Requests scheduled | Confirm traffic passed through routing. | Count by model, fairness ID and priority; not client successes. |
llm_d_epp_request_error_total | Router-observed errors | Retain failures in the comparison. | Counts by workload and error code. |
llm_d_epp_request_running | Requests counted in flight | Check router request accounting. | Current count; separate from engine-running requests. |
llm_d_epp_request_ttft_seconds | Router first-response delay | Compare router timing with client timing. | Seconds to first response byte; histogram. |
llm_d_epp_request_streaming_tpot_seconds | Router average decode time per token | Compare streaming cadence across matched runs. | Per-request (duration − TTFT) / (output tokens − 1), in seconds. This histogram contains request averages, not individual chunk gaps. |
llm_d_epp_request_duration_seconds | Router request duration | Locate delay across the request path. | Router-observed duration in seconds; histogram. |
llm_d_epp_scheduler_attempts_total | Scheduling attempts | Identify the selected backend. | Status and endpoint identity; attempts are not completions. |
llm_d_epp_flow_control_requests_total | Admission outcomes | Verify the admission path was used. | Dispatch/rejection outcomes by priority and pool. |
llm_d_epp_flow_control_queue_size | Requests waiting in admission | Detect waiting before model dispatch. | Queued request count by workload. |
llm_d_epp_flow_control_queue_bytes | Queued request bytes | Check the configured byte budget. | Sum of queued request sizes. This is not total process memory. |
llm_d_epp_flow_control_request_queue_duration_seconds | Time in admission | Measure the cost of waiting. | Seconds from enqueue to outcome; histogram. |
llm_d_epp_flow_control_pool_saturation | Detector saturation signal | Compare the signal with the loaded ceiling. | Pool saturation ratio at the stated router pin. Compare with the loaded policy ceiling. |
llm_d_epp_per_endpoint_queue_size | Router view of engine waiting | Compare router input with engine observations. | Waiting requests per mapped endpoint. |
llm_d_epp_datalayer_poll_errors_total | Metric-source polling failures | Detect missing inputs to the router. | Error count by source type. |
llm_d_epp_datalayer_extract_errors_total | Metric extraction failures | Detect unusable collected inputs. | Error count by source/extractor type. |
llm_d_epp_flow_control_stale_endpoints | Missing or old utilization samples | Separate stale input from actual pressure. | Count from the latest detector evaluation; not all-stage proof. |
vLLM
Scrape each engine’s metrics endpoint; retain pod and engine identity. Connect to the pod · Save a scrape.
| Metric / field | Plain English | Why needed | What it shows |
|---|---|---|---|
vllm:request_success_total | Finished requests by reason | Reconcile engine outcomes with the client. | Counts include abort/error reasons; check normal completions separately. |
vllm:generation_tokens_total | Generated tokens | Confirm work reached the engine. | Cumulative generated-token count. |
vllm:num_requests_running | Engine-running requests | Observe active model work. | Current requests in execution batches. |
vllm:num_requests_waiting | Engine-waiting requests | Distinguish engine waiting from admission waiting. | Current engine queue length. |
vllm:kv_cache_usage_perc | KV-cache occupancy | Observe engine cache pressure. | Fraction from 0 to 1; not total GPU memory. |
vllm:time_to_first_token_seconds | Engine first-token delay | Compare engine timing with client timing. | Seconds to first token; histogram. |
vllm:request_time_per_output_token_seconds | Engine average decode time per token | Compare engine timing with the client. | Per-request average time per output token, in seconds. Keep separate from the engine inter-token gap histogram. |
vllm:e2e_request_latency_seconds | Engine request duration | Measure server-side latency. | Total server latency in seconds; histogram. |
vllm:request_queue_time_seconds | Engine queue time | Locate waiting inside the server. | Seconds waiting in the engine; histogram. |
vllm:prefix_cache_queries_total | Tokens checked for prefix reuse | Interpret cache-hit counts. | Cumulative queried-token count. |
vllm:prefix_cache_hits_total | Tokens found in prefix cache | Track cache differences between runs. | Hit-token count; compare with queries over the same window. |
AIPerf
Read the client results in the AIPerf run’s output directory. Run and export command.
| Metric / field | Plain English | Why needed | What it shows |
|---|---|---|---|
request_count | Successful client records | Confirm completed client work. | Successful-record count in native exports. |
error_request_count | Failed client records | Prevent failed requests disappearing from results. | Error-record count; retain error details. |
time_to_first_token | Client first-content delay | Measure what the caller experiences. | Per-request first-token timing; read exported units. |
inter_token_latency | Client average decode time per token | Evaluate the streaming token-latency criterion. | Per-request (request latency − TTFT) / tokens after the first content chunk. With unavailable or invalid first-chunk usage, the divisor is output tokens − 1. Requires at least two output tokens. Read exported units. This is not an individual-gap distribution. |
request_latency | Client request duration | Measure end-to-end user latency. | Per-request duration; read exported units. |
input_sequence_length | Actual prompt length | Compare input-size distributions across runs. | Input-token count under the runner’s tokenizer. |
output_sequence_length | Actual response length | Keep the amount of generated work comparable. | Output-token count under the runner’s tokenizer. |
Kubernetes API
Read pod state from the Kubernetes API: kubectl --context "$CTX" -n "$NS" get pod "$POD" -o json.
| Metric / field | Plain English | Why needed | What it shows |
|---|---|---|---|
status.conditions | Pod readiness | Verify the test path stays available. | Ready condition and transitions. |
status.containerStatuses[].restartCount | Container restarts | Detect interrupted or reset processes. | Restart count per container. |
status.containerStatuses[].lastState | Previous termination state | Investigate serving faults. | Previous termination reason and timing. |
status.containerStatuses[].imageID | Running image identity | Match evidence to the deployed software. | Runtime image identifier/digest. |
Streaming measurement boundaries. AIPerf reports client timing, the router reports its observed response timing, and vLLM reports engine timing. Compare matched requests and intervals. Do not subtract unrelated percentiles. For individual gaps, the router exposes llm_d_epp_request_streaming_itl_seconds for consecutive response body chunks. These chunks can contain multiple tokens.
Evidence for conditional controls
Collect router series from each verified Endpoint Picker metrics endpoint and engine series from each model server. Preserve request, workload, priority, tenant and endpoint identity in supporting records. Optional series require their producer to be active.
| Component / producer | Exact keys or evidence | Use and limits |
|---|---|---|
| Client · AIPerf records | request_counterror_request_counttime_to_first_tokenrequest_latencyoutput_sequence_length | Successful/failed records, first-token/end-to-end timing and output length. Retain all attempts, including no-content errors with absent latency. Compute per-class attainment with the agreed denominator. |
| Router · Endpoint Picker | llm_d_epp_flow_control_requests_total | Separate Dispatched, capacity rejection, TTL expiry, no-endpoint and context-cancellation outcomes. Dispatch is not client completion. |
| Router · Endpoint Picker | llm_d_epp_flow_control_capacity_utilization_requestsllm_d_epp_flow_control_capacity_utilization_bytesllm_d_epp_flow_control_global_capacity_utilization_requestsllm_d_epp_flow_control_global_capacity_utilization_bytes | Per-band versus global pending-capacity fractions. Combine with queue_size, queue_bytes and request_queue_duration_seconds; these are not engine-running counts. |
| Router · ceiling-policy trace | Loaded policy parametersactive prioritiescomputed ceiling at dispatch | Join computed ceiling and admission decisions to pool saturation. No dedicated policy-ceiling Prometheus series was verified at this pin; capture diagnostic traces rather than inventing a key. |
| Router · conditional eviction | llm_d_epp_flow_control_revocations_issued_totalllm_d_epp_flow_control_revocations_totalllm_d_epp_flow_control_revocation_confirmation_seconds | Issued revocations, confirmed/timed_out terminal outcomes and confirmation time. Confirm truncated streams, lost work and engine release using matched request traces. |
| Router · conditional eviction | llm_d_epp_flow_control_reclaim_targetllm_d_epp_flow_control_pending_reclaim | Reclamation deficit and pending debits in saturation units. These are internal observations, not configurable pacing knobs. |
| Router · predicted-latency-producer | llm_d_epp_request_predicted_ttft_secondsllm_d_epp_request_predicted_tpot_secondsllm_d_epp_request_slo_violation_total | Predicted seconds and observed violation counts. Compare request-level predictions with actual client outcomes; inspect missing predictions and service errors. A histogram alone does not establish calibration. |
| Router · predicted-latency-producer | llm_d_epp_request_ttft_prediction_duration_secondsllm_d_epp_request_tpot_prediction_duration_seconds | Prediction computation time, not inference latency. Inspect training/prediction service calls and readiness separately. |
| Router · program-aware-fairness | llm_d_epp_program_aware_jains_fairness_indexllm_d_epp_program_aware_avg_wait_time_millisecondsllm_d_epp_program_aware_attained_service_tokens | Policy diagnostics: Jain’s index over average queue waits, milliseconds of wait and weighted attained service. Attained service is LAS-only, updated on completion with lazy decay. It is absent under turn-priority. A flat exported value does not prove accounting has stopped. The index does not measure completed-token shares. Reconcile with per-program successful completions, token shares and oldest pending age. |
| Conditional distributed syncer | Integration-specific peer-update age, publish failures and per-endpoint request/token totals | No stock registered distributed syncer was established. Inspect the actual implementation before naming series; prove convergence and stale-peer recovery. Missing instrumentation is unresolved evidence, not zero errors. |
| Engine · vLLM | vllm:request_prefill_time_secondsvllm:request_decode_time_secondsvllm:num_preemptions_totalvllm:request_num_preemptions | Processing-stage histograms, cumulative engine preemptions, and per-request preemption histogram. Engine preemption is distinct from router cancellation. |
Absent metrics are not zero. Keep errors and requests without content in the attempted-request denominator even when no latency value exists. Return to optional failure controls.
Gateway and optional infrastructure exporters
Identify the installed gateway and CPU/GPU collectors before naming their metrics.
Histograms expose _bucket, _sum and _count. Missing is not zero; one smoke request will not exercise every row. Admission metrics are conditional on the active implementation.