Metrics by component

Scrape each producer pod. Retain namespace, pod UID and engine identity. These are reference names—match the deployed version.

Endpoint Picker

Scrape each Endpoint Picker pod’s verified metrics URL. Connect to the pod · Save a scrape.

Metric / fieldPlain EnglishWhy neededWhat it shows
llm_d_epp_infoBuild identityMatch behavior to the running build.Commit/build reference; confirm image digest if unknown.
llm_d_epp_request_totalRequests scheduledConfirm traffic passed through routing.Count by model, fairness ID and priority; not client successes.
llm_d_epp_request_error_totalRouter-observed errorsRetain failures in the comparison.Counts by workload and error code.
llm_d_epp_request_runningRequests counted in flightCheck router request accounting.Current count; separate from engine-running requests.
llm_d_epp_request_ttft_secondsRouter first-response delayCompare router timing with client timing.Seconds to first response byte; histogram.
llm_d_epp_request_streaming_tpot_secondsRouter average decode time per tokenCompare streaming cadence across matched runs.Per-request (duration − TTFT) / (output tokens − 1), in seconds. This histogram contains request averages, not individual chunk gaps.
llm_d_epp_request_duration_secondsRouter request durationLocate delay across the request path.Router-observed duration in seconds; histogram.
llm_d_epp_scheduler_attempts_totalScheduling attemptsIdentify the selected backend.Status and endpoint identity; attempts are not completions.
llm_d_epp_flow_control_requests_totalAdmission outcomesVerify the admission path was used.Dispatch/rejection outcomes by priority and pool.
llm_d_epp_flow_control_queue_sizeRequests waiting in admissionDetect waiting before model dispatch.Queued request count by workload.
llm_d_epp_flow_control_queue_bytesQueued request bytesCheck the configured byte budget.Sum of queued request sizes. This is not total process memory.
llm_d_epp_flow_control_request_queue_duration_secondsTime in admissionMeasure the cost of waiting.Seconds from enqueue to outcome; histogram.
llm_d_epp_flow_control_pool_saturationDetector saturation signalCompare the signal with the loaded ceiling.Pool saturation ratio at the stated router pin. Compare with the loaded policy ceiling.
llm_d_epp_per_endpoint_queue_sizeRouter view of engine waitingCompare router input with engine observations.Waiting requests per mapped endpoint.
llm_d_epp_datalayer_poll_errors_totalMetric-source polling failuresDetect missing inputs to the router.Error count by source type.
llm_d_epp_datalayer_extract_errors_totalMetric extraction failuresDetect unusable collected inputs.Error count by source/extractor type.
llm_d_epp_flow_control_stale_endpointsMissing or old utilization samplesSeparate stale input from actual pressure.Count from the latest detector evaluation; not all-stage proof.

vLLM

Scrape each engine’s metrics endpoint; retain pod and engine identity. Connect to the pod · Save a scrape.

Metric / fieldPlain EnglishWhy neededWhat it shows
vllm:request_success_totalFinished requests by reasonReconcile engine outcomes with the client.Counts include abort/error reasons; check normal completions separately.
vllm:generation_tokens_totalGenerated tokensConfirm work reached the engine.Cumulative generated-token count.
vllm:num_requests_runningEngine-running requestsObserve active model work.Current requests in execution batches.
vllm:num_requests_waitingEngine-waiting requestsDistinguish engine waiting from admission waiting.Current engine queue length.
vllm:kv_cache_usage_percKV-cache occupancyObserve engine cache pressure.Fraction from 0 to 1; not total GPU memory.
vllm:time_to_first_token_secondsEngine first-token delayCompare engine timing with client timing.Seconds to first token; histogram.
vllm:request_time_per_output_token_secondsEngine average decode time per tokenCompare engine timing with the client.Per-request average time per output token, in seconds. Keep separate from the engine inter-token gap histogram.
vllm:e2e_request_latency_secondsEngine request durationMeasure server-side latency.Total server latency in seconds; histogram.
vllm:request_queue_time_secondsEngine queue timeLocate waiting inside the server.Seconds waiting in the engine; histogram.
vllm:prefix_cache_queries_totalTokens checked for prefix reuseInterpret cache-hit counts.Cumulative queried-token count.
vllm:prefix_cache_hits_totalTokens found in prefix cacheTrack cache differences between runs.Hit-token count; compare with queries over the same window.

AIPerf

Read the client results in the AIPerf run’s output directory. Run and export command.

Metric / fieldPlain EnglishWhy neededWhat it shows
request_countSuccessful client recordsConfirm completed client work.Successful-record count in native exports.
error_request_countFailed client recordsPrevent failed requests disappearing from results.Error-record count; retain error details.
time_to_first_tokenClient first-content delayMeasure what the caller experiences.Per-request first-token timing; read exported units.
inter_token_latencyClient average decode time per tokenEvaluate the streaming token-latency criterion.Per-request (request latency − TTFT) / tokens after the first content chunk. With unavailable or invalid first-chunk usage, the divisor is output tokens − 1. Requires at least two output tokens. Read exported units. This is not an individual-gap distribution.
request_latencyClient request durationMeasure end-to-end user latency.Per-request duration; read exported units.
input_sequence_lengthActual prompt lengthCompare input-size distributions across runs.Input-token count under the runner’s tokenizer.
output_sequence_lengthActual response lengthKeep the amount of generated work comparable.Output-token count under the runner’s tokenizer.

Kubernetes API

Read pod state from the Kubernetes API: kubectl --context "$CTX" -n "$NS" get pod "$POD" -o json.

Metric / fieldPlain EnglishWhy neededWhat it shows
status.conditionsPod readinessVerify the test path stays available.Ready condition and transitions.
status.containerStatuses[].restartCountContainer restartsDetect interrupted or reset processes.Restart count per container.
status.containerStatuses[].lastStatePrevious termination stateInvestigate serving faults.Previous termination reason and timing.
status.containerStatuses[].imageIDRunning image identityMatch evidence to the deployed software.Runtime image identifier/digest.

Streaming measurement boundaries. AIPerf reports client timing, the router reports its observed response timing, and vLLM reports engine timing. Compare matched requests and intervals. Do not subtract unrelated percentiles. For individual gaps, the router exposes llm_d_epp_request_streaming_itl_seconds for consecutive response body chunks. These chunks can contain multiple tokens.

Evidence for conditional controls

Collect router series from each verified Endpoint Picker metrics endpoint and engine series from each model server. Preserve request, workload, priority, tenant and endpoint identity in supporting records. Optional series require their producer to be active.

Component / producerExact keys or evidenceUse and limits
Client · AIPerf recordsrequest_count
error_request_count
time_to_first_token
request_latency
output_sequence_length
Successful/failed records, first-token/end-to-end timing and output length. Retain all attempts, including no-content errors with absent latency. Compute per-class attainment with the agreed denominator.
Router · Endpoint Pickerllm_d_epp_flow_control_requests_totalSeparate Dispatched, capacity rejection, TTL expiry, no-endpoint and context-cancellation outcomes. Dispatch is not client completion.
Router · Endpoint Pickerllm_d_epp_flow_control_capacity_utilization_requests
llm_d_epp_flow_control_capacity_utilization_bytes
llm_d_epp_flow_control_global_capacity_utilization_requests
llm_d_epp_flow_control_global_capacity_utilization_bytes
Per-band versus global pending-capacity fractions. Combine with queue_size, queue_bytes and request_queue_duration_seconds; these are not engine-running counts.
Router · ceiling-policy traceLoaded policy parameters
active priorities
computed ceiling at dispatch
Join computed ceiling and admission decisions to pool saturation. No dedicated policy-ceiling Prometheus series was verified at this pin; capture diagnostic traces rather than inventing a key.
Router · conditional evictionllm_d_epp_flow_control_revocations_issued_total
llm_d_epp_flow_control_revocations_total
llm_d_epp_flow_control_revocation_confirmation_seconds
Issued revocations, confirmed/timed_out terminal outcomes and confirmation time. Confirm truncated streams, lost work and engine release using matched request traces.
Router · conditional evictionllm_d_epp_flow_control_reclaim_target
llm_d_epp_flow_control_pending_reclaim
Reclamation deficit and pending debits in saturation units. These are internal observations, not configurable pacing knobs.
Router · predicted-latency-producerllm_d_epp_request_predicted_ttft_seconds
llm_d_epp_request_predicted_tpot_seconds
llm_d_epp_request_slo_violation_total
Predicted seconds and observed violation counts. Compare request-level predictions with actual client outcomes; inspect missing predictions and service errors. A histogram alone does not establish calibration.
Router · predicted-latency-producerllm_d_epp_request_ttft_prediction_duration_seconds
llm_d_epp_request_tpot_prediction_duration_seconds
Prediction computation time, not inference latency. Inspect training/prediction service calls and readiness separately.
Router · program-aware-fairnessllm_d_epp_program_aware_jains_fairness_index
llm_d_epp_program_aware_avg_wait_time_milliseconds

llm_d_epp_program_aware_attained_service_tokens
Policy diagnostics: Jain’s index over average queue waits, milliseconds of wait and weighted attained service. Attained service is LAS-only, updated on completion with lazy decay. It is absent under turn-priority. A flat exported value does not prove accounting has stopped. The index does not measure completed-token shares. Reconcile with per-program successful completions, token shares and oldest pending age.
Conditional distributed syncerIntegration-specific peer-update age, publish failures and per-endpoint request/token totalsNo stock registered distributed syncer was established. Inspect the actual implementation before naming series; prove convergence and stale-peer recovery. Missing instrumentation is unresolved evidence, not zero errors.
Engine · vLLMvllm:request_prefill_time_seconds
vllm:request_decode_time_seconds
vllm:num_preemptions_total
vllm:request_num_preemptions
Processing-stage histograms, cumulative engine preemptions, and per-request preemption histogram. Engine preemption is distinct from router cancellation.

Absent metrics are not zero. Keep errors and requests without content in the attempted-request denominator even when no latency value exists. Return to optional failure controls.

Gateway and optional infrastructure exporters

Identify the installed gateway and CPU/GPU collectors before naming their metrics.

Histograms expose _bucket, _sum and _count. Missing is not zero; one smoke request will not exercise every row. Admission metrics are conditional on the active implementation.