Verify resources and configuration

Lead the first test

  1. Identify the deployed path. Confirm versions and route from the gateway through the router to the model servers.
  2. Prove collection first. Get a successful metrics response from each producer pod. Resolve transport and authentication separately.
  3. Send one bounded request. Save the response and client timing, then reconcile before/after server evidence.
  4. Decide whether to proceed. Start the baseline only when the request, collection and classification checks pass.

Cluster access · HTTP / HTTPS / authentication · Capture metrics · Test request · Troubleshooting · Sources

Use these commands on the operator’s approved workstation. Deployed versions, names, certificates and permissions must be discovered. Local fixture tests do not establish access to your cluster.

1. Establish access and save the deployment identity

Prerequisites: Bash, kubectl, curl and Python 3. Use the installed AIPerf version for client metrics. AWS CLI and the approved AWS profile are needed only for EKS authentication. Start with a fresh private evidence directory.

# Bash. Fill these from the operator's environment, not the example runs.
set +x
umask 077
: "${CTX:?Select the approved kubeconfig context}" "${NS:?Workload namespace}"
: "${EVIDENCE_DIR:?New private directory for this session}"
mkdir "$EVIDENCE_DIR" || exit 1
kubectl config get-contexts
kubectl --context "$CTX" --request-timeout=10s get namespace "$NS"
kubectl --context "$CTX" auth can-i get pods -n "$NS"
kubectl --context "$CTX" auth can-i create pods --subresource=portforward -n "$NS"
kubectl --context "$CTX" version -o yaml > "$EVIDENCE_DIR/kubernetes-version.yaml"
kubectl --context "$CTX" -n "$NS" get pods -o wide

Expected: the intended context and namespace are accessible. Required access checks say yes. A Kubernetes API Unauthorized or timeout must be resolved before debugging model metrics.

EKS: expired login, wrong role, or private API endpoint

The AWS identity used by kubectl must have cluster access. EKS login is separate from gateway API credentials and metrics-reader credentials. If the context is missing, use the operator’s approved kubeconfig setup procedure.

: "${AWS_PROFILE:?Approved profile}" "${AWS_REGION:?Cluster region}" "${CLUSTER:?Cluster name}"
aws --profile "$AWS_PROFILE" sts get-caller-identity
aws --profile "$AWS_PROFILE" --region "$AWS_REGION" eks describe-cluster \
  --name "$CLUSTER" --query 'cluster.{endpoint:endpoint,version:version,access:accessConfig,network:resourcesVpcConfig}'
# If this is an SSO profile and its session expired:
# aws sso login --profile "$AWS_PROFILE"

Unauthorized: confirm the profile/role used by the kubeconfig exec configuration and its EKS access entry or legacy mapping with the cluster owner. Timeout or DNS failure: use the approved VPN, VPC workstation or bastion for a private API. Refreshing a token does not fix missing network access.

EKS kubeconfig setup · EKS troubleshooting

Map each producer before selecting a metrics port
: "${POD:?Exact producer pod}" "${CONTAINER:?Container in that pod}"
kubectl --context "$CTX" -n "$NS" get pod "$POD" -o json > "$EVIDENCE_DIR/$POD.identity.json"
kubectl --context "$CTX" -n "$NS" get pod "$POD" \
  -o jsonpath='{.metadata.uid}{"\n"}{range .spec.containers[*]}{.name}{"\n"}{.image}{"\n"}{.ports}{"\n"}{.args}{"\n"}{end}{range .status.containerStatuses[*]}{.name}{" "}{.imageID}{" restarts="}{.restartCount}{"\n"}{end}'
kubectl --context "$CTX" -n "$NS" logs "$POD" -c "$CONTAINER" \
  --since=15m --tail=200 --timestamps > "$EVIDENCE_DIR/$POD.startup.log"
kubectl --context "$CTX" -n "$NS" get services -o wide
kubectl --context "$CTX" -n "$NS" get endpointslices.discovery.k8s.io -o wide

Read the selected container’s args, listener logs, mounted configuration and service target ports. A declared containerPort is a hint, not proof that a process listens there. Record namespace, pod UID, component, image digest, metrics scheme/port/path and auth method for each router and engine pod. Gateway and GPU exporters have their own version-specific endpoints. Do not infer their metrics from the model API URL.

Use discovered names in Bash. Repeat namespace/pod checks for every component in the test path.

Command values

CTX, NS, DEPLOYMENT, POD, SERVICE, RESOURCE, OBJECT, ROUTER_POD, ROUTER_CONTAINER, CONFIGMAP, LOCAL_PORT, METRICS_PORT come from the actual deployment. Gateway/pool resource types depend on installed APIs. Save outputs privately.

Resource / configurationVerifyCommand
1. Cluster and runnerCorrect context; Kubernetes and AIPerf versions.
kubectl --context "$CTX" version -o yaml
aiperf --version
aiperf profile --help
2. Deployment and podsModel, replicas, images/digests, ports and readiness.
kubectl --context "$CTX" -n "$NS" get deployment "$DEPLOYMENT" -o yaml
kubectl --context "$CTX" -n "$NS" get pods -o wide
kubectl --context "$CTX" -n "$NS" get pod "$POD" -o json
3. Service and endpointsService ports match the intended serving pods.
kubectl --context "$CTX" -n "$NS" get service "$SERVICE" -o yaml
kubectl --context "$CTX" -n "$NS" get endpointslices.discovery.k8s.io -o yaml
4. Gateway, route and poolTrace the API URL to the correct backend and Endpoint Picker. Use installed resource names.
kubectl --context "$CTX" api-resources --verbs=get -o wide
kubectl --context "$CTX" -n "$NS" get "$RESOURCE" "$OBJECT" -o yaml
5. Active router configurationMatch mounted config to process arguments/logs; check admission, detector, filters and request classification.
kubectl --context "$CTX" -n "$NS" get pod "$ROUTER_POD" -o json
kubectl --context "$CTX" -n "$NS" get configmap "$CONFIGMAP" -o yaml
kubectl --context "$CTX" -n "$NS" logs "$ROUTER_POD" -c "$ROUTER_CONTAINER" --timestamps
6. Metric listenerOne verified URL per producer pod, with the right port, TLS and auth.
kubectl --context "$CTX" -n "$NS" port-forward --address 127.0.0.1 "pod/$POD" "$LOCAL_PORT:$METRICS_PORT"

2. Connect to the actual metrics listener

Forward one exact pod at a time. Keep this terminal open. Use a different local port for each simultaneous producer.

: "${POD:?Producer pod}" "${LOCAL_PORT:?Unused local port}" "${METRICS_PORT:?Verified pod listener port}"
kubectl --context "$CTX" -n "$NS" port-forward --address 127.0.0.1 \
  "pod/$POD" "$LOCAL_PORT:$METRICS_PORT"

Expected: Forwarding from 127.0.0.1:…. Port-forward removes the need to reach the pod network from your laptop. It does not bypass endpoint authentication. A pod restart ends the forward.

Version check: router 30f06d9c defaults to metrics port 9090 and endpoint authentication enabled. Metrics TLS is controlled by --metrics-cert-dir and optional --metrics-client-ca-file. --secure-serving describes a different listener. Other builds and file-discovery mode can differ. Use their actual arguments and logs. Listener options · Metrics server wiring.

HTTP, trusted HTTPS, and certificate hostname matching

In terminal B, re-enter the selected context, namespace and local port from terminal A. Shell variables and arrays do not transfer between terminals. Choose one verified transport. Empty arrays are intentional. Do not send bearer tokens to an unidentified endpoint.

set +u
: "${CTX:?Re-enter approved context}" "${NS:?Re-enter namespace}" "${LOCAL_PORT:?Re-enter forwarded local port}"
CURL_ARGS=(--noproxy '*')
# Plain HTTP listener reached only through the loopback port-forward:
METRICS_URL="http://127.0.0.1:$LOCAL_PORT/metrics"
# For HTTPS instead, set the approved server CA and certificate DNS name:
# METRICS_URL="https://$METRICS_DNS:$LOCAL_PORT/metrics"
# CURL_ARGS=(--noproxy '*' --cacert "$METRICS_CA" --resolve "$METRICS_DNS:$LOCAL_PORT:127.0.0.1")
# If the server requires mutual TLS, append the approved client identity:
# CURL_ARGS+=(--cert "$METRICS_CLIENT_CERT" --key "$METRICS_CLIENT_KEY")

Use the actual metrics path if it differs from /metrics. A hostname mismatch through localhost needs the certificate DNS name plus --resolve. An unknown CA needs the approved CA bundle. The Kubernetes API CA is not automatically the metrics CA. Do not make --insecure the collection configuration. For a non-forwarded endpoint, omit the loopback --resolve mapping and the local-only --noproxy override. Use the approved network/proxy path. Each curl command starts with --disable so a personal curlrc cannot silently add retries or redirects.

401 / 403: obtain the right metrics identity

401: missing, expired or unacceptable credentials. 403: the identity is recognized but access is denied. A runner or router service account does not automatically have permission to read metrics. Use the existing approved reader account. Requesting its short-lived token is a credential operation.

set +x
umask 077
: "${READER_NS:?Metrics reader namespace}" "${READER_SA:?Approved metrics reader account}"
: "${TOKEN_HEADER_FILE:?New private header-file path}"
# Ask the cluster owner if TokenRequest or impersonation checks are not allowed.
METRICS_TOKEN=$(kubectl --context "$CTX" -n "$READER_NS" create token "$READER_SA" --duration=15m) || exit 1
( set -o noclobber; printf 'Authorization: Bearer %s\n' "$METRICS_TOKEN" > "$TOKEN_HEADER_FILE" ) || exit 1
unset METRICS_TOKEN
CURL_ARGS+=(--header "@$TOKEN_HEADER_FILE")

The requested duration may be adjusted by the server. Renew an expired token before a long test. Do not print, commit or share the header file. An EKS IAM token or inference API key is not interchangeable with this service-account token.

# Requires permission to impersonate these service accounts. Otherwise ask an admin.
kubectl --context "$CTX" auth can-i get /metrics \
  --as="system:serviceaccount:$READER_NS:$READER_SA"
# Authenticated metrics also needs the serving router's delegated-auth permissions:
: "${ROUTER_NS:?}" "${ROUTER_SA:?Actual router service account}"
kubectl --context "$CTX" auth can-i create tokenreviews.authentication.k8s.io \
  --as="system:serviceaccount:$ROUTER_NS:$ROUTER_SA"
kubectl --context "$CTX" auth can-i create subjectaccessreviews.authorization.k8s.io \
  --as="system:serviceaccount:$ROUTER_NS:$ROUTER_SA"

Expected: yes for the needed checks. An impersonation error means this diagnostic could not run. It does not establish that the target account lacks access. If reader authorization is missing, the administrator should grant get on the non-resource URL /metrics using a ClusterRole and a binding to that specific reader. Do not grant cluster-admin or disable metrics authentication.

TokenRequest · Permission checks

A configured feature is not proof it was loaded. Gate-off or priority 0 does not establish an unconstrained baseline.

7. One-request smoke check

  1. Save a timestamped metric scrape from every mapped producer before the request.
  2. Run one request through the intended gateway; retain success or failure.
  3. Save post-request scrapes and reconcile client, router, engine and pod health.
Collection command — repeat per pod before and after

Use a verified URL and a new filename containing namespace, pod and phase. CURL_ARGS holds approved CA/auth options; an empty array is valid for a verified no-auth endpoint.

# Bash. Capture status and body separately. Do not enable shell tracing.
set +u
umask 077
: "${METRICS_URL:?}" "${OUT:?New private filename prefix including pod and phase}"
declare -p CURL_ARGS >/dev/null || exit 1
( set -o noclobber; : > "$OUT.lock" ) || exit 1
date -u +%Y-%m-%dT%H:%M:%SZ > "$OUT.start.utc"
SCRAPE_STATUS=0
curl --disable "${CURL_ARGS[@]}" --silent --show-error --connect-timeout 5 --max-time 15 \
  --output "$OUT.prom" --write-out '%{http_code}\n' "$METRICS_URL" \
  > "$OUT.http-status" 2> "$OUT.error" || SCRAPE_STATUS=$?
printf '%s\n' "$SCRAPE_STATUS" > "$OUT.exit-status"
date -u +%Y-%m-%dT%H:%M:%SZ > "$OUT.end.utc"
python3 - "$OUT" <<'CHECK'
from pathlib import Path
import re,sys
p=sys.argv[1]
status=Path(p+'.http-status').read_text().strip()
rc=Path(p+'.exit-status').read_text().strip()
if status!='200' or rc!='0':
    raise SystemExit('SCRAPE FAILED: HTTP '+status+', curl '+rc+'. Read the saved error and troubleshooting table.')
text=Path(p+'.prom').read_text()
names=sorted(set(re.findall(r'^([a-zA-Z_:][a-zA-Z0-9_:]*)(?:\{[^\n]*\})?\s+(?:[-+]?\d|NaN|[-+]?Inf)',text,re.M)))
if not names or text.lstrip().startswith('<'):
    raise SystemExit('SCRAPE FAILED: expected metric samples, not a login/HTML/empty response.')
Path(p+'.names.txt').write_text('\n'.join(names)+'\n')
print('SCRAPE OK:',len(names),'sample names. Now check required producer names and labels.')
CHECK
AIPerf command

Use the served model, a local tokenizer, small agreed input/output sizes, a timeout and a new output directory. AIPERF_AUTH_ARGS contains only approved auth/classification options. Initialize AIPERF_AUTH_ARGS=() only for a verified endpoint needing none. Otherwise use the installed runner’s documented credential mechanism and --header options for the confirmed request mapping. Verify no extra warmup, repeated profiles or gateway retries.

# Traffic-generating; run after the setup checks. Bash.
set +u
: "${GATEWAY_URL:?}" "${MODEL:?}" "${TOKENIZER:?}" "${ISL:?}" "${OSL:?}" "${TIMEOUT:?}" "${RUN_DIR:?New private directory}"
declare -p AIPERF_AUTH_ARGS >/dev/null || exit 1
umask 077
mkdir "$RUN_DIR" || exit 1
SMOKE_STATUS=0
date -u +%Y-%m-%dT%H:%M:%SZ > "$RUN_DIR/probe-start.utc"
aiperf profile --url "$GATEWAY_URL" --model "$MODEL" --tokenizer "$TOKENIZER"   --endpoint-type chat --streaming --request-count 1 --concurrency 1   --isl "$ISL" --osl "$OSL" --request-timeout-seconds "$TIMEOUT"   --no-server-metrics "${AIPERF_AUTH_ARGS[@]}"   --output-artifact-dir "$RUN_DIR/smoke" || SMOKE_STATUS=$?
printf '%s\n' "$SMOKE_STATUS" > "$RUN_DIR/probe-exit-status.txt"
date -u +%Y-%m-%dT%H:%M:%SZ > "$RUN_DIR/probe-end.utc"
# Take the post-request scrapes even when SMOKE_STATUS is nonzero.

Per-pod scrapes supply the server evidence here. AIPerf auto-discovery is disabled. If using Prometheus, verify its ingestion too.

Pass: the request completed normally, required scrapes succeeded, intended workload classification is confirmed, and pod health is unchanged. Reconcile router/engine counter or histogram-count changes over the same interval. In a shared cluster, background traffic prevents an exact +1 assertion without matching request traces. Explain extra requests, failures, missing series or resets before a sweep.

One request cannot prove saturation, fairness, all-replica behavior or every metric. Short-lived gauges may not change between scrapes.

Alternative: one curl request to isolate gateway/API failures

Use this when AIPerf setup is the uncertain layer. Choose this or the AIPerf smoke test first. Running both sends two requests. This probe expects a streaming OpenAI-compatible text chat response. Confirm the served model, full API path and supported token-limit field with the operator.

GATEWAY_CURL_ARGS=() is valid only for a verified unauthenticated endpoint. Otherwise set the approved gateway CA and authentication options, using a private header file as in the metrics example. Gateway credentials are separate from metrics credentials. Add the intended objective and fairness headers only after confirming their configured mapping.

# Sends exactly one client request. Use the full approved chat-completions URL.
set +u
set +x
umask 077
: "${CHAT_URL:?Full supported API URL}" "${MODEL:?Served model name}"
: "${TOKEN_LIMIT_FIELD:?max_tokens or max_completion_tokens, as supported}"
: "${OUTPUT_LIMIT:?Approved positive token budget}" "${TIMEOUT:?Approved seconds}"
: "${SMOKE_DIR:?New private directory}"
declare -p GATEWAY_CURL_ARGS >/dev/null || exit 1
mkdir "$SMOKE_DIR" || exit 1
python3 - "$MODEL" "$TOKEN_LIMIT_FIELD" "$OUTPUT_LIMIT" "$SMOKE_DIR/request.json" <<'REQUEST'
import json,sys
model,key,limit,path=sys.argv[1:]
if key not in ('max_tokens','max_completion_tokens') or int(limit)<1:
    raise SystemExit('Choose the supported output-limit field and a positive budget.')
with open(path,'w') as f:
    json.dump({'model':model,'messages':[{'role':'user','content':'Reply with the single word ready.'}],
               'stream':True,key:int(limit)},f)
REQUEST
if [ "$?" -ne 0 ]; then exit 1; fi
date -u +%Y-%m-%dT%H:%M:%SZ > "$SMOKE_DIR/start.utc"
REQUEST_STATUS=0
curl --disable "${GATEWAY_CURL_ARGS[@]}" --silent --show-error --no-buffer \
  --connect-timeout 5 --max-time "$TIMEOUT" \
  --header 'Content-Type: application/json' --data-binary "@$SMOKE_DIR/request.json" \
  --output "$SMOKE_DIR/response.sse" --write-out '%{http_code}\n%{time_starttransfer}\n%{time_total}\n' \
  "$CHAT_URL" > "$SMOKE_DIR/transport.txt" 2> "$SMOKE_DIR/curl.error" || REQUEST_STATUS=$?
printf '%s\n' "$REQUEST_STATUS" > "$SMOKE_DIR/curl.status"
date -u +%Y-%m-%dT%H:%M:%SZ > "$SMOKE_DIR/end.utc"
# Always collect post-request metrics, even if this validation fails.
python3 - "$SMOKE_DIR" <<'RESPONSE'
from pathlib import Path
import json,sys
p=Path(sys.argv[1])
status=(p/'transport.txt').read_text().splitlines()
if (p/'curl.status').read_text().strip()!='0' or not status or status[0]!='200':
    raise SystemExit('REQUEST FAILED: inspect saved HTTP status, error and response body.')
content=[];reasons=[];errors=[]
for line in (p/'response.sse').read_text().splitlines():
    if not line.startswith('data:'):continue
    data=line[5:].strip()
    if not data or data=='[DONE]':continue
    try:item=json.loads(data)
    except json.JSONDecodeError:
        errors.append('Malformed SSE JSON');continue
    if not isinstance(item,dict):
        errors.append('Unexpected event shape');continue
    if item.get('error'):errors.append('API error event')
    for choice in item.get('choices',[]):
        text=choice.get('delta',{}).get('content')
        if isinstance(text,str):content.append(text)
        if choice.get('finish_reason'):reasons.append(choice['finish_reason'])
result={'http_status':status[0],'has_visible_content':bool(''.join(content).strip()),
        'finish_reasons':reasons,'errors':errors}
(p/'validation.json').write_text(json.dumps(result,indent=2))
if errors or not result['has_visible_content'] or not reasons or any(r!='stop' for r in reasons):
    raise SystemExit('NOT READY: incomplete, non-text, rejected or token-limited response. Inspect validation.json before another request.')
print('REQUEST OK: visible content and normal completion. Reconcile server metrics and pod health next.')
RESPONSE

API smoke pass: HTTP200, visible content, a normal completion and no error event. A token-limited or reasoning-only response is not a complete text smoke result. Agree a sufficient bounded output budget before retrying. time_starttransfer is first response byte, not client TTFT. Use AIPerf for benchmark timing and token metrics. Gateway retry policies may cause additional backend attempts even though curl sends one request.

Confirm this request received the intended priority and tenant identity
  1. Before sending: record the selected InferenceObjective, its configured priority and a unique approved test fairness identity. Use the same request headers in the client and gateway.
  2. Send: x-llm-d-inference-objective and x-llm-d-inference-fairness-id using their verified values. Check gateway propagation. Keep credentials in the private gateway header file.
  3. Observe: correlate the request ID/time with router traces or logs showing the resolved priority and fairness identity. Where emitted, cross-check the matching priority and fairness_id labels in router request/flow-control series. Label cardinality limits can aggregate identities.
  4. Decide: matching values pass classification. A mismatch requires fixing the mapping/headers. Missing trace or label evidence leaves classification unverified. A request with omitted headers establishes API-path success only unless the fallback assignment was explicitly checked.

Workload classification decisions and YAML. Do not begin a priority or fairness comparison using unverified assignments.

Collect the evidence each decision needs

  1. Every producer: one verified URL and raw scrape per router and engine pod. Preserve all exposed samples and labels, not only a few grep matches. The metric catalog identifies required versus conditional producers.
  2. Before / during / after: capture timestamps and pod UID/restarts. For a load test, use an approved repeated interval through the load and drain periods. Two snapshots cannot recover transient queue gauges or accurate tail latency.
  3. Client: retain all AIPerf attempts, errors and per-request timing/token records. Router dispatches and engine finishes are not substitutes for client successes.
  4. Unavailable: absent is not zero. A single request will not expose saturation, eviction, fairness or every conditional series. Record which conclusion is blocked.
Prometheus exists: verify ingestion and export a time window

Use the approved Prometheus URL and its own CA/auth settings in PROM_CURL_ARGS. Initialize PROM_CURL_ARGS=() only when that endpoint needs no additional CA/auth options. Otherwise populate it with the approved options as for direct scrapes. Start/end are UTC RFC3339 timestamps or Unix seconds. Select the actual discovered job, instance and pod labels. ServiceMonitor/PodMonitor resources exist only when their CRDs are installed.

set +u
: "${PROM_URL:?Verified Prometheus base URL}" "${PROM_OUT:?New private output prefix}"
: "${START:?UTC window start}" "${END:?UTC window end}" "${STEP:?Chosen query step}"
: "${PROMQL:?Expression with the actual producer labels}"
declare -p PROM_CURL_ARGS >/dev/null || exit 1
curl --disable "${PROM_CURL_ARGS[@]}" --fail --silent --show-error --max-time 30 \
  "$PROM_URL/api/v1/targets?state=active" > "$PROM_OUT.targets.json"
curl --disable "${PROM_CURL_ARGS[@]}" --fail --silent --show-error --max-time 30 --get \
  "$PROM_URL/api/v1/query_range" --data-urlencode "query=$PROMQL" \
  --data-urlencode "start=$START" --data-urlencode "end=$END" \
  --data-urlencode "step=$STEP" > "$PROM_OUT.range.json"

Check: API status: success, a nonempty result and fresh sample timestamps. Inspect target health/lastError and query up with the selected labels. Use counter increases over matched windows, account for resets, and retain histogram buckets for percentile calculations. A query step cannot restore samples never scraped. For each required histogram, export its _bucket, _sum and _count series.

Prometheus HTTP API

Direct scrapes work but the router cannot see engine metrics

Direct collection, Prometheus ingestion and router-to-engine polling are three separate paths. At this router pin, inspect metrics-data-source parameters scheme, path, caCertPath, clientCertPath, clientKeyPath and insecureSkipVerify, plus the configured engine address/port and extractor names. A gateway bearer token cannot simply be added to an unsupported plugin field. If the installed source cannot authenticate to the engine listener, that is an integration prerequisite.

Correlate llm_d_epp_datalayer_poll_errors_total, llm_d_epp_datalayer_extract_errors_total, llm_d_epp_flow_control_stale_endpoints and router logs with engine availability. Pinned engine-metrics source.

Stop at the failed layer

SymptomNext checkReady when
kubectl Unauthorized / ForbiddenCheck EKS identity and cluster access. A valid AWS login can still lack Kubernetes permissions.Namespace and required resource permissions pass.
kubectl timeout / DNS failureCheck approved VPN/VPC path, cluster endpoint access settings and DNS.Kubernetes API responds before testing pods.
Port-forward denied / disconnectedCheck create pods/portforward permission, chosen pod readiness and restarts. Restart forward after a pod replacement.Forward stays running for the test window.
curl 7 / connection refusedCheck forwarding terminal, local port and actual listener. Wrong target port is not an auth failure.TCP connection reaches the intended listener.
curl 35 / TLS wrong version / empty replyConfirm HTTP versus HTTPS from listener args/logs. HTTP to TLS may return400 or close the connection.Transport succeeds with the verified scheme.
curl 60 / certificate failureUse approved CA, certificate DNS name and --resolve for a local forward. For mTLS use the approved client cert/key.Certificate verification succeeds.
HTTP401 / HTTP403Check token expiry/audience, /metrics reader RBAC and router delegated TokenReview/SubjectAccessReview access. Read router auth errors.Authorized reader gets200 and metric samples.
HTTP404 / HTML / redirectConfirm port and path belong to metrics. An ingress/login response is not a scrape. Do not forward credentials through unverified redirects.200 body contains the intended producer’s samples.
HTTP200 but required metric absentCompare deployed version with the catalog. Check feature engagement, streaming support and whether an observation has occurred. Histograms use _bucket/_sum/_count.Required names exist or their absence has a documented cause and test limit.
Direct scrape works, Prometheus is emptyCheck Prometheus target discovery, selectors, target scheme/port/auth/TLS, scrape errors and relabeling.Target up=1, fresh samples, expected labels and histogram buckets.
Engine scrape works, router reports stale inputsRouter-to-engine polling is a separate network and credential path. Inspect its loaded data-source scheme/path/TLS settings and poll/extract errors.Router observes fresh, correctly mapped engine input.
Smoke fails with400/404/422Check served model name, full API path, request schema and supported output-limit field.One supported request returns a complete response.
Smoke returns429/503 or times outSave error body. Inspect route/pool health, router admission/TTL/no-endpoint outcomes, engine readiness/queue and logs. Do not increase load.The failure is explained and a corrected single request succeeds.
Kubernetes commands for the failing pod and route
: "${POD:?Failing pod}" "${CONTAINER:?Failing container}"
kubectl --context "$CTX" -n "$NS" describe pod "$POD"
kubectl --context "$CTX" -n "$NS" logs "$POD" -c "$CONTAINER" --since=10m --tail=200 --timestamps
# If restartCount increased, inspect the previous container instance:
# kubectl --context "$CTX" -n "$NS" logs "$POD" -c "$CONTAINER" --previous --tail=200 --timestamps
kubectl --context "$CTX" -n "$NS" get events --field-selector "involvedObject.name=$POD" --sort-by=.lastTimestamp
kubectl --context "$CTX" -n "$NS" get networkpolicies.networking.k8s.io
# After identifying the service and route types/names:
kubectl --context "$CTX" -n "$NS" get service "$SERVICE" -o yaml
kubectl --context "$CTX" -n "$NS" get endpointslices.discovery.k8s.io \
  -l "kubernetes.io/service-name=$SERVICE" -o yaml
kubectl --context "$CTX" -n "$NS" get "$RESOURCE" "$OBJECT" -o yaml

Check route Accepted/ResolvedRefs conditions where supported, service selectors/targetPort, ready endpoint addresses and the intended inference pool. If namespace-scoped permission is missing, ask the owning operator to run the same read-only checks. Do not install a debug pod or change policies during diagnosis without their approval.

Sources and verification scope

These pins describe the guide’s reference behavior. Match the deployed image digest and version before treating any name or default as applicable. The Example tab links recorded run evidence. A local rehearsal cannot certify customer networking, permissions or versions.