flow-control-benchmarks

Utilization detector calibration

Business question

When should queue depth or KV-cache pressure activate flow control?

Answer. Use queue depth to react to requests waiting inside vLLM and a KV threshold to react to memory pressure; both activation points must be calibrated because earlier intervention can add queueing latency.

Visual summary

Utilization detector calibration tested serving path

Utilization detector calibration benchmark results

Tested configuration

Replay this package with Flow Control Flight Recorder

Queue-depth result

Queue depth 8 produced the lowest p95 TTFT in this short-request closed-loop calibration, with 1.6% lower median throughput than queue depth 5.

Queue depth Repeats Steady RPS p95 TTFT p99 TTFT
5 3 47.6 1,863 ms 2,142 ms
8 3 46.8 1,607 ms 1,864 ms

This does not make 8 the production default. A higher threshold waits for a larger backend queue before activating policy, so the later open-loop scenarios compare earlier thresholds 2 and 5 under bursty, mixed-priority traffic.

KV-cache result

The KV-cache sweep intentionally used 28,672 input tokens, 256 output tokens, and concurrency 96 to create memory pressure. Threshold 0.8 activated policy earlier than 0.75 and had lower p95 TTFT, but neither improved TTFT over the flow-control-off control.

KV setting Repeats p95 TTFT Peak policy queue
Flow control off 3 16,938 ms 0 requests
Threshold 0.8 3 21,559 ms 23-27 requests
Threshold 0.75 3 28,540 ms 28-29 requests

Threshold 0.8 is retained as the primary memory-pressure calibration point. This sweep verifies detector activation and exposes its cost; it does not prove a latency SLO.

Method

Evidence

File Contents
summary.csv One row per retained detector run with explicit units and evidence role.
request-results.csv One sanitized row per request.
traffic-samples.csv Issued, completed, and outstanding requests over time.
system-metrics.csv Curated queue, saturation, vLLM, KV-cache, and preemption metrics.
run-evidence.csv Metrics, routes, headers, cache state, engagement, and proof gates.
analysis.json Matched medians, decision, and claim boundary.

Scope

These closed-loop sweeps calibrate activation points. Customer-facing behavior is established by the open-loop production scenarios, where thresholds are tested with noisy traffic, mixed priorities, and varying request shapes.

Reproduce

These cache-off closed-loop sweeps used the current pipeline/benchmark.py. Queue-depth thresholds were 1, 2, 4, 5, and 8. KV-cache thresholds were 0.50, 0.60, 0.70, 0.75, 0.80, 0.90, and 1.00. Selected points and boundaries ran three times.

OUTPUT_DIR=${OUTPUT_DIR:-results/utilization-detector-calibration}
PROMPT_CACHE_DIR=${PROMPT_CACHE_DIR:?Set the generated prompt-cache directory}
SCENARIO_FILE=${SCENARIO_FILE:-benchmark-data/upstream-flow-control-v0.9.0/utilization-detector-calibration/queue-depth-scenario.json}
SCENARIO_FILTER=${SCENARIO_FILTER:-utilization_queue_depth_calibration}

pipeline/run-in-cluster.sh \
  "$OUTPUT_DIR" "$OUTPUT_DIR/live-status.json" "$PROMPT_CACHE_DIR" \
  "$SCENARIO_FILE" -- --scenario-filter "$SCENARIO_FILTER" \
  --prompt-pool-size 24 --warmup-duration 30 --warmup-concurrency 2 \
  --steady-state-trim-s 30 --metric-sample-interval-s 0.5 \
  --vllm-prefix-caching off --traffic-seed 42 --arrival-mode closed_loop

For the KV-cache sweep, set SCENARIO_FILE to kv-threshold-scenario.json and SCENARIO_FILTER to utilization_kv_threshold_calibration. Set the tested threshold in the Endpoint Picker before each point. Both exact traffic definitions are published beside this README; run-config.json records the images and metric cadence.