When should queue depth or KV-cache pressure activate flow control?
Answer. Use queue depth to react to requests waiting inside vLLM and a KV threshold to react to memory pressure; both activation points must be calibrated because earlier intervention can add queueing latency.
Replay this package with Flow Control Flight Recorder
Queue depth 8 produced the lowest p95 TTFT in this short-request closed-loop calibration, with 1.6% lower median throughput than queue depth 5.
| Queue depth | Repeats | Steady RPS | p95 TTFT | p99 TTFT |
|---|---|---|---|---|
| 5 | 3 | 47.6 | 1,863 ms | 2,142 ms |
| 8 | 3 | 46.8 | 1,607 ms | 1,864 ms |
This does not make 8 the production default. A higher threshold waits for a larger backend queue before activating policy, so the later open-loop scenarios compare earlier thresholds 2 and 5 under bursty, mixed-priority traffic.
The KV-cache sweep intentionally used 28,672 input tokens, 256 output tokens, and concurrency 96 to create memory pressure. Threshold 0.8 activated policy earlier than 0.75 and had lower p95 TTFT, but neither improved TTFT over the flow-control-off control.
| KV setting | Repeats | p95 TTFT | Peak policy queue |
|---|---|---|---|
| Flow control off | 3 | 16,938 ms | 0 requests |
| Threshold 0.8 | 3 | 21,559 ms | 23-27 requests |
| Threshold 0.75 | 3 | 28,540 ms | 28-29 requests |
Threshold 0.8 is retained as the primary memory-pressure calibration point. This sweep verifies detector activation and exposes its cost; it does not prove a latency SLO.
| File | Contents |
|---|---|
summary.csv |
One row per retained detector run with explicit units and evidence role. |
request-results.csv |
One sanitized row per request. |
traffic-samples.csv |
Issued, completed, and outstanding requests over time. |
system-metrics.csv |
Curated queue, saturation, vLLM, KV-cache, and preemption metrics. |
run-evidence.csv |
Metrics, routes, headers, cache state, engagement, and proof gates. |
analysis.json |
Matched medians, decision, and claim boundary. |
These closed-loop sweeps calibrate activation points. Customer-facing behavior is established by the open-loop production scenarios, where thresholds are tested with noisy traffic, mixed priorities, and varying request shapes.
These cache-off closed-loop sweeps used the current pipeline/benchmark.py. Queue-depth thresholds were 1, 2, 4, 5, and 8. KV-cache thresholds were 0.50, 0.60, 0.70, 0.75, 0.80, 0.90, and 1.00. Selected points and boundaries ran three times.
OUTPUT_DIR=${OUTPUT_DIR:-results/utilization-detector-calibration}
PROMPT_CACHE_DIR=${PROMPT_CACHE_DIR:?Set the generated prompt-cache directory}
SCENARIO_FILE=${SCENARIO_FILE:-benchmark-data/upstream-flow-control-v0.9.0/utilization-detector-calibration/queue-depth-scenario.json}
SCENARIO_FILTER=${SCENARIO_FILTER:-utilization_queue_depth_calibration}
pipeline/run-in-cluster.sh \
"$OUTPUT_DIR" "$OUTPUT_DIR/live-status.json" "$PROMPT_CACHE_DIR" \
"$SCENARIO_FILE" -- --scenario-filter "$SCENARIO_FILTER" \
--prompt-pool-size 24 --warmup-duration 30 --warmup-concurrency 2 \
--steady-state-trim-s 30 --metric-sample-interval-s 0.5 \
--vllm-prefix-caching off --traffic-seed 42 --arrival-mode closed_loop
For the KV-cache sweep, set SCENARIO_FILE to kv-threshold-scenario.json and SCENARIO_FILTER to utilization_kv_threshold_calibration. Set the tested threshold in the Endpoint Picker before each point. Both exact traffic definitions are published beside this README; run-config.json records the images and metric cadence.