flow-control-benchmarks

Long stability

Business question

Does the selected priority-protection configuration recover after repeated production-shaped surges?

Answer. The queue drained after both surges, realtime latency returned to its earlier range, and all 14,889 requests completed.

Visual summary

Long stability tested serving path

Long stability benchmark results

Tested configuration

Replay this package with Flow Control Flight Recorder

Result

All 14,889 requests completed during the 30-minute mixed-workload run. Flow control engaged in both surges, and the policy queue returned to zero afterward.

Premium p95 TTFT reached 1,760 ms in the first surge and 1,227 ms in the second. It returned to 127 ms and 279 ms in the two recovery windows and ended at 290 ms. Both surges remained below the predeclared 3,000 ms test guardrail.

Window Premium p95 TTFT Maximum policy queue
Baseline 299 ms 0 requests
Surge 1 1,760 ms 39 requests
Recovery 1 127 ms 8 requests
Surge 2 1,227 ms 27 requests
Recovery 2 279 ms 41 requests
Final 290 ms 0 requests

Peak vLLM waiting was 29 requests, peak KV-cache use was 23.8%, and vLLM recorded no preemptions. The Endpoint Picker did not restart.

Method

The complete traffic and system configuration is in run-config.json.

Evidence

File Contents
summary.csv Overall latency, throughput, and status results by workload.
window-summary.csv Latency, throughput, and status results for each surge and recovery window.
request-results.csv One sanitized row per request with TTFT, TPOT, latency, token counts, and status.
traffic-samples.csv Issued, completed, and outstanding requests by workload over time.
system-metrics.csv Curated Endpoint Picker and vLLM metrics throughout the run.
run-evidence.csv Schedule, header, route, metric, cache, flow-control, and data-quality gates.
analysis.json Stability decision, window results, and claim boundary.

Scope

This is one 30-minute, single-model, cache-off run. It supports recovery after repeated surges but does not promise a fixed TTFT for every load.

Reproduce

This package used GuideLLM 0.7.0 for one 1,800-second run. The configuration was request-count admission at 128 requests, 10% headroom, random routing, one model replica, four GuideLLM workers per tenant, and cache off.

python3 pipeline/guidellm_trace.py --scenario-file benchmark-data/upstream-flow-control-v0.9.0/long-stability/scenario.json --scenario mixed_workload_long_stability --out-dir /tmp/long-stability --traffic-seed 20260809
python3 pipeline/run_guidellm_scenario.py --manifest /tmp/long-stability/manifest.json --run-dir results/long-stability --prefix long-stability --namespace "${NAMESPACE:-flow-control}" --runner-pod "${RUNNER_POD:-flow-control-benchmark-runner}" --expected-detector concurrency-detector --expected-concurrency-mode requests --expected-max-concurrency 128 --expected-headroom 0.10 --expected-picker random-picker --expected-prefix-cache off --expected-model-replicas 1 --http-version 1 --guidellm-worker-processes 4 --drain-after-done --drain-timeout-s 600 --recover-multiline-sse

scenario.json contains both surges and the recovery windows. run-config.json records the full service configuration.