flow-control-benchmarks

Batch isolation

Business question

Can realtime and standard work retain lower TTFT while batch traffic shares the same model server?

Answer. realtime and Standard recorded lower p95 TTFT than Batch in every repeat, but their repeat spreads were too wide for stable point estimates.

Visual summary

Batch isolation tested serving path

Batch isolation benchmark results

Tested configuration

Replay this package with Flow Control Flight Recorder

What the benchmark showed

Workload Median surge p95 TTFT
realtime 442 ms
Standard 515 ms
Batch 13,077 ms

realtime and standard traffic retained lower TTFT while batch absorbed more of the queue. Every request succeeded, flow control engaged during each run, and prefix caching remained off.

The ordering was consistent, but realtime p95 TTFT ranged from 371 to 669 ms (1.81× spread) and standard ranged from 436 to 1,017 ms (2.33×). Both exceed the 1.5× repeat-stability gate. The medians remain visible as directional evidence and are not used as stable headline estimates.

Run inventory

Evidence

File Contents
summary.csv Per-run outcomes, throughput, TTFT, end-to-end latency, and TPOT.
window-summary.csv Baseline, surge, and recovery metrics.
request-results.csv One sanitized row per request.
traffic-samples.csv Issued, completed, and outstanding requests over time.
system-metrics.csv Queue, saturation, vLLM, KV-cache, preemption, and cache metrics.
run-evidence.csv Headers, route counts, cache state, flow-control engagement, and proof gates.
run-config.json Images, topology, engine settings, detector settings, and traffic method.
analysis.json Medians, ranges, run inventory, and claim boundary.

The queue-depth-2 result is a single-run calibration. The package does not use it as a matched detector comparison.

Reproduce

This scenario used GuideLLM 0.7.0 with request-count admission at 128 requests, 15% headroom, random routing, one model replica, and cache off. Three selected repeats used the same deterministic traffic schedule.

python3 pipeline/guidellm_trace.py --scenario-file benchmark-data/upstream-flow-control-v0.9.0/production-scenarios/batch-isolation/scenario.json --scenario batch_isolation --out-dir /tmp/batch-isolation --traffic-seed 42
python3 pipeline/run_guidellm_scenario.py --manifest /tmp/batch-isolation/manifest.json --run-dir results/batch-isolation --prefix batch-isolation --namespace "${NAMESPACE:-flow-control}" --runner-pod "${RUNNER_POD:-flow-control-benchmark-runner}" --expected-detector concurrency-detector --expected-concurrency-mode requests --expected-max-concurrency 128 --expected-headroom 0.15 --expected-picker random-picker --expected-prefix-cache off --expected-model-replicas 1 --http-version 1 --guidellm-worker-processes 4 --drain-after-done --drain-timeout-s 300 --recover-multiline-sse

scenario.json contains only the batch-isolation traffic. run-config.json records the tested images and settings.