Can realtime and standard work retain lower TTFT while batch traffic shares the same model server?
Answer. realtime and Standard recorded lower p95 TTFT than Batch in every repeat, but their repeat spreads were too wide for stable point estimates.
Replay this package with Flow Control Flight Recorder
| Workload | Median surge p95 TTFT |
|---|---|
| realtime | 442 ms |
| Standard | 515 ms |
| Batch | 13,077 ms |
realtime and standard traffic retained lower TTFT while batch absorbed more of the queue. Every request succeeded, flow control engaged during each run, and prefix caching remained off.
The ordering was consistent, but realtime p95 TTFT ranged from 371 to 669 ms (1.81× spread) and standard ranged from 436 to 1,017 ms (2.33×). Both exceed the 1.5× repeat-stability gate. The medians remain visible as directional evidence and are not used as stable headline estimates.
| File | Contents |
|---|---|
summary.csv |
Per-run outcomes, throughput, TTFT, end-to-end latency, and TPOT. |
window-summary.csv |
Baseline, surge, and recovery metrics. |
request-results.csv |
One sanitized row per request. |
traffic-samples.csv |
Issued, completed, and outstanding requests over time. |
system-metrics.csv |
Queue, saturation, vLLM, KV-cache, preemption, and cache metrics. |
run-evidence.csv |
Headers, route counts, cache state, flow-control engagement, and proof gates. |
run-config.json |
Images, topology, engine settings, detector settings, and traffic method. |
analysis.json |
Medians, ranges, run inventory, and claim boundary. |
The queue-depth-2 result is a single-run calibration. The package does not use it as a matched detector comparison.
This scenario used GuideLLM 0.7.0 with request-count admission at 128 requests, 15% headroom, random routing, one model replica, and cache off. Three selected repeats used the same deterministic traffic schedule.
python3 pipeline/guidellm_trace.py --scenario-file benchmark-data/upstream-flow-control-v0.9.0/production-scenarios/batch-isolation/scenario.json --scenario batch_isolation --out-dir /tmp/batch-isolation --traffic-seed 42
python3 pipeline/run_guidellm_scenario.py --manifest /tmp/batch-isolation/manifest.json --run-dir results/batch-isolation --prefix batch-isolation --namespace "${NAMESPACE:-flow-control}" --runner-pod "${RUNNER_POD:-flow-control-benchmark-runner}" --expected-detector concurrency-detector --expected-concurrency-mode requests --expected-max-concurrency 128 --expected-headroom 0.15 --expected-picker random-picker --expected-prefix-cache off --expected-model-replicas 1 --http-version 1 --guidellm-worker-processes 4 --drain-after-done --drain-timeout-s 300 --recover-multiline-sse
scenario.json contains only the batch-isolation traffic. run-config.json records the tested images and settings.