flow-control-benchmarks

Mixed production workload

Business question

Which admission method better protects realtime traffic when chat, agentic, long-context, and batch work share one model server?

Answer. Request-count admission produced lower realtime latency, while input-token admission produced lower long-context and batch latency.

Visual summary

Mixed production workload tested serving path

Mixed production workload benchmark results

Tested configuration

Replay this package with Flow Control Flight Recorder

Result

Request-count admission gave realtime chat the lower surge latency. Input-token admission spread latency more evenly across the four workloads.

Admission method realtime p95 TTFT Agentic p95 TTFT Long-context p95 TTFT Batch p95 TTFT Peak vLLM waiting
Request count: 128 with 10% headroom 1,994 ms 2,745 ms 5,150 ms 8,654 ms 16 requests
Input tokens: 75,000 2,914 ms 2,836 ms 3,076 ms 2,832 ms 43 requests

Request-count admission lowered realtime p95 TTFT by 919 ms and moved more of the wait to lower-priority work. Input-token admission reduced long-context and batch latency, while more requests waited inside vLLM.

All 9,132 requests succeeded. Flow control engaged in all six runs. Prefix caching was disabled, cache counters remained zero, and vLLM recorded no preemptions.

Method

The complete configuration is in run-config.json.

Evidence

File Contents
summary.csv Per-run TTFT, TPOT, queue, KV cache, and detector results.
window-summary.csv Baseline, surge, and recovery results by workload.
request-results.csv One sanitized row per request with timing, tokens, and status.
traffic-samples.csv Issued, completed, and outstanding requests over time.
system-metrics.csv Endpoint Picker and vLLM metrics collected during every run.
run-evidence.csv Schedule, header, route, cache, flow-control, and data-quality checks.
analysis.json Per-run values, medians, ranges, and claim boundary.

Scope

The configuration choice depends on workload shape and business priority. These latency values apply to the tested model, GPU, traffic, and single-model setup; they do not define a general service-level objective.

Reproduce

This package used GuideLLM 0.7.0, one model replica, random routing, and cache off. Three matched repeats compared request-count admission at 128 requests and 10% headroom with input-token admission at 75,000 tokens and no headroom.

python3 pipeline/guidellm_trace.py \
  --scenario-file benchmark-data/upstream-flow-control-v0.9.0/mixed-production-workload/scenario.json \
  --scenario mixed_production_request_shapes --out-dir /tmp/mixed-production \
  --traffic-seed 42

for REPEAT in 1 2 3; do
  python3 pipeline/run_guidellm_scenario.py \
    --manifest /tmp/mixed-production/manifest.json \
    --run-dir "results/mixed-production/request-count/repeat-$REPEAT" \
    --prefix "mixed-production-request-count-repeat-$REPEAT" \
    --namespace "${NAMESPACE:-flow-control}" \
    --runner-pod "${RUNNER_POD:-flow-control-benchmark-runner}" \
    --expected-detector concurrency-detector \
    --expected-concurrency-mode requests --expected-max-concurrency 128 \
    --expected-add-estimated-output-tokens false --expected-headroom 0.10 \
    --expected-picker random-picker --expected-prefix-cache off \
    --expected-model-replicas 1 --http-version 1 \
    --guidellm-worker-processes 4 --drain-after-done \
    --drain-timeout-s 300 --recover-multiline-sse

  python3 pipeline/run_guidellm_scenario.py \
    --manifest /tmp/mixed-production/manifest.json \
    --run-dir "results/mixed-production/input-token/repeat-$REPEAT" \
    --prefix "mixed-production-input-token-repeat-$REPEAT" \
    --namespace "${NAMESPACE:-flow-control}" \
    --runner-pod "${RUNNER_POD:-flow-control-benchmark-runner}" \
    --expected-detector concurrency-detector \
    --expected-concurrency-mode tokens --expected-max-token-concurrency 75000 \
    --expected-add-estimated-output-tokens false --expected-headroom 0.00 \
    --expected-picker random-picker --expected-prefix-cache off \
    --expected-model-replicas 1 --http-version 1 \
    --guidellm-worker-processes 4 --drain-after-done \
    --drain-timeout-s 300 --recover-multiline-sse
done

The four tenant shapes and rates are in scenario.json. The two exact admission configurations are in run-config.json.