Which admission method better protects realtime traffic when chat, agentic, long-context, and batch work share one model server?
Answer. Request-count admission produced lower realtime latency, while input-token admission produced lower long-context and batch latency.
Replay this package with Flow Control Flight Recorder
Request-count admission gave realtime chat the lower surge latency. Input-token admission spread latency more evenly across the four workloads.
| Admission method | realtime p95 TTFT | Agentic p95 TTFT | Long-context p95 TTFT | Batch p95 TTFT | Peak vLLM waiting |
|---|---|---|---|---|---|
| Request count: 128 with 10% headroom | 1,994 ms | 2,745 ms | 5,150 ms | 8,654 ms | 16 requests |
| Input tokens: 75,000 | 2,914 ms | 2,836 ms | 3,076 ms | 2,832 ms | 43 requests |
Request-count admission lowered realtime p95 TTFT by 919 ms and moved more of the wait to lower-priority work. Input-token admission reduced long-context and batch latency, while more requests waited inside vLLM.
All 9,132 requests succeeded. Flow control engaged in all six runs. Prefix caching was disabled, cache counters remained zero, and vLLM recorded no preemptions.
max-num-seqs=128, max-num-batched-tokens=8192, and a 32,768-token model limit.The complete configuration is in run-config.json.
| File | Contents |
|---|---|
summary.csv |
Per-run TTFT, TPOT, queue, KV cache, and detector results. |
window-summary.csv |
Baseline, surge, and recovery results by workload. |
request-results.csv |
One sanitized row per request with timing, tokens, and status. |
traffic-samples.csv |
Issued, completed, and outstanding requests over time. |
system-metrics.csv |
Endpoint Picker and vLLM metrics collected during every run. |
run-evidence.csv |
Schedule, header, route, cache, flow-control, and data-quality checks. |
analysis.json |
Per-run values, medians, ranges, and claim boundary. |
The configuration choice depends on workload shape and business priority. These latency values apply to the tested model, GPU, traffic, and single-model setup; they do not define a general service-level objective.
This package used GuideLLM 0.7.0, one model replica, random routing, and cache off. Three matched repeats compared request-count admission at 128 requests and 10% headroom with input-token admission at 75,000 tokens and no headroom.
python3 pipeline/guidellm_trace.py \
--scenario-file benchmark-data/upstream-flow-control-v0.9.0/mixed-production-workload/scenario.json \
--scenario mixed_production_request_shapes --out-dir /tmp/mixed-production \
--traffic-seed 42
for REPEAT in 1 2 3; do
python3 pipeline/run_guidellm_scenario.py \
--manifest /tmp/mixed-production/manifest.json \
--run-dir "results/mixed-production/request-count/repeat-$REPEAT" \
--prefix "mixed-production-request-count-repeat-$REPEAT" \
--namespace "${NAMESPACE:-flow-control}" \
--runner-pod "${RUNNER_POD:-flow-control-benchmark-runner}" \
--expected-detector concurrency-detector \
--expected-concurrency-mode requests --expected-max-concurrency 128 \
--expected-add-estimated-output-tokens false --expected-headroom 0.10 \
--expected-picker random-picker --expected-prefix-cache off \
--expected-model-replicas 1 --http-version 1 \
--guidellm-worker-processes 4 --drain-after-done \
--drain-timeout-s 300 --recover-multiline-sse
python3 pipeline/run_guidellm_scenario.py \
--manifest /tmp/mixed-production/manifest.json \
--run-dir "results/mixed-production/input-token/repeat-$REPEAT" \
--prefix "mixed-production-input-token-repeat-$REPEAT" \
--namespace "${NAMESPACE:-flow-control}" \
--runner-pod "${RUNNER_POD:-flow-control-benchmark-runner}" \
--expected-detector concurrency-detector \
--expected-concurrency-mode tokens --expected-max-token-concurrency 75000 \
--expected-add-estimated-output-tokens false --expected-headroom 0.00 \
--expected-picker random-picker --expected-prefix-cache off \
--expected-model-replicas 1 --http-version 1 \
--guidellm-worker-processes 4 --drain-after-done \
--drain-timeout-s 300 --recover-multiline-sse
done
The four tenant shapes and rates are in scenario.json. The two exact admission configurations are in run-config.json.