Can a model pool add replicas without losing per-GPU efficiency or worsening premium latency under the same offered load per GPU?
Answer. In the tested range, served throughput per GPU stayed within 0.6% from one to four replicas, and four replicas served all offered requests.
Replay this package with Flow Control Flight Recorder
Scaling from one to four model replicas kept served throughput per GPU within 0.6% and did not worsen median premium burst p95 TTFT. Four replicas served four times the offered traffic with no HTTP 429 responses in the three tested runs.
The smaller pools exposed a capacity boundary: five realtime requests received HTTP 429 responses at one replica, and one standard long-context request received an HTTP 429 response at two replicas. The data supports the scale-out behavior; it does not prove rejection-free service at every pool size.
| Model replicas | Offered requests | Success | Median premium burst p95 TTFT | Served requests/s/GPU |
|---|---|---|---|---|
| 1 | 3,498 | 99.86% | 374 ms | 4.293 |
| 2 | 7,014 | 99.99% | 340 ms | 4.317 |
| 4 | 13,950 | 100% | 322 ms | 4.302 |
Flow control engaged in all nine runs. Prefix caching was disabled, cache counters remained zero, and vLLM recorded no preemptions. The Endpoint Picker did not restart.
max-num-seqs=128, max-num-batched-tokens=8192, and a 32,768-token
model limit.The complete configuration is in run-config.json.
| File | Contents |
|---|---|
summary.csv |
Run-level latency, throughput, status, queue, KV cache, and restart results. |
request-results.csv |
One sanitized row per request with TTFT, TPOT, latency, token counts, and status. |
traffic-samples.csv |
Issued, completed, and outstanding requests by tenant over time. |
system-metrics.csv |
Curated Endpoint Picker and per-model vLLM metrics collected during every run. |
run-evidence.csv |
Schedule, header, route, metric, cache, flow-control, and data-quality gates. |
analysis.json |
Per-run values, topology medians, decision limits, and claim boundary. |
This package tests one Endpoint Picker with one, two, or four model replicas. It does not test multiple Endpoint Picker replicas, set an absolute service-level objective, or show rejection-free service at every pool size.
This package used GuideLLM 0.7.0 with one Endpoint Picker and one, two, or four model replicas. Offered load scaled with the replica count. Each topology ran three times with seed 18, exact input-token admission at 20,000 tokens per replica, 25% headroom, random routing, and cache off.
SCENARIO=${SCENARIO:?Set one_replica_scaled_load, two_replica_scaled_load, or four_replica_scaled_load}
MODEL_REPLICAS=${MODEL_REPLICAS:?Set 1, 2, or 4 to match the deployed pool}
WORKERS=$((8 * MODEL_REPLICAS))
python3 pipeline/guidellm_trace.py \
--scenario-file benchmark-data/upstream-flow-control-v0.9.0/multi-replica-scaling/scenario.json \
--scenario "$SCENARIO" --out-dir "/tmp/$SCENARIO" --traffic-seed 18
python3 pipeline/run_guidellm_scenario.py \
--manifest "/tmp/$SCENARIO/manifest.json" \
--run-dir "results/model-pool-scale/$SCENARIO" --prefix "$SCENARIO" \
--namespace "${NAMESPACE:-flow-control}" \
--runner-pod "${RUNNER_POD:-flow-control-benchmark-runner}" \
--expected-detector concurrency-detector --expected-concurrency-mode tokens \
--expected-max-concurrency 1000000 --expected-max-token-concurrency 20000 \
--expected-add-estimated-output-tokens false --expected-headroom 0.25 \
--expected-picker random-picker --expected-prefix-cache off \
--expected-model-replicas "$MODEL_REPLICAS" --http-version 1 \
--guidellm-worker-processes "$WORKERS" --guidellm-mp-poll-interval-s 0.01 \
--drain-after-done --drain-timeout-s 300 --recover-multiline-sse
scenario.json defines the matched per-GPU traffic. run-config.json defines the topology matrix.