flow-control-benchmarks

Model pool scaling

Business question

Can a model pool add replicas without losing per-GPU efficiency or worsening premium latency under the same offered load per GPU?

Answer. In the tested range, served throughput per GPU stayed within 0.6% from one to four replicas, and four replicas served all offered requests.

Visual summary

Model pool scaling tested serving path

Model pool scaling benchmark results

Tested configuration

Replay this package with Flow Control Flight Recorder

Result

Scaling from one to four model replicas kept served throughput per GPU within 0.6% and did not worsen median premium burst p95 TTFT. Four replicas served four times the offered traffic with no HTTP 429 responses in the three tested runs.

The smaller pools exposed a capacity boundary: five realtime requests received HTTP 429 responses at one replica, and one standard long-context request received an HTTP 429 response at two replicas. The data supports the scale-out behavior; it does not prove rejection-free service at every pool size.

Model replicas Offered requests Success Median premium burst p95 TTFT Served requests/s/GPU
1 3,498 99.86% 374 ms 4.293
2 7,014 99.99% 340 ms 4.317
4 13,950 100% 322 ms 4.302

Flow control engaged in all nine runs. Prefix caching was disabled, cache counters remained zero, and vLLM recorded no preemptions. The Endpoint Picker did not restart.

Method

The complete configuration is in run-config.json.

Evidence

File Contents
summary.csv Run-level latency, throughput, status, queue, KV cache, and restart results.
request-results.csv One sanitized row per request with TTFT, TPOT, latency, token counts, and status.
traffic-samples.csv Issued, completed, and outstanding requests by tenant over time.
system-metrics.csv Curated Endpoint Picker and per-model vLLM metrics collected during every run.
run-evidence.csv Schedule, header, route, metric, cache, flow-control, and data-quality gates.
analysis.json Per-run values, topology medians, decision limits, and claim boundary.

Scope

This package tests one Endpoint Picker with one, two, or four model replicas. It does not test multiple Endpoint Picker replicas, set an absolute service-level objective, or show rejection-free service at every pool size.

Reproduce

This package used GuideLLM 0.7.0 with one Endpoint Picker and one, two, or four model replicas. Offered load scaled with the replica count. Each topology ran three times with seed 18, exact input-token admission at 20,000 tokens per replica, 25% headroom, random routing, and cache off.

SCENARIO=${SCENARIO:?Set one_replica_scaled_load, two_replica_scaled_load, or four_replica_scaled_load}
MODEL_REPLICAS=${MODEL_REPLICAS:?Set 1, 2, or 4 to match the deployed pool}
WORKERS=$((8 * MODEL_REPLICAS))

python3 pipeline/guidellm_trace.py \
  --scenario-file benchmark-data/upstream-flow-control-v0.9.0/multi-replica-scaling/scenario.json \
  --scenario "$SCENARIO" --out-dir "/tmp/$SCENARIO" --traffic-seed 18
python3 pipeline/run_guidellm_scenario.py \
  --manifest "/tmp/$SCENARIO/manifest.json" \
  --run-dir "results/model-pool-scale/$SCENARIO" --prefix "$SCENARIO" \
  --namespace "${NAMESPACE:-flow-control}" \
  --runner-pod "${RUNNER_POD:-flow-control-benchmark-runner}" \
  --expected-detector concurrency-detector --expected-concurrency-mode tokens \
  --expected-max-concurrency 1000000 --expected-max-token-concurrency 20000 \
  --expected-add-estimated-output-tokens false --expected-headroom 0.25 \
  --expected-picker random-picker --expected-prefix-cache off \
  --expected-model-replicas "$MODEL_REPLICAS" --http-version 1 \
  --guidellm-worker-processes "$WORKERS" --guidellm-mp-poll-interval-s 0.01 \
  --drain-after-done --drain-timeout-s 300 --recover-multiline-sse

scenario.json defines the matched per-GPU traffic. run-config.json defines the topology matrix.