Flow control benchmarksBenchmark-Walkthrough
Home All evidence

Measuring flow control
under pressure.

This walkthrough follows the original RHAII 3.4 utilization-detector campaign. The current v0.9.0 benchmark extends it with admission sweeps, production-shaped traffic, mixed workloads, scaling, routing, and stability tests.

Request path

Each tenant tags its requests with a tier and a tenant id, flow control orders them by priority, and one shared vLLM pool serves the result.

TENANTSpremiumpriority 100standardpriority 0batchpriority -10tier header x-llm-d-inference-objectivetenant header x-llm-d-inference-fairness-idENDPOINT PICKERflow controlpriority-aware admission queue10000-10admits by priority while the pool has room;queues instead of rejecting when it does notunmatched objective → priority 0;matched priority is independentof current pool pressureSHARED POOLvLLM, 1x H100, TP 1openai/gpt-oss-20bmax-num-seqs 128max-num-batched-tokens 8192max-model-len 32768gpu-memory-utilization 0.90prefix caching off
Three tenants share one H100. Flow control admits by priority and queues rather than rejects when the pool is full.

How priority is decided

The Endpoint Picker resolves the priority before admission. Queued requests are ordered by policy, and the saturation check determines whether they can dispatch.

1 · REQUEST ARRIVES x-llm-d-inference- objective its tier x-llm-d-inference- fairness-id 2 · RESOLVE PRIORITY Endpoint Picker binds the objective to an InferenceObjective: premium 100 standard 0 batch −10 · unbound → 0 3 · SATURATED? utilization detector = max( QueueDepth / 4, KVCache / 0.8 ) below 1.0 → eligible at 1.0 → remain queued 4 · ORDER across bands: strict priority 100 → 0 → −10 within a band: round-robin
A matching objective supplies the integer priority regardless of saturation. The utilization detector compares queue depth and key-value (KV) cache pressure with their thresholds; queued work dispatches when the admission gate permits it.

Three states of the same mechanism

A · POOL HAS CAPACITY detector below 1 · queue empty premium standard batch engine has room Priority is resolved; little waiting is needed. All three tiers can dispatch promptly with no backlog; the same policy still applies. B · SATURATED, MIXED TIERS detector ≥ 1 · strict priority queue, front → back premium standard batch next slot 1st 2nd 3rd Requests dispatch in priority order. When the gate permits dispatch, the highest nonempty band is served first. Lower bands wait while higher work remains. Measured p95 TTFT: 1117 ms vs 1406 ms at 128 output tokens. C · SATURATED, SAME BAND three tenants, all priority 100 round-robin, taking turns A B C then back to A No band jumps another. All three sit at priority 100, so ordering rotates whose request dispatches next, A then B then C. Bounds a bursty tenant's share. Nonempty tenant queues share the available dispatch turns.
With room and no backlog, dispatch can be prompt. During congestion, strict priority selects the highest nonempty band when the gate permits dispatch; round-robin shares its turns among nonempty tenant queues.

The knobs

The reference campaign used six settings to control engine capacity, saturation detection, and dispatch order.

KnobOwnerWhat it doesSet to
max-num-seqsvLLM engineMax requests the engine runs at once, the running-batch ceiling128
queueDepthThresholdutilization detectorQueue length that counts as saturated4
kvCacheUtilThresholdutilization detectorKV-cache fraction that counts as saturated0.8
maxConcurrencyconcurrency detectorEndpoint-picker cap on in-flight requests before they reach vLLMswept 32–128
headroomconcurrency detectorAllowed burst above the cap0 (default)
priority bandsflow-control orderingThe dispatch order, not a detector: premium 100, standard 0, batch −10100 / 0 / −10

The scenarios run the utilization detector, which ships enabled and reads engine telemetry: the worse of QueueDepth / queueDepthThreshold and KVCacheUsage / kvCacheUtilThreshold per endpoint, averaged across the pool, with backpressure once that ratio crosses one.

The concurrency detector uses maxConcurrency to cap in-flight load before requests reach vLLM. Configuring it alongside vLLM's max-num-seqs demonstrates how admission control and engine capacity interact at two layers.

How we chose the operating point and configuration

The benchmark kept a set number of requests in flight, sending another when one finished. In-flight requests include those queued or processing. We tested concurrency settings from 32 to 200 in two passes for the original RHAII 3.4 campaign. The chart and table show pass 1.

Operating point: throughput peaks at 128, latency climbs past itPass 1 · single tenant · 512 input / 128 output tokens · 180 s per point.01530456005001,0001,5002,000Throughput (rps)p95 TTFT (ms)cap = 128 (set)24.137.145.552.047.150.62617449489671,4551,814326496128160200in-flight requests
Pass 1: 52.0 requests/s and p95 TTFT of 967 ms at 128 in-flight requests.

We chose 128 in-flight requests because both passes reached their highest measured throughput there. Increasing concurrency to 160 or 200 raised first-token latency without improving throughput. For another deployment, choose a tested load that meets its latency target; latency at peak throughput may already exceed that target.

The sweep kept max-num-seqs=128 and flow control enabled with queue-depth threshold 4.

View pass 1 data
Concurrency (requests)Throughput (requests/s)p50 TTFT (ms)p95 TTFT (ms)Mean running requestsp95 waiting requests
3224.1171261284
6437.13077445423
9645.54279488245
12852.05669679968
16047.1708145512232
20050.61523181412272

The running and waiting counts are measured in vLLM. Running is an average; waiting is a 95th percentile across sampled times. They are not simultaneous counts and should not be added together.

At 128 in-flight requests, pass 2 served 54.1 requests/s with p95 TTFT of 940 ms. The two-pass medians used in the README chart are 53.0 requests/s and 954 ms.

Input fixed at 512, output swept 64 / 128 / 512

Input is fixed at 512 tokens; output is swept across 64, 128, and 512, with 128 the headline. This stands in for short interactive work such as chat turns, retrieval-augmented generation (RAG) answers, and tool calls, where a brief reply makes time to first token the felt latency. Output length is the limiter under saturation: a longer generation holds its GPU slot for more decode steps.

Three request shapes, one swept dimension Input 512 tokens, held constant. Output is the swept dimension. 64 out 512 in 128 out 512 in headline 512 out 512 in
The 128-token shape is the headline; the full ladder appears in the service-tiers result.
ChoiceValueWhy
Input length512 in (fixed)Short interactive work; held constant across every scenario
Output length64 / 128 / 512 (swept)Output is the saturation limiter, so we swept it; 128 is the headline
max-num-seqs128 (set)The engine's running-request cap; requests queue after this limit is reached
max-num-batched-tokens8192Total tokens scheduled per engine step
Prefix cachingoffLatencies reflect scheduling, not a warm cache
Repeats3 countedEach headline is the median of the three per-repeat p95s, steady-state trimmed, with the min–max range shown

Each point ran for 180 s. Pass 1 tested concurrency in ascending order (32, 64, 96, 128, 160, 200); pass 2 used 96, 200, 64, 128, 32, 160. Evidence: pass 1 · pass 2.

Client concurrency and engine capacity are different controls

With max-num-seqs=128, vLLM can run up to 128 requests at once. Additional requests wait, even if the hardware might support a higher limit.

The sweep below changed client concurrency while the engine cap remained 128. It measured how offered concurrency affected throughput and latency; it did not test an engine cap of 64.

In pass 1, throughput rose from 24.1 requests per second at 32 in-flight requests to 52.0 at 128. The benchmark used 128 as the selected client operating point for this model, GPU, and request shape.

in-flight requests throughput, requests per second 0204056 326496128 24.1 37.1 45.5 52.0 rises through 128 — engine cap held at 128 64 in-flight requests ≈ 37 rps · measured 128 in-flight requests ≈ 52 rps · selected point client-concurrency sweep
At 64 in-flight requests, the run serves about 37 requests per second. At 128, it reaches 52 requests per second. These points do not establish the effect of changing the engine or admission cap.
pass 1: 32/64/96/128 concurrent → 24.1/37.1/45.5/52.0 requests per second

The binding cap depends on the workload. max-num-seqs limits how many requests run; max-num-batched-tokens limits the total tokens scheduled in each engine step; the key-value (KV) cache limits how much context the pool can hold. Whichever fills first sets the effective ceiling. This client-concurrency sweep held those engine limits fixed; it did not isolate which resource would limit a differently configured engine.

The maxConcurrency detector applies a similar admission cap at the Endpoint Picker layer. At admission caps of 32, 48, 64, and 96, measured mean running counts were 29, 39, 46, and 61. That gives us a second tuning layer: engine capacity is set in vLLM, and request admission can be capped before the request reaches vLLM.

Setup and verification

All runs follow the same basic setup.

ArgValue
Modelopenai/gpt-oss-20b
Hardwareone H100, TP 1
max-num-seqs128
max-num-batched-tokens8192
max-model-len32768
gpu-memory-utilization0.90
enable_prefix_cachingFalse
input / output tokens512 in / 64,128,512 out

Engine args are set in the serving pod spec and verified in the pod startup log. Prefix caching is confirmed by vllm:prefix_cache_queries_total staying flat, and the request shape comes from the client config — input fixed, output swept.

Prefix caching off

Automatic Prefix Caching (APC) is off, so reused prompt prefixes do not explain the measured latency differences. vLLM still uses its KV cache while generating each response; end-to-end latency includes both router and engine work.

Per-request headers

Each request carries two headers.

HeaderCarries
x-llm-d-inference-objectivethe priority tier, premium, standard, or batch
x-llm-d-inference-fairness-idthe tenant, used for same-band fairness

The older x-gateway-inference-objective and x-gateway-inference-fairness-id forms are deprecated aliases and still accepted; the harness sent them; they resolve identically.

How a tier becomes a priority

The Endpoint Picker resolves the header against a matching objective before admission, independently of saturation. We verify premium resolves to 100 before every run.

TierResolves to
premium100
standard0
batch-10

Verification

Two conditions are checked before any latency is trusted.

CheckWherePasses when
Priority routingEndpoint Picker /metricsqueue_duration_seconds{priority="100"} present for premium
Pool saturatedvLLM /metricsnum_requests_running pinned at 128 and num_requests_waiting > 0

Service tiers

When the pool saturates, priority decides who waits. The runs below hold the pool at the operating point with flow control on and measure the separation between tiers.

Mean time per output token stayed at 19–21 ms in every repeat, read from vllm:inter_token_latency_seconds: the queue decides who starts; the pace of tokens already flowing barely moves.

Offered load over the run: standard surges, premium holdsIn-flight requests per tier. Standard floods mid-run while premium keeps a steady low rate.02505007501,000warm-upcool-down0s30s60s90s120stime (elapsed seconds)premiumstandard
Standard surges past the batch cap mid-run while premium holds a low steady rate.
Service tiers, flow control on: premium clears first on a saturated pool p95 TTFT, median of three per-repeat percentiles, 512 in / 128 out, steady-state trimmed. 500 1,000 1,500 0 1117 ms Premium range 1056–1211 ms · spread 1.15x 1406 ms Standard queued behind premium each cycle ~1.25x separation Premium p50 is 374 ms while the pool stays pinned at its 128-slot cap with a real queue behind it.
With flow control on and the pool saturated, premium p95 TTFT settles at 1117 ms against 1406 ms for standard, a ~1.25x separation that held across all three repeats.
Premium p95 TTFT per repeat Each repeat's own steady-state p95; the headline is the median of the three. 400 800 1,200 0 median 1117 ms 1211 1117 1056 repeat 1 repeat 2 repeat 3 The spread across repeats is 1.15x, inside the ≤1.5x stability gate this round applies to every headline.
The three repeats land within a 1.15x band, so the 1117 ms headline is a stable scheduling result, not a lucky draw.
View the data
RepeatPremium p95Premium p50Steady samples
r011211 ms510 ms4108
r021117 ms374 ms4420
r031056 ms345 ms4537
median1117 ms374 ms—

At roughly 48 aggregate requests per second, premium and standard compete for the shared pool. The measured premium steady-state tail was ~1.1 s; end-to-end TTFT alone does not identify which scheduling or engine wait produced the separation.

What we learned An earlier cut of this result pooled every repeat into one percentile, and a startup transient in one repeat produced a headline the steady state could not reproduce. Every headline now passes four guards: per-repeat percentiles, never pooled across repeats; a steady-state trim before counting; a spread gate, max/min ≤ 1.5x across repeats; and saturation verified from the serving metrics, not assumed.

Output length ladder

The 64- and 512-output-token arms are being restated with per-repeat percentiles; their earlier pooled numbers are withdrawn, and the 128-output result above is the restated headline.

The 128-output-token result shows measured tier separation under this workload. An SLA requires explicit latency targets and validation across the intended workloads and load levels.

Tier protection holds on two replicas

The same service-tiers scenario on a two-replica pool with flow control on kept premium ahead of standard, so the ordering survives scale-out.

The two-replica percentiles are being restated with per-repeat percentiles; the ordering and zero rejections stand. Data in benchmark-data/rhaii-3.4-flow-control/multi-replica-tiers.

Batch isolation

In these runs, flow control shifted batch overload from rejections to queueing. The rejection counts alone do not establish unchanged interactive latency.

Offered load over the run: batch ramps in behind the interactive tiersIn-flight requests per tier. Batch floods the pool after the interactive tiers are established.02505007501,000warm-upcool-down0s30s60s90s120stime (elapsed seconds)premiumstandardbatch
Batch ramps in behind the interactive tiers and floods the pool well past the batch cap.
Rejected requests under a batch floodA low-priority batch tenant floods a saturated pool. Flow control off sheds the overflow withHTTP 429s. Flow control on queues it instead.48,224FLOW CONTROL OFFrequests rejected with 4290FLOW CONTROL ONrequests rejectedqueued
Same configured client-concurrency schedule and GPU: flow control off records 48,224 batch HTTP 429s across the three repeats; flow control on records zero. These closed-loop clients issue fewer requests when queueing slows completion, so this is not an equal-arrival-rate comparison.

Batch’s own TTFT percentiles under the flood are being restated with per-repeat percentiles. The mechanism is unchanged: held work waits longer because it queues behind the interactive tiers instead of being shed.

Overnight document and report pipelines can fill the same GPUs that serve interactive traffic by day: the platform holds the work until capacity exists instead of pushing retries back to the application. Rejections at every output length are shown below.

View the data
Output tokens429s off429s on
6444,4850
12848,2240
51253,9400

Consolidation

Two premium tenants share one GPU at higher utilization while a lower-priority tenant floods the pool. This scenario prices whether the interactive tenants keep their tail through the flood.

Offered load over the run: two premium tenants steady, standard floodsIn-flight requests. Two premium tenants hold a packed steady load; a standard tenant floods mid-run.0300600900warm-upcool-down0s30s60s90s120stime (elapsed seconds)premiumstandard
Two premium tenants hold a steady packed load while a standard tenant floods the shared pool mid-run.

The consolidation percentiles are being restated with per-repeat percentiles; the earlier pooled numbers are withdrawn. This scenario explores sharing one GPU while protecting interactive latency. The withdrawn percentiles do not establish SLA preservation, a hardware saving, or a universal noisy-neighbor guarantee.

Same-band fairness

Within one priority band, round-robin bounds a bursty tenant’s share of dispatch turns. Three premium tenants ran at the same priority while tenant A burst to roughly ten times its peers’ rate.

Fairness under a big burst: the burster is served, its excess waits Three tenants, one band, flow control on. Median of three 300 s repeats, steady-state trimmed. SERVED RATE (RPS) 40 80 82 6 6 A (burster) B (peer) C (peer) served through the burst · no starvation TIME QUEUED IN THE ENDPOINT PICKER (MS) 25 50 45 17.5 17.3 A (burster) B (peer) C (peer) the burster waits ~2.6x longer for its turns
Round-robin gives every flow a turn, so the burster is served at 82 requests per second and nobody starves, but its excess waits in its own queue, ~2.6x the peers’ queue time, while the peers hold flat.
View the data
TenantOffered patternServed rateQueue time
A (burster)steady, then a sustained burst82 rps45 ms
B (peer)steady6 rps17.5 ms
C (peer)steady6 rps17.3 ms

Round-robin’s unit of fairness is the dispatch turn, so it bounds the burster’s share of turns rather than the peers’ latency. Every request from every tenant completed with zero errors in all three repeats.

The lesson is that fairness alone is not enough if the vLLM pod is allowed to hold more work than it can process with low TTFT. Round-robin fairness decides which tenant is dequeued next, but the vLLM configuration determines how much work can accumulate inside the engine.

Two saturation detectors

Both detectors gate the same flow-control admission queue. They differ only in how they decide the pool is saturated, and that difference decides whether the running batch can be thinned at all.

The utilization detector, the shipped default, ran every scenario above. It reads engine telemetry, queue depth and KV-cache usage, and computes saturation as the worse of the two ratios, averaged across the pool. Because it reads telemetry, it reacts after queue or KV pressure has built, and because it only applies backpressure, it cannot cap the running batch: the engine still fills to max-num-seqs, 128, and the detector decides who waits behind it.

The concurrency detector is the maxConcurrency variant. It computes saturation as in-flight load over pool capacity and hooks the request lifecycle synchronously, reacting to a new request before telemetry updates. maxConcurrency is an endpoint-picker config that caps in-flight requests before they reach vLLM.

Important note: the concurrency detector counts requests, not memory. A few large requests can still fill KV memory and cause preemption or swapping, so KV-cache utilization, preemption metrics, and memory-related settings should be tracked alongside maxConcurrency.

UTILIZATION DETECTOR (SHIPPED) reads queue depth + KV telemetry requests running batch 128 fills to max-num-seqs priority-ordered queue premium ahead of standard telemetry utilization detector reads it after the queue has formed Decides who waits. Cannot cap the batch. CONCURRENCY DETECTOR (UPSTREAM) caps in-flight load before the engine requests maxConcurrency cap running batch < 128 thinned by the cap synchronous: reacts before telemetry updates counts requests, not memory; large requests can still fill KV cache and cause preemption Caps in-flight load. Thins the batch.
Same admission queue, two ways to call saturation: the utilization detector lets the batch fill to 128; the concurrency detector caps admissions before the engine.

We characterized the concurrency detector by sweeping maxConcurrency under a single saturating load, so the chart below is a tuning result for this traffic shape.

Concurrency detector: maxConcurrency changes the tradeoffSame load at every cap. Premium p95 TTFT was lowest at maxConcurrency 48; standard improved as the cap loosened.02505007501,000Premium p95 TTFT (ms)568461582761907461 ms03,0006,0009,00012,000Standard p95 TTFT (ms)12,7227,8085,2073,7142,69732486496128maxConcurrency (in-flight cap)
Premium p95 TTFT was lowest at 461 ms with maxConcurrency 48. Standard p95 improved from 12,722 ms at cap 32 to 2,697 ms at cap 128. The standard panel spans thirteen times the premium panel's range.
View the data
maxConcurrency (requests)Running requestsPremium p95 TTFTStandard p95 TTFT
3229568 ms12722 ms
4839461 ms7808 ms
6446582 ms5207 ms
9661761 ms3714 ms
12870907 ms2697 ms

Flow control orders the queue, so premium is dispatched before standard regardless of the cap. What maxConcurrency changes is engine occupancy: lower maxConcurrency generally protects TTFT by allowing fewer requests to run inside the vLLM pod at once. However, the result still depends on request shape: long prefill, long decode, or large-context requests can still affect TTFT.

This section characterizes maxConcurrency under one saturating load.

Limits

Round-robin fairness handles tenants within one priority band. Flow control matters most when protecting one priority from another.