Measuring flow control
under pressure.
This walkthrough follows the original RHAII 3.4 utilization-detector campaign. The current v0.9.0 benchmark extends it with admission sweeps, production-shaped traffic, mixed workloads, scaling, routing, and stability tests.
- Service tiers. On a saturated pool with flow control on, premium p95 time to first token (TTFT) holds at 1117 ms against 1406 ms for standard, a ~1.25x separation, stable across repeats.
- Batch isolation. With the same configured client-concurrency schedule, the flow-control-off runs recorded 48k+ batch rejections; the on runs recorded zero. Queueing slowed the closed-loop clients, so the two arms generated different request counts.
- Fairness. In the measured same-priority scenario, round-robin shared dispatch turns with peer tenants while one tenant overloaded the pool. This does not guarantee latency or progress across different priority bands.
- Limits. Round-robin fairness handles tenants within one priority band. Flow control matters most when protecting one priority from another.
Request path
Each tenant tags its requests with a tier and a tenant id, flow control orders them by priority, and one shared vLLM pool serves the result.
How priority is decided
The Endpoint Picker resolves the priority before admission. Queued requests are ordered by policy, and the saturation check determines whether they can dispatch.
Three states of the same mechanism
The knobs
The reference campaign used six settings to control engine capacity, saturation detection, and dispatch order.
| Knob | Owner | What it does | Set to |
|---|---|---|---|
| max-num-seqs | vLLM engine | Max requests the engine runs at once, the running-batch ceiling | 128 |
| queueDepthThreshold | utilization detector | Queue length that counts as saturated | 4 |
| kvCacheUtilThreshold | utilization detector | KV-cache fraction that counts as saturated | 0.8 |
| maxConcurrency | concurrency detector | Endpoint-picker cap on in-flight requests before they reach vLLM | swept 32–128 |
| headroom | concurrency detector | Allowed burst above the cap | 0 (default) |
| priority bands | flow-control ordering | The dispatch order, not a detector: premium 100, standard 0, batch −10 | 100 / 0 / −10 |
The scenarios run the utilization detector, which ships enabled and reads engine telemetry: the worse of QueueDepth / queueDepthThreshold and KVCacheUsage / kvCacheUtilThreshold per endpoint, averaged across the pool, with backpressure once that ratio crosses one.
The concurrency detector uses maxConcurrency to cap in-flight load before requests reach vLLM. Configuring it alongside vLLM's max-num-seqs demonstrates how admission control and engine capacity interact at two layers.
How we chose the operating point and configuration
The benchmark kept a set number of requests in flight, sending another when one finished. In-flight requests include those queued or processing. We tested concurrency settings from 32 to 200 in two passes for the original RHAII 3.4 campaign. The chart and table show pass 1.
We chose 128 in-flight requests because both passes reached their highest measured throughput there. Increasing concurrency to 160 or 200 raised first-token latency without improving throughput. For another deployment, choose a tested load that meets its latency target; latency at peak throughput may already exceed that target.
The sweep kept max-num-seqs=128 and flow control enabled with queue-depth threshold 4.
View pass 1 data
| Concurrency (requests) | Throughput (requests/s) | p50 TTFT (ms) | p95 TTFT (ms) | Mean running requests | p95 waiting requests |
|---|---|---|---|---|---|
| 32 | 24.1 | 171 | 261 | 28 | 4 |
| 64 | 37.1 | 307 | 744 | 54 | 23 |
| 96 | 45.5 | 427 | 948 | 82 | 45 |
| 128 | 52.0 | 566 | 967 | 99 | 68 |
| 160 | 47.1 | 708 | 1455 | 122 | 32 |
| 200 | 50.6 | 1523 | 1814 | 122 | 72 |
The running and waiting counts are measured in vLLM. Running is an average; waiting is a 95th percentile across sampled times. They are not simultaneous counts and should not be added together.
At 128 in-flight requests, pass 2 served 54.1 requests/s with p95 TTFT of 940 ms. The two-pass medians used in the README chart are 53.0 requests/s and 954 ms.
Input fixed at 512, output swept 64 / 128 / 512
Input is fixed at 512 tokens; output is swept across 64, 128, and 512, with 128 the headline. This stands in for short interactive work such as chat turns, retrieval-augmented generation (RAG) answers, and tool calls, where a brief reply makes time to first token the felt latency. Output length is the limiter under saturation: a longer generation holds its GPU slot for more decode steps.
| Choice | Value | Why |
|---|---|---|
| Input length | 512 in (fixed) | Short interactive work; held constant across every scenario |
| Output length | 64 / 128 / 512 (swept) | Output is the saturation limiter, so we swept it; 128 is the headline |
| max-num-seqs | 128 (set) | The engine's running-request cap; requests queue after this limit is reached |
| max-num-batched-tokens | 8192 | Total tokens scheduled per engine step |
| Prefix caching | off | Latencies reflect scheduling, not a warm cache |
| Repeats | 3 counted | Each headline is the median of the three per-repeat p95s, steady-state trimmed, with the min–max range shown |
Each point ran for 180 s. Pass 1 tested concurrency in ascending order (32, 64, 96, 128, 160, 200); pass 2 used 96, 200, 64, 128, 32, 160. Evidence: pass 1 · pass 2.
Client concurrency and engine capacity are different controls
With max-num-seqs=128, vLLM can run up to 128 requests at once. Additional requests wait, even if the hardware might support a higher limit.
The sweep below changed client concurrency while the engine cap remained 128. It measured how offered concurrency affected throughput and latency; it did not test an engine cap of 64.
In pass 1, throughput rose from 24.1 requests per second at 32 in-flight requests to 52.0 at 128. The benchmark used 128 as the selected client operating point for this model, GPU, and request shape.
- 64 in-flight requests: the run serves about 37 requests per second.
- 128 in-flight requests: the run reaches about 52 requests per second with the same engine configuration.
- Tuning choice: if the goal is lower premium TTFT, different configurations may need to be considered, including accepting lower total throughput by admitting fewer requests into vLLM at once.
The binding cap depends on the workload. max-num-seqs limits how many requests run; max-num-batched-tokens limits the total tokens scheduled in each engine step; the key-value (KV) cache limits how much context the pool can hold. Whichever fills first sets the effective ceiling. This client-concurrency sweep held those engine limits fixed; it did not isolate which resource would limit a differently configured engine.
The maxConcurrency detector applies a similar admission cap at the Endpoint Picker layer. At admission caps of 32, 48, 64, and 96, measured mean running counts were 29, 39, 46, and 61. That gives us a second tuning layer: engine capacity is set in vLLM, and request admission can be capped before the request reaches vLLM.
Setup and verification
All runs follow the same basic setup.
| Arg | Value |
|---|---|
| Model | openai/gpt-oss-20b |
| Hardware | one H100, TP 1 |
| max-num-seqs | 128 |
| max-num-batched-tokens | 8192 |
| max-model-len | 32768 |
| gpu-memory-utilization | 0.90 |
| enable_prefix_caching | False |
| input / output tokens | 512 in / 64,128,512 out |
Engine args are set in the serving pod spec and verified in the pod startup log. Prefix caching is confirmed by vllm:prefix_cache_queries_total staying flat, and the request shape comes from the client config — input fixed, output swept.
Prefix caching off
Automatic Prefix Caching (APC) is off, so reused prompt prefixes do not explain the measured latency differences. vLLM still uses its KV cache while generating each response; end-to-end latency includes both router and engine work.
Per-request headers
Each request carries two headers.
| Header | Carries |
|---|---|
| x-llm-d-inference-objective | the priority tier, premium, standard, or batch |
| x-llm-d-inference-fairness-id | the tenant, used for same-band fairness |
The older x-gateway-inference-objective and x-gateway-inference-fairness-id forms are deprecated aliases and still accepted; the harness sent them; they resolve identically.
How a tier becomes a priority
The Endpoint Picker resolves the header against a matching objective before admission, independently of saturation. We verify premium resolves to 100 before every run.
| Tier | Resolves to |
|---|---|
| premium | 100 |
| standard | 0 |
| batch | -10 |
Verification
Two conditions are checked before any latency is trusted.
| Check | Where | Passes when |
|---|---|---|
| Priority routing | Endpoint Picker /metrics | queue_duration_seconds{priority="100"} present for premium |
| Pool saturated | vLLM /metrics | num_requests_running pinned at 128 and num_requests_waiting > 0 |
Service tiers
When the pool saturates, priority decides who waits. The runs below hold the pool at the operating point with flow control on and measure the separation between tiers.
Mean time per output token stayed at 19–21 ms in every repeat, read from vllm:inter_token_latency_seconds: the queue decides who starts; the pace of tokens already flowing barely moves.
View the data
| Repeat | Premium p95 | Premium p50 | Steady samples |
|---|---|---|---|
| r01 | 1211 ms | 510 ms | 4108 |
| r02 | 1117 ms | 374 ms | 4420 |
| r03 | 1056 ms | 345 ms | 4537 |
| median | 1117 ms | 374 ms | — |
At roughly 48 aggregate requests per second, premium and standard compete for the shared pool. The measured premium steady-state tail was ~1.1 s; end-to-end TTFT alone does not identify which scheduling or engine wait produced the separation.
Output length ladder
The 64- and 512-output-token arms are being restated with per-repeat percentiles; their earlier pooled numbers are withdrawn, and the 128-output result above is the restated headline.
The 128-output-token result shows measured tier separation under this workload. An SLA requires explicit latency targets and validation across the intended workloads and load levels.
Tier protection holds on two replicas
The same service-tiers scenario on a two-replica pool with flow control on kept premium ahead of standard, so the ordering survives scale-out.
The two-replica percentiles are being restated with per-repeat percentiles; the ordering and zero rejections stand. Data in benchmark-data/rhaii-3.4-flow-control/multi-replica-tiers.
Batch isolation
In these runs, flow control shifted batch overload from rejections to queueing. The rejection counts alone do not establish unchanged interactive latency.
Batch’s own TTFT percentiles under the flood are being restated with per-repeat percentiles. The mechanism is unchanged: held work waits longer because it queues behind the interactive tiers instead of being shed.
Overnight document and report pipelines can fill the same GPUs that serve interactive traffic by day: the platform holds the work until capacity exists instead of pushing retries back to the application. Rejections at every output length are shown below.
View the data
| Output tokens | 429s off | 429s on |
|---|---|---|
| 64 | 44,485 | 0 |
| 128 | 48,224 | 0 |
| 512 | 53,940 | 0 |
Consolidation
Two premium tenants share one GPU at higher utilization while a lower-priority tenant floods the pool. This scenario prices whether the interactive tenants keep their tail through the flood.
The consolidation percentiles are being restated with per-repeat percentiles; the earlier pooled numbers are withdrawn. This scenario explores sharing one GPU while protecting interactive latency. The withdrawn percentiles do not establish SLA preservation, a hardware saving, or a universal noisy-neighbor guarantee.
Same-band fairness
Within one priority band, round-robin bounds a bursty tenant’s share of dispatch turns. Three premium tenants ran at the same priority while tenant A burst to roughly ten times its peers’ rate.
View the data
| Tenant | Offered pattern | Served rate | Queue time |
|---|---|---|---|
| A (burster) | steady, then a sustained burst | 82 rps | 45 ms |
| B (peer) | steady | 6 rps | 17.5 ms |
| C (peer) | steady | 6 rps | 17.3 ms |
Round-robin’s unit of fairness is the dispatch turn, so it bounds the burster’s share of turns rather than the peers’ latency. Every request from every tenant completed with zero errors in all three repeats.
The lesson is that fairness alone is not enough if the vLLM pod is allowed to hold more work than it can process with low TTFT. Round-robin fairness decides which tenant is dequeued next, but the vLLM configuration determines how much work can accumulate inside the engine.
Two saturation detectors
Both detectors gate the same flow-control admission queue. They differ only in how they decide the pool is saturated, and that difference decides whether the running batch can be thinned at all.
The utilization detector, the shipped default, ran every scenario above. It reads engine telemetry, queue depth and KV-cache usage, and computes saturation as the worse of the two ratios, averaged across the pool. Because it reads telemetry, it reacts after queue or KV pressure has built, and because it only applies backpressure, it cannot cap the running batch: the engine still fills to max-num-seqs, 128, and the detector decides who waits behind it.
The concurrency detector is the maxConcurrency variant. It computes saturation as in-flight load over pool capacity and hooks the request lifecycle synchronously, reacting to a new request before telemetry updates. maxConcurrency is an endpoint-picker config that caps in-flight requests before they reach vLLM.
Important note: the concurrency detector counts requests, not memory. A few large requests can still fill KV memory and cause preemption or swapping, so KV-cache utilization, preemption metrics, and memory-related settings should be tracked alongside maxConcurrency.
We characterized the concurrency detector by sweeping maxConcurrency under a single saturating load, so the chart below is a tuning result for this traffic shape.
View the data
| maxConcurrency (requests) | Running requests | Premium p95 TTFT | Standard p95 TTFT |
|---|---|---|---|
| 32 | 29 | 568 ms | 12722 ms |
| 48 | 39 | 461 ms | 7808 ms |
| 64 | 46 | 582 ms | 5207 ms |
| 96 | 61 | 761 ms | 3714 ms |
| 128 | 70 | 907 ms | 2697 ms |
Flow control orders the queue, so premium is dispatched before standard regardless of the cap. What maxConcurrency changes is engine occupancy: lower maxConcurrency generally protects TTFT by allowing fewer requests to run inside the vLLM pod at once. However, the result still depends on request shape: long prefill, long decode, or large-context requests can still affect TTFT.
- The concurrency detector changed the premium tail: in this sweep, the concurrency detector produced its lowest premium p95 TTFT at 461 ms with
maxConcurrency48. - maxConcurrency tunes the tradeoff. The lowest premium p95 TTFT occurred at
maxConcurrency48. Tighter caps made premium queue against the cap; looser caps admitted more work into vLLM and increased premium TTFT. Standard traffic improved as the cap loosened. - Output length sets how hard the cap bites. A longer generation holds each admitted slot for more decode steps, so at a fixed cap a longer output pushes more work into the queue and lengthens every tail. The output-length ladder in the tiers scenario is being restated with the same per-repeat method.
This section characterizes maxConcurrency under one saturating load.
Limits
Round-robin fairness handles tenants within one priority band. Flow control matters most when protecting one priority from another.
- Priority isolation protects one band from another. Premium p95 TTFT held ~1.25x ahead of standard, 1117 ms against 1406 ms, on a pool pinned at its cap, and 44k to 54k batch rejections became 0 at every output length.
- Round-robin fairness bounds a bursty tenant’s share of dispatch turns inside its own band. Under a big burst its queue time ran ~2.6x its peers’ while the peers held flat and nobody starved. It does not lower peers’ latency, which the shared running batch sets.
- The absolute latencies track one H100. On other hardware the numbers move and the behavior holds: the tier with an objective kept it, and deferrable work waited instead of failing.