Flow control benchmark

Flow control benchmark: Endpoint Picker v0.9.0

Can one shared model pool enforce priority and fairness under bursty, production-shaped traffic?

Answer: Across three mixed-priority scenarios, lower-priority work absorbed most of the queuing delay. In a separate same-priority test, two peer tenants remained below 700 ms p95 TTFT while a bursting tenant waited longer. Separate tests identified the tradeoffs among request-count, token-count, and backend-pressure admission signals.

The tests showed two distinct protections: priority separation and same-priority fairness.

Lower-priority work absorbed most of the queuing delay in three mixed-priority scenarios. In the fourth, configured round-robin fairness kept peer tenants moving within one priority band. The latency-sensitive priority tiers (Platinum, Gold, and Silver), realtime tenants A and B in consolidation, and peer tenants B and C in same-priority fairness met the 1.5× repeat-stability gate for surge p95 TTFT across three repeats. Batch isolation remained directional evidence.

Measured latency under surge Across the three repeat-stable scenarios, latency-sensitive tiers or peer tenants recorded median p95 time to first token (TTFT) from 404 to 570 ms. Flow control engaged in every retained run
Per-GPU throughput held Served requests per second (RPS) per GPU varied by at most 0.6% across one, two, and four model replicas. Nine matched runs; HTTP 429: 5/3,498, 1/7,014, and 0/13,950

How the benchmark chose its configuration

Twelve evidence suites form one decision path. Calibration established a production starting point; later suites tested where it held, where another signal helped, and where stable v0.9 admission control stopped.

Configuration carried into the four production scenarios

vLLM concurrency128 maximum sequences
vLLM token budget8,192 maximum batched tokens
Endpoint Picker admission128 in-flight requests with 10% detector headroom; 15% in batch isolation
Queue policies testedRound-robin fairness across tenant flows; first-come-first-served ordering within each flow

Set engine capacity

Find an execution point that keeps the GPU fed without paying a large tail-latency penalty.

Carried forward: 128 sequences and an 8,192-token scheduler budget.

Harden the result

Check whether the selected starting point survives more replicas, repeated surges, and a different routing strategy.

Boundary: scale and recovery held in the tested topology; prefix-aware routing did not justify a default change.

The production scenarios used one H100, one model replica, random routing, and prefix caching off unless a package states otherwise. Each linked suite records its own method, repeats, and claim boundary.

What happened under production-shaped surges?

Three scenarios ask how priority tiers share one model pool. A fourth asks whether peers within one priority band keep receiving dispatch turns.

Range-plot key:Dot: median runLine: minimum–maximum across three repeatsPale area: under 1 second, shown as a reference—not an SLO

Priority tiers

Platinum, Gold, and Silver recorded subsecond median p95 TTFT; Bronze Batch reached 13.3 seconds.

Dot: median of each run's surge-window p95 TTFT. Line: observed minimum–maximum across three repeats. Ranges: Platinum 368–486 ms; Gold 414–594 ms; Silver 592–778 ms; Bronze Batch 11,458–17,728 ms. Open evidence package.

Batch isolation

realtime and Standard recorded lower p95 TTFT than Batch in every repeat, but their repeat spread exceeded the stability gate.

Directional evidence. realtime ranged from 371–669 ms (1.81×); Standard ranged from 436–1,017 ms (2.33×), above the 1.5× stability gate. Batch ranged from 11,425–15,242 ms. Open evidence package.

Consolidation

During the consolidation surge, realtime tenants A and B remained near 500 milliseconds p95 time to first token while the lower-priority Standard burst reached about 25,900 milliseconds.
The figure combines median outstanding requests across three repeats with the median of each run's 100–175 second surge-window p95 TTFT. Median surge p95 TTFT: realtime A 509 ms; realtime B 556 ms; Standard 25,892 ms. Exact ranges: realtime A 503–558 ms; realtime B 505–599 ms; Standard 22,833–27,307 ms. Configuration: one H100, one model replica, request-count limit 128, 10% detector headroom, prefix cache off. This is descriptive evidence from the tested single-GPU topology, not a universal latency guarantee. Open evidence package.

Same-priority fairness

With round-robin explicitly configured, peers B and C stayed below 700 ms while tenant A generated the burst.

Dot: median of each run's surge-window p95 TTFT. Line: observed minimum–maximum across three repeats. Peer ranges across three repeats: B 508–619 ms; C 563–675 ms. Round-robin was the tested policy, not the stable v0.9 default. Open evidence package.

Priority tiers, Consolidation, and Same-priority fairness used 10% detector headroom; Batch isolation used 15%. Flow control engaged and policy queues were active in every retained run. The pale area in the three range plots ends at one second as a reading aid, not a declared service-level objective.

Traffic sent during each surge

Each plot shows requests sent per second from 60 to 210 seconds and uses its own labeled y-axis range. The similar shapes are intentional: every scenario used the same surge window so differences in latency reflect the traffic mix and policy behavior. Prefix caching was off.

Configuration evidence

The Endpoint Picker controls what may enter; vLLM controls how much work may run. The decision map above links every suite, while this section keeps the mechanism and detailed calibration available for review.

These sweeps selected a starting configuration for the production tests. vLLM limits control how much work runs at once; Endpoint Picker limits control when new work queues. Tighter admission reduced realtime latency by shifting more waiting to lower-priority work.

Requests enter the Endpoint Picker with an inference objective and a fairness identity, represented by a tenant label in these tests. The EPP combines the fairness identity and priority into a flow key, stores each flow in its own queue within a priority band, and selects a flow queue and its head request before dispatch.
Priority selects a band. Fairness selects a flow within that band; in these runs, each fairness identity represented one tenant. Ordering selects that flow queue's head request. These benchmarks explicitly configured round-robin fairness; stable v0.9 defaults to global-strict fairness.
Connected scatterplot showing throughput and p95 time to first token across request caps from 16 through 160, with the selected cap of 128 highlighted.
Closed-loop request-count calibration on one H100. Moving from cap 128 to 160 added 3% steady throughput while p95 TTFT increased 17%. Caps 16–96 were single calibration runs; caps 128 and 160 used three repeats. Open the evidence package.
Seven calibration figures and selection details

vLLM active-request limit: 128 kept tail latency lower

Allowing 192 instead of 128 active requests added 1.4 requests/s, while p99 TTFT rose from 1,888 to 2,613 ms.

Max sequences limits how many requests vLLM can actively process. Three matched runs were completed at each setting.

vLLM token budget: 8,192 balanced latency and throughput

The 8,192-token budget produced the highest throughput and lowest p95 TTFT among the three tested settings.

Max batched tokens limits how many prompt and generated tokens vLLM may schedule in one step. The selected point was repeated three times.

Endpoint Picker request limit: 128 reduced first-token latency

Limiting in-flight requests to 128 reduced p95 TTFT by 363 ms with 2.9% less throughput than the 160-request limit.

The request limit caps in-flight work at the Endpoint Picker. Both y-axes are zoomed to the observed range so the two tested settings remain legible.

Waiting-queue threshold: eight queued requests produced lower TTFT

Activating flow control at eight waiting requests lowered p95 TTFT by 256 ms and throughput by 1.6% versus depth five.

Queue depth is the number of requests waiting inside vLLM; it is not request size. Both y-axes are zoomed to the observed range.

Lower request limits protected premium latency but made standard traffic wait longer

The 48-request limit produced the lowest premium p95 TTFT. Increasing the limit admitted more standard work sooner, but premium latency rose.

Both lines use the same logarithmic y-axis in milliseconds. This is a latency tradeoff between priority tiers, not a throughput comparison.

KV-cache thresholds activated flow control as model memory filled

The 0.8 threshold produced lower p95 TTFT than 0.75, while both thresholds were slower than the flow-control-off calibration.

The utilization detector uses this threshold to react to model-memory pressure. The result is a safety calibration, not a latency optimum.

Token-count admission handled mixed request sizes better than the output estimate

Counting exact prompt tokens served 21.4 requests/s; adding the output estimate reduced throughput to 6.8 requests/s.

Token-count admission is concurrencyMode: tokens: it estimates load from prompt size instead of treating every request equally. Request count and exact input tokens have three matched repeats.
Selected settings
LayerSettingUse
vLLM execution128 maximum sequences; 8,192 maximum batched tokensLimits active requests and the token budget per scheduler step
Endpoint Picker admission128 in-flight requests; 10% detector headroomUses a relaxed endpoint-filter boundary above the 128-request saturation reference
Batch isolation128 in-flight requests; 15% detector headroomUses a larger endpoint-filter cushion; it does not reserve capacity by priority
Size-aware optionExact input-token countAccounts for prompt size when request counts misstate load

Choosing an admission signal

The admission signal determines when the Endpoint Picker stops sending more work. The right signal depends on whether request count, prompt size, backend queueing, or model-memory pressure best represents the workload.

In-flight requestsRequestConcurrencyDetector: requests modeUse for similarly sized traffic when protecting realtime TTFT is the first objective. It produced the lowest realtime latency in the matched consolidation and fairness comparisons.
Input tokensRequestConcurrencyDetector: tokens modeTest when prompt sizes vary materially. It improved long-context and batch latency in the mixed workload, but it did not consistently improve the paired long-context realtime comparison.
Waiting queueUtilizationDetector: queue-depth signalUse as a reactive signal for backend queue growth. In the matched production comparisons it reacted later than request count and produced higher realtime TTFT.
KV-cache pressureUtilizationDetector: KV-cache signalUse as a memory-pressure guardrail for long prompts and large KV-cache demand. The sweep calibrated activation thresholds; it did not establish a latency-optimal default.

Headroom changes replica eligibility, not priority capacity

endpoint filtering boundary = detector capacity × (1 + headroom)

For the selected request-count setting, 128 requests with 10% headroom produces a filtering boundary of 140.8. The extra margin tolerates brief per-replica overshoot; it does not hold capacity aside for a priority tier.

v0.9 scope: This report covers stable v0.9's uniform pool-level usage limit and per-replica detector headroom. Priority-specific reserve and in-flight eviction are outside this report.

Detector headroomsaturationDetector.headroomRelaxes the endpoint filter above the detector's request or token limit. It tolerates brief overshoot but does not reserve capacity by priority.
Uniform reservestatic-usage-limit-policy.thresholdSets one dispatch ceiling for every priority. A value below 1.0 holds the same fraction of capacity back from all traffic.

Which admission signal protected realtime latency sooner: in-flight requests or the vLLM waiting queue?

In-flight request limitStops dispatch before admitted requests exceed the configured limit.
Waiting-queue thresholdStops dispatch after requests have already formed a queue inside vLLM.
In-flight token limitAccounts for prompt size and was tested separately under mixed and long-context traffic.

The plots compare request count and queue depth with the same prompts, traffic, model, GPU, and three repeats. Token-count results appear under Workload behavior.

Consolidation: realtime p95 TTFT by admission signal

Request count acted before a vLLM queue formed; both realtime tenants remained near 0.5 seconds.

0 ms3,000 ms6,000 ms

In-flight requests: limit 128

realtime A
509 ms
realtime B
556 ms

vLLM waiting queue: threshold 2

realtime A
4,711 ms
realtime B
4,567 ms

vLLM waiting queue: threshold 5

realtime A
5,117 ms
realtime B
4,906 ms
Dot: median p95 TTFT. Line: observed minimum to maximum across three repeats.

Same-priority fairness: peer p95 TTFT by admission signal

Request count preserved both peers near 0.5 seconds while one same-priority tenant generated the surge.

0 ms3,000 ms6,000 ms

In-flight requests: limit 128

Peer B
527 ms
Peer C
570 ms

vLLM waiting queue: threshold 2

Peer B
5,023 ms
Peer C
4,519 ms
Dot: median p95 TTFT. Line: observed minimum to maximum across three repeats.

Answer: In both production comparisons, the in-flight request limit protected realtime latency sooner. Median p95 TTFT remained near 0.5 seconds; waiting-queue detection reacted after backend queueing began, and median p95 TTFT rose to roughly 4.5–5.1 seconds.

Workload behavior

Request size, response length, and work already running inside vLLM changed which admission setting worked best.

Admission control cannot reclaim work already running in vLLM

realtime p95 TTFT rose from 133 ms to 15,378 ms after running Batch occupied vLLM capacity.

100 ms1 s10 s20 s
realtime only
133 ms
Batch running
15,378 ms
Log scale. Three matched repeats per condition. This test defines the boundary of stable v0.9 admission control; after-dispatch recovery is outside this report.

The agentic workload produced higher TTFT while per-token latency stayed similar

The agentic test used larger prompts, longer outputs, and a lower request rate. It recorded higher TTFT than chat, but this comparison does not isolate one cause.

Surge p95 TTFT (ms)

Chat
420 ms
Agentic
1,352 ms

Surge p95 TPOT (ms/token)

Chat
35.0 ms/token
Agentic
30.4 ms/token
Three single-tenant repeats per request shape; all requests completed.

Admission method changed which priority absorbed the wait

Request count lowered premium p95 TTFT; input-token admission produced more even latency across tiers.

Request countExact input tokens
1,500 ms9,000 ms
Premium
1,994 / 2,914 ms
Gold
2,745 / 2,836 ms
Standard
5,150 / 3,076 ms
Batch
8,654 / 2,832 ms

Use request count when premium latency is the first objective. Test input-token admission when more even service across request sizes matters.

Exact input-token admission did not consistently improve long-context TTFT

Across eight paired seeds, neither admission method was consistently faster.

Request countExact input tokens
250 ms500 ms
101
350 / 343 ms
102
340 / 341 ms
103
290 / 314 ms
104
351 / 324 ms
105
278 / 291 ms
106
382 / 356 ms
107
345 / 362 ms
108
492 / 366 ms
Paired p95 TTFT by seed. Mean difference: 16.4 ms; 95% confidence interval: -23.9 to 56.8 ms.

Scale, recovery, and routing

The selected configuration was tested across larger model pools, repeated surges, and cache-aware routing.

Per-GPU throughput held as the model pool grew

Served RPS per GPU varied by at most 0.6% across one, two, and four replicas.

Premium p95 TTFT ranges: 366–416 ms, 332–349 ms, and 306–349 ms. HTTP 429: 5/3,498 at one replica, 1/7,014 at two, and 0/13,950 at four. Throughput scaled cleanly; sparse 429s prevent a rejection-free claim at every pool size.

realtime latency recovered after both sustained surges

Flow control engaged twice; final p95 TTFT returned to 290 ms from a 299 ms baseline, with zero preemptions.

One 30-minute run, 14,889 completed requests, cache off.

Prefix-aware routing helped some workloads but not the full mix

The tested prompts did not create enough additional prefix reuse to offset uneven routing: realtime and batch improved, while standard long-context traffic slowed.

RandomPrefix-aware
500 ms60,000 ms · log scale
realtime
1,195 / 936 ms
Agentic
4,364 / 3,946 ms
Standard
10,324 / 12,635 ms
Batch
49,763 / 44,462 ms
Overall p95 TTFT by workload across three repeats per routing mode. Standard was 22% slower with prefix-aware routing. Cache hit rate did not increase, and median route imbalance rose from 0.9% to 19.1%.

Random routing remains the reference setting for this workload because prefix-aware routing did not improve latency consistently across tiers.

Method and claim boundaries
AreaMethodBoundary
Production scenariosOpen-loop Poisson traffic, noisy sinusoidal phases, three repeats, cache offOne H100 and one model replica
Detector comparisonThree matched repeats for request count and queue depthDescriptive medians and ranges; queue-depth-5 fairness remains calibration-only
Calibration sweepsClosed-loop fixed concurrencyConfiguration selection; production traffic carries the latency claim
ScaleOne Endpoint Picker with one, two, and four model replicasScope: one Endpoint Picker
Prefix routingTwo model replicas, cache on, shared-prefix trafficHigh repeat variance; no default change
MetricsRequests, policy queues, vLLM running and waiting, KV cache, preemptions, routesAbsolute latency depends on model, hardware, request shape, and load

Evidence packages

Each package includes configuration, request data, traffic samples, system metrics, proof gates, and claim limits.

Engine capacity and configurationvLLM sequence and token limits
Request and token admissionrequest and prompt-token limits
Utilization detector calibrationqueue depth and KV-cache pressure
Production scenariospriority, isolation, consolidation, fairness
Batch interferenceafter-dispatch admission boundary
Selected workload shapeschat and agentic traffic
Mixed production workloadrequest-count and token-count admission
Long-context admissioneight paired seeds
Model-pool scalingone, two, and four replicas
Long stabilitytwo surges and recovery
Prefix-cache routingrandom and prefix-aware routing