How llm-d flow control schedules shared inference
Flow control is the Endpoint Picker's policy layer for multi-tenant inference. Endpoint Picker queues apply configured priority, fairness, and dispatch ceilings as model-worker pressure rises.
The guide covers request classification, queueing, dispatch, capacity controls, batch eviction, configuration, and operating metrics.
Implementation explanations use the source snapshot checked September 16, 2026. Historical defaults and measured results retain their version labels; source inspection does not verify a deployed image.
Why LLM traffic breaks rate limiting
Traditional gateways limit requests per second and assume every request costs about the same. LLM requests break that assumption.
Cost is driven by input length and by the autoregressive decode loop. One request can consume orders of magnitude more compute and memory than another at the same request rate.
Tenant interference
Large prompts from one tenant can occupy the prefill stage, consume key-value (KV) cache, fill queue space, and increase queue time for other tenants.
Requests cannot move after dispatch
Once a request enters a model server's local queue, the Endpoint Picker cannot route the request to another replica.
Priority inversion
A first-come, first-served model scheduler can place realtime work behind batch requests that arrived earlier.
Request sizes vary
A long-context request can use much more compute and memory than a short request, although each request adds one to a request count.
Flow control provides tunable queue-depth and KV-cache settings for a given traffic pattern and operating goal.
What flow control does at the pool
| Flow-control behavior | System boundary |
|---|---|
Flow control uses Endpoint Picker queues to apply priority, fairness, and dispatch ceilings. Flow control stays active below the ceiling. Requests still enter its queues and can dispatch as soon as they are eligible. At the ceiling, they wait in the Endpoint Picker. The Endpoint Picker selects a model-server pod. The scheduler inside that pod decides when the request runs. |
Flow control changes when and in what order requests dispatch. Total model capacity still comes from the workers. When demand exceeds that capacity, lower-priority work absorbs more of the wait so higher-priority work can dispatch first. |
How pool saturation is calculated
In this selected utilization-detector example, the Endpoint Picker calculates each endpoint's score from reported waiting-queue depth and KV-cache use, then averages those scores. The illustrated thresholds are waiting 4 and KV 0.80. Detector choice and thresholds are configurable; the optional current-build reference below also explains separate prefill/decode stages.
Request path
The gateway classifies the request. The Endpoint Picker resolves its flow, queues it when required, applies the dispatch policy, and selects a model replica.
x-llm-d-inference-objective and x-llm-d-inference-fairness-id are the canonical header names. The older x-gateway-inference-objective and x-gateway-inference-fairness-id forms are accepted as deprecated aliases.How priority is assigned
The request header names an InferenceObjective. The Endpoint Picker resolves that objective to the priority stored in the FlowKey. A missing or invalid objective receives the configured default priority.
Flow-control architecture
The Endpoint Picker layer manages the queues and routing to each pod, which is scored at dispatch time.
Priority, fairness, and ordering
Each dispatch cycle checks priority, then fairness, then ordering. Priority selection is fixed. Fairness and ordering use policy plugins.
| Tier | Question | Who decides | Options |
|---|---|---|---|
| 1 · Priority | Which band goes first? | Fixed. The highest-priority band with queued work is checked first | Bands come from InferenceObjective, negative values mark sheddable work |
| 2 · Fairness | Which tenant inside the band? | Fairness policy plugin | Round-robin across tenants, or global strict ordering |
| 3 · Ordering | Which request from that tenant? | Ordering policy plugin | First-come first-served, earliest deadline first, or a deadline from a service-level-objective (SLO) header |
Use the same band with round-robin fairness when workloads should share dispatch turns. Use separate bands when one requires priority over another.
The dispatch cycle
The dispatch loop checks each band’s applicable ceiling and moves at most one request per pass. The figure simplifies this to one shared static ceiling.
Shared-static-ceiling example
Fairness policies under the same burst
Round-robin assigns dispatch turns across nonempty flows within one priority band. Global-strict compares flow heads using their ordering policy; with FCFS, it selects the earliest enqueued request across the band.
program-aware-fairness is an advanced alternative.Ordering policies on the same queue
Illustrative ordering comparisons below show different signals for four waiting requests. EDF uses absolute queue-expiry times. The SLO example assumes equal received timestamps, so its TTFT targets also determine deadline order.
Tenants inside a band
The first request creates a tenant queue. A lease keeps the queue active, and the Endpoint Picker removes the queue after the configured idle period.
Where the waiting goes
Excess requests wait in Endpoint Picker policy queues or inside vLLM. Priority and fairness apply while requests remain in the Endpoint Picker.
Scheduling behavior to verify
Measure dispatch order, served share, queue time, and rejection behavior under saturation.
When requests are refused
The Endpoint Picker adds the refusal reason to the x-llm-d-request-dropped-reason response header.
Endpoint Picker outcomes map to the client responses shown in the table.
Historical v0.9.0 wire mappings. This table preserves the tested version. The current-build reference below distinguishes a contended pool from an empty pool.
| What happened | Status | Reason on the wire |
|---|---|---|
| A band's queue limit was hit, controlled load shedding | 429 | rejected-saturated |
| A request waited past its TTL | 503 | rejected-ttl-expired |
| The client disconnected while queued | 503 | rejected-context-cancelled |
| Graceful shutdown drained the queue, retryable | 503 | rejected-shutting-down |
The configuration surface
The default column is historical: router v0.9.0, commit 5f4e762f341a5196393ce79f8a57c3e1900c4a6b. Experimental rows describe separate later benchmark builds, not features tested by the older campaign. Tested examples retain their published values; current-source behavior is documented separately below.
| Setting | v0.9 default or feature state | Tested examples | What it changes |
|---|---|---|---|
| Queue, priority, and fairness | |||
priorityBands[].priority | configured per band | 100 / 0 / −10 | Sets which priority band is considered first. |
fairnessPolicyRef (per band) | global-strict | round-robin | Decides which tenant gets the next turn inside a priority band. |
orderingPolicyRef (per band) | fcfs | fcfs | Orders requests inside one tenant's queue. |
maxBytes (per band) | 1 GB each; effective cap | unset → 1 GB × 3 bands | Bounds queued request data in each band. |
maxBytes (global) | unset | 10 GiB; did not bind | Bounds queued request data across all bands. |
maxRequests | unlimited | unset | Bounds the number of queued requests. |
defaultRequestTTL | 0 · client context only | 60s | Limits how long a request may wait in the Endpoint Picker. |
expiryCleanupInterval | 1s | 1s | Sets how often expired queued requests are removed. |
flowGCTimeout · priorityBandGCTimeout | idle flows and bands collected | defaults | Removes inactive flow-control state. |
| Saturation detector and per-replica headroom | |||
flowControl.saturationDetector.pluginRef | utilization-detector | utilization and concurrency | Selects how model-replica pressure is measured. Both detector types produce the pool-saturation signal. |
utilization-detector.queueDepthThreshold | 5 requests | 2 / 5 / 8 requests | Sets the ideal waiting-queue depth for each model replica. |
utilization-detector.kvCacheUtilThreshold | 0.80 · 80% | 0.50–0.90 | Sets the ideal KV-cache use for each model replica. |
concurrency-detector.concurrencyMode | requests | requests and tokens | Chooses whether pressure is counted as in-flight requests or estimated in-flight tokens. |
concurrency-detector.maxConcurrency | 100 requests | 48–160 requests | Sets ideal in-flight request capacity per model replica when mode is requests. |
concurrency-detector.maxTokenConcurrency | 1,000,000 tokens | 20,000–75,000 tokens | Sets ideal in-flight token capacity per model replica when mode is tokens. |
headroom on either detector | 0.0 · 0% | 0–25% | Raises the selected detector's per-replica filtering boundary above its ideal limit. Priority-specific dispatch ceilings use the usage-limit policy. |
| Priority dispatch ceilings · experimental beyond the v0.9 static policy | |||
flowControl.usageLimitPolicyPluginRef | static policy at 1.0 | static and experimental priority holdback | Selects the pool-admission policy that decides whether a priority may dispatch. |
static-usage-limit-policy.threshold | 1.0; full ceiling | 1.0 in v0.9 scenarios | Sets one pool-saturation ceiling for every priority. |
priority-holdback-policy.minCeiling / maxCeiling | experimental build | 0.50 / 1.0 | Holds lower-priority dispatch at lower total-pool saturation; it does not allocate private GPU memory. |
priority-holdback-policy.shape / domain | experimental build | linear / rank | Controls how ceilings are distributed across the configured priorities. |
| Batch eviction and retry · experimental | |||
flowControl.enableEviction | experimental build | true in the experimental eviction-and-retry tests | Lets the Endpoint Picker stop negative-priority work already running in vLLM when higher-priority work is blocked. |
| Eviction eligibility | request priority must be below 0 | batch priority −10 | Limits reclamation to work explicitly marked as lower-priority and sheddable. |
| Retry owner | external client or processor | Async Processor | The client handles the HTTP 429, starts a new request, and prevents duplicate final results. |
Optional: current-build reference
Source checked September 16, 2026 at router bb2113e4ecd79b049c7322164794ca9ea30b8cbb. These controls describe that source snapshot. Availability does not mean an older benchmark exercised them, and a deployment must be checked against its own image and effective configuration.
Two endpoints: which priority may dispatch, and where?
Selected request-count example: endpoints A and B each have maxConcurrency: 100; nominal pool capacity is 200 accounted requests. With headroom: 0.1, each endpoint's filter threshold is 110. Priority holdback assigns the high band ceiling 1.0 and low band 0.8.
| Loads | Pool | Gate | Endpoint |
|---|---|---|---|
| A 109 + B 61 | 170 / 200 = 0.85 | High may advance; low waits | Both below 110 |
| A 110 + B 60 | 170 / 200 = 0.85 | High may advance; low waits | A excluded; B remains |
| A 80 + B 80 | 160 / 200 = 0.80 | Low waits at its ceiling | Both remain |
A lower-priority ceiling tests total observed pool saturation, not a private queue size or GPU-memory allocation. Headroom leaves the denominator at 200 and creates no physical capacity. The filter uses a strict below-threshold check and returns the original candidates if all would be excluded. Passing these checks permits scheduling; it does not guarantee immediate engine execution.
The loader injects the selected detector into scheduling profiles when it implements filtering. Its absence from handwritten profile YAML alone does not establish inactivity.
What each detector measures
Requests: total accounted requests divided by endpoint count × request limit. Useful for call concurrency; a short call and a long call count equally. Tokens: accounted tokens divided by endpoint count × token limit. Useful for estimated work; it is not a direct GPU utilization or KV-memory measurement. Hybrid: average the larger normalized request/token pressure at each endpoint. Utilization: average the larger waiting-queue/KV pressure at each endpoint, subject to the configured staleness policy.
The accounting unit is an endpoint, not automatically a GPU, pod, or data-parallel rank. Current utilization defaults are waiting threshold 5, KV threshold 0.8, stale after 200 ms, and stale endpoints treated as saturated. The earlier illustrated threshold 4 is a selected example.
Token accounting depends on the in-flight producer and tokenizer. Cached-prefix discounts depend on the selected prefix producer. Optional output estimates use request metadata and an estimate cap; that cap does not limit the response length. With output estimates disabled, first response chunks release prompt-token accounting even while decoding continues. A busy decoder can therefore have zero accounted prompt tokens.
Ceilings, fairness, and deadlines answer different questions
Ceiling: may this priority dispatch at this pressure? Static applies one threshold. Priority holdback spreads ceilings between minCeiling and maxCeiling. Soft-reflective ceilings respond to saturation, priority rank, and cycle state; they do not read measured latency misses.
Fairness: which flow gets a turn within a band? Round-robin is the selected teaching example. Current defaults are global-strict fairness, FCFS ordering, and a static ceiling. Ordering: which request comes first under that policy? FCFS uses enqueue time; EDF uses enqueue time plus queue TTL; slo-deadline-ordering-policy uses received time plus TTFT metadata. Ordering does not guarantee a deadline will be met.
Prefill/decode stages and multiple Endpoint Pickers
The controller applies one selected detector configuration separately to prefill and decode endpoint partitions, then uses the higher stage pressure. Prefill 0.65 and decode 0.90 produce effective pressure 0.90. Combined endpoints contribute to both stages. Do not average all stage endpoints into one undifferentiated fleet.
Model replicas and Endpoint Picker processes are different scaling choices. Cross-picker accounting requires synchronization wiring and does not establish globally shared fairness queues. Recalibrate when hardware, parallelism, workload, or stage ratios change.
Identity, queue outcomes, and cancellation
Objectives model-a-high and model-b-high can both use priority 100. Two InferenceObjective objects named high cannot coexist in the same namespace just because their pool references differ. Missing objectives or priorities fall back to 0. Fairness identity uses an explicit header, then a produced agent identity if available, then the default ID. The gateway must propagate trusted identity; these fields are not authentication.
Dispatch waiting, queue-capacity rejection, queue expiry, client cancellation, and in-flight eviction are separate outcomes. Queue budgets limit router memory or waiting time, not GPU capacity. At this source pin, capacity or TTL expiry against a contended pool maps to 429; an empty pool maps to 503. The historical v0.9 table above retains its TTL 503 mapping. Client disconnects may prevent delivery of any HTTP response.
Priority changes future dispatch order. Optional eviction can terminate eligible lower-priority work already dispatched; the selected sheddable filter accepts priorities below 0. Cancellation must propagate through the gateway to the engine, and stream termination does not prove immediate resource recovery. The client owns retries and duplicate-result handling. For an agent task containing many inference calls and tool pauses, cancelling one call does not automatically discard the entire task; retry logic must account for tool side effects.
Two checks: priority ceiling and endpoint headroom
The priority ceiling decides whether queued work may advance at the observed pool saturation. Headroom changes the selected detector's endpoint filtering threshold when a destination is chosen. Both utilization and concurrency detectors support headroom; neither creates physical capacity.
Version and configuration details
Stable v0.9: the static usage-limit policy provides one pool-wide ceiling for all priorities. The default threshold is 1.0. Headroom defaults to 0.0 for both saturation detectors. The experimental priority-holdback policy adds priority-specific ceilings.
Batch eviction interrupts lower-priority work for retry
In the linked eviction experiment, blocked higher-priority demand triggered eviction of eligible negative-priority work. Envoy reset the upstream stream and returned HTTP 429; the Async Processor retried as a new request. This optional path is separate from ordinary dispatch priority.
Configuration patterns
Combine queue limits, admission policy, detector settings, and eviction according to the workload requirement.
Multi-tenant isolation · benchmarked
Effect Round-robin gives each active tenant dispatch turns within the priority band.
Tradeoff Requests consume Endpoint Picker queue memory while they wait.
SLO-driven interactive
Effect An in-flight request limit controls how much work reaches vLLM. SLO deadlines can order queued requests.
Tradeoff Lower in-flight limits can reduce throughput. Clients must provide the SLO header.
Batch-heavy consolidation
Effect Per-band limits bound queued batch request count or bytes.
Tradeoff For a nonempty contended pool, the Endpoint Picker returns HTTP 429 when the batch band reaches its limit.
Protect realtime before batch dispatches · benchmarked
Effect Lower-priority work stops dispatching at a lower total-pool saturation, allowing higher-priority dispatch below its own ceiling.
Tradeoff Batch may use less GPU capacity during quiet periods. Calibrate the ceilings for each model and hardware configuration.
Recover capacity after batch starts · benchmarked
Effect Eviction stops eligible lower-priority generation when realtime demand needs capacity occupied by batch.
Tradeoff Retry recomputes interrupted tokens. Engine metrics may report the released capacity after a delay. The retry owner must prevent duplicate final results.
Max throughput, minimum machinery
Effect First-come, first-served ordering preserves request arrival order with minimal policy work.
Tradeoff A tenant that sends more requests receives a larger share of dispatches under global-strict fairness.
Operating flow control
Configuration to tune
Band priorities in InferenceObjective map business tiers to dispatch order. Per-band limits bound how much low-priority work can queue.
Request TTL limits queue time.
The detector threshold defines the pressure signal; the applicable priority ceiling decides when dispatch pauses. maxRequests and maxBytes cap the Endpoint Picker queue.
What to watch
pool_saturation reports the detector's pool saturation value. Dispatch pauses when the value reaches the applicable ceiling.
queue_size reports queued requests per tenant and priority band. Autoscaling systems can use sustained queue growth as a demand signal.
How to verify flow control is active
Verify Endpoint Picker routing, Endpoint Picker queueing, and vLLM pressure during the same saturated interval.
| Check | Where | Pass when |
|---|---|---|
| Priority routing | Endpoint Picker metrics | The policy-queue metric includes the expected priority bands and tenant identifiers |
| Flow control engaged | Endpoint Picker metrics | Pool saturation reaches the configured boundary and the policy queue rises above zero requests |
| Engine pressure | vLLM metrics | Running, waiting, KV-cache, and preemption metrics explain what the model workers experienced |
Sweep each hardware, model, and workload combination to find its observed latency/throughput knee. Treat it as a calibration candidate, then choose detector limits and priority ceilings for the operating goal. A measured knee is not an automatically selected universal boundary.
The published benchmark evidence pages include configuration sweeps, production traffic, scale tests, and claim boundaries.