How llm-d flow control schedules shared inference

Flow control is the Endpoint Picker's policy layer for multi-tenant inference. Endpoint Picker queues apply configured priority, fairness, and dispatch ceilings as model-worker pressure rises.

The guide covers request classification, queueing, dispatch, capacity controls, batch eviction, configuration, and operating metrics.

Implementation explanations use the source snapshot checked September 16, 2026. Historical defaults and measured results retain their version labels; source inspection does not verify a deployed image.

Why LLM traffic breaks rate limiting

Traditional gateways limit requests per second and assume every request costs about the same. LLM requests break that assumption.

Cost is driven by input length and by the autoregressive decode loop. One request can consume orders of magnitude more compute and memory than another at the same request rate.

Same request rate, very different cost 256-token prompt, 64-token reply occupies a batch slot briefly 2048-token prompt, 512-token reply occupies compute and KV cache for the whole decode loop Longer prompts and replies can require more compute and KV-cache capacity. Illustration · bar lengths are schematic.

Tenant interference

Large prompts from one tenant can occupy the prefill stage, consume key-value (KV) cache, fill queue space, and increase queue time for other tenants.

Requests cannot move after dispatch

Once a request enters a model server's local queue, the Endpoint Picker cannot route the request to another replica.

Priority inversion

A first-come, first-served model scheduler can place realtime work behind batch requests that arrived earlier.

Request sizes vary

A long-context request can use much more compute and memory than a short request, although each request adds one to a request count.

Flow control provides tunable queue-depth and KV-cache settings for a given traffic pattern and operating goal.

What flow control does at the pool

Flow-control behaviorSystem boundary

Flow control uses Endpoint Picker queues to apply priority, fairness, and dispatch ceilings.

Flow control stays active below the ceiling. Requests still enter its queues and can dispatch as soon as they are eligible. At the ceiling, they wait in the Endpoint Picker.

The Endpoint Picker selects a model-server pod. The scheduler inside that pod decides when the request runs.

Flow control changes when and in what order requests dispatch. Total model capacity still comes from the workers.

When demand exceeds that capacity, lower-priority work absorbs more of the wait so higher-priority work can dispatch first.

How saturation changes request dispatch AVAILABLE CAPACITY Requests Endpoint Picker DISPATCH OPEN vLLM POOL SATURATED New requests PRIORITY QUEUES HIGHER PRIORITY STANDARD BATCH SATURATION AT CEILING Dispatch paused vLLM
Selected example: one static ceiling for all bands. At that ceiling all bands wait; below it, the highest eligible request advances. Other policies can assign different ceilings to different priorities.
Holding new dispatches can reduce additional engine contention. Its effect on time per output token depends on the workload and requires measurement.

How pool saturation is calculated

In this selected utilization-detector example, the Endpoint Picker calculates each endpoint's score from reported waiting-queue depth and KV-cache use, then averages those scores. The illustrated thresholds are waiting 4 and KV 0.80. Detector choice and thresholds are configurable; the optional current-build reference below also explains separate prefill/decode stages.

queue depth ÷ threshold waiting ÷ 4 KV-cache use ÷ limit usage ÷ 0.80 max endpoint score most constrained resource average pool saturation pool_saturation With the saturated staleness policy, stale metrics score 1.0. The detector also filters overloaded endpoints out of routing, failing open if none survive the filter.

Request path

The gateway classifies the request. The Endpoint Picker resolves its flow, queues it when required, applies the dispatch policy, and selects a model replica.

1 · TAG The trusted gateway labels traffic x-llm-d-inference-objective → priority x-llm-d-inference-fairness-id → tenant 2 · CLASSIFY FlowKey = (tenant, priority) Every FlowKey gets its own queue. Missing objective defaults to priority 0. 3 · QUEUE BY POLICY Priority bands, then fairness, then order Band 100 · premium · dispatches first tenant-a queue tenant-b queue selected fairness rotates flows · ordering sorts work Band 0 · standard · waits behind premium tenant-c queue Band -10 · sheddable · bounded queue batch queue Per-band request and byte limits bound Endpoint Picker queue memory. HTTP mappings depend on the build and pool state. See the versioned outcome table below. 4 · THE GATE Priority ceiling check Saturation is a gradient, 0 → 1+. Under the ceiling → dispatch. At or over → the cycle halts. Re-measured every cycle from queuedepth and KV cache, or in-flight count. 5 · LATE BINDING The scheduler picks late The endpoint is chosen at dispatch, where capacity or prefix cache is. vLLM model servers The detector measures pressure; the selected priority ceiling controls dispatch.Endpoint filtering checks individual destinations during scheduling. Trust boundary: the gateway injects headers from identity. Client values do not set them.
x-llm-d-inference-objective and x-llm-d-inference-fairness-id are the canonical header names. The older x-gateway-inference-objective and x-gateway-inference-fairness-id forms are accepted as deprecated aliases.

How priority is assigned

The request header names an InferenceObjective. The Endpoint Picker resolves that objective to the priority stored in the FlowKey. A missing or invalid objective receives the configured default priority.

x-llm-d-inference-objective: gpt-oss-premium InferenceObjective lookup objective present with priority? found missing or nil its declared priority premium → 100 defaultPriority 0 · fallback for this request A missing objective or unset priority uses 0 for that request. Other requests keep their independently resolved priorities.

Flow-control architecture

The Endpoint Picker layer manages the queues and routing to each pod, which is scored at dispatch time.

Clients many tenants Gateway injects the two headers Endpoint Picker flow control queues bands · fairness · ordering · gate scheduler, late binding picks the best endpoint at dispatch time Server 1 · saturated execution busy, requests waiting Server 2 · has room holds this prompt's prefix cache Server 3 · busy execution busy, queue draining Held requests remain unassigned while queued. Endpoint selection occurs when a request leaves the queue.

Priority, fairness, and ordering

Each dispatch cycle checks priority, then fairness, then ordering. Priority selection is fixed. Fairness and ordering use policy plugins.

TierQuestionWho decidesOptions
1 · PriorityWhich band goes first?Fixed. The highest-priority band with queued work is checked firstBands come from InferenceObjective, negative values mark sheddable work
2 · FairnessWhich tenant inside the band?Fairness policy pluginRound-robin across tenants, or global strict ordering
3 · OrderingWhich request from that tenant?Ordering policy pluginFirst-come first-served, earliest deadline first, or a deadline from a service-level-objective (SLO) header
Each cycle considers higher priorities first. At a reached ceiling, dispatch pauses for that cycle; when pressure falls, eligible work can advance.

Use the same band with round-robin fairness when workloads should share dispatch turns. Use separate bands when one requires priority over another.

The dispatch cycle

The dispatch loop checks each band’s applicable ceiling and moves at most one request per pass. The figure simplifies this to one shared static ceiling.

Shared-static-ceiling example

measure saturation a gradient, 0 → 1+ check the ceiling usage-limit policy under walk the bands highest priority first dispatch one item next cycle at or over halt the whole cycle this dispatch cycle stops The halt prevents priority inversion. Dispatching lower-priority work would deepen the saturation a blocked high band is waiting out.
A successful cycle dispatches one request from the highest eligible priority band. The benchmarked v0.9 configuration used the same saturation ceiling for every band.

Fairness policies under the same burst

Round-robin assigns dispatch turns across nonempty flows within one priority band. Global-strict compares flow heads using their ordering policy; with FCFS, it selects the earliest enqueued request across the band.

Round-robin · turns rotate across tenants tenant-a · bursting tenant-b · steady tenant-c · steady cursor rotates turns: a · b · c · a · b · c · a · b · c … Equal dispatch turns while flows remainbacklogged and eligible. a's excess waits in a's own queue
Global-strict · the default flow heads compared using FCFS Earlier enqueued requests dispatch first across eligible flow heads. lowest overhead · no isolation
With continuously backlogged eligible flows, round-robin rotates dispatch opportunities. It does not guarantee equal completions, token throughput, latency, or progress under sustained higher-priority demand. Global-strict with FCFS gives earlier enqueued requests precedence across the band. These are dispatch policies; program-aware-fairness is an advanced alternative.

Ordering policies on the same queue

Illustrative ordering comparisons below show different signals for four waiting requests. EDF uses absolute queue-expiry times. The SLO example assumes equal received timestamps, so its TTFT targets also determine deadline order.

fcfs · first come, first served · benchmarked A · arrived 1st B · 2nd C · 3rd D · 4th serves A B C D using arrival order permits: predictability · ignores: urgency edf · earliest deadline first A · no deadline B · deadline 40s C · deadline 2s D · deadline 15s serves C D B, then A using the closest deadline earlier expiry first · completion before expiry is not guaranteed slodeadline · deadline from an SLO header A · no header B · ttft 500ms C · ttft 2s D · ttft 1s serves B D C, then A using the client SLO permits: per-request targets · requires: header discipline SLO deadline = received time + TTFT target; ordering does not enforce the target
FCFS uses enqueue time. EDF uses enqueue time plus effective queue TTL. SLO ordering uses received time plus TTFT metadata; missing or invalid metadata sorts last. Fairness selects a flow; global-strict also compares flow heads using the ordering policy. A queue TTL is not a TTFT objective.
plugins/flowcontrol/ordering · fcfs · edf · slodeadline

Tenants inside a band

The first request creates a tenant queue. A lease keeps the queue active, and the Endpoint Picker removes the queue after the configured idle period.

each queue tracks its own length and byte size the whole time created first request arrives active · leased requests hold the lease idle · timer running flowGCTimeout counts down collected queue removed a new request re-pins the tenant's queue
Unused flows are collected after their last lease is released and the idle timeout expires. Configured priority bands stay registered.

Where the waiting goes

Excess requests wait in Endpoint Picker policy queues or inside vLLM. Priority and fairness apply while requests remain in the Endpoint Picker.

Where requests wait ENDPOINT PICKER QUEUE Priority selects the band. Fairness selects the tenant. Ordering selects the request. Queue time is measured per tenant and priority band. llm_d_epp_flow_control_request_queue_duration_seconds vLLM WAITING / RUNNING After dispatch, vLLM decides when the request executes. It may wait behind existing work before joining a batch. The Endpoint Picker does not schedule vLLM’s running work. vllm:num_requests_running, vllm:num_requests_waiting Detector settings define pressure; the applicable priority ceiling determines when dispatch pauses. Changing thresholds can move waiting between router and engine. Measure latency and throughput together.

Scheduling behavior to verify

Measure dispatch order, served share, queue time, and rejection behavior under saturation.

Priority
Each cycle considers higher priorities first. At a reached ceiling, dispatch pauses for that cycle; when pressure falls, eligible work can advance.
Fairness
Round-robin gives tenants in the same priority band equal dispatch turns while each tenant has queued work.
Queue time
Queue-duration metrics identify which tenant and priority band absorbed the delay.
Batch limits
A bounded low-priority band limits queued batch work. With a nonempty contended pool, requests beyond the limit receive HTTP 429.
Latency depends on the model, hardware, traffic shape, and configuration. The policy controls dispatch order, tenant selection, request ordering, and queue limits when demand exceeds capacity.

When requests are refused

The Endpoint Picker adds the refusal reason to the x-llm-d-request-dropped-reason response header.

Endpoint Picker outcomes map to the client responses shown in the table.

request EnqueueAndWait capacity check bytes · requests RejectedCapacity queue limits hit · rejected before queueing RejectedOther pre-enqueue failure fits queued in its flow waits for its three picks Dispatched unblocked → scheduler → endpoint EvictedTTL waited past its TTL EvictedContextCancelled the client gave up while queued EvictedOther e.g. shutdown while queued After dispatch, the evictor tracks eligible in-flight requests. Queue policies continue to manage queued work.

Historical v0.9.0 wire mappings. This table preserves the tested version. The current-build reference below distinguishes a contended pool from an empty pool.

What happenedStatusReason on the wire
A band's queue limit was hit, controlled load shedding429rejected-saturated
A request waited past its TTL503rejected-ttl-expired
The client disconnected while queued503rejected-context-cancelled
Graceful shutdown drained the queue, retryable503rejected-shutting-down
Queues are in-memory. A dispatch pause itself does not discard queued work, but TTL expiry, cancellation, shutdown, or process loss can terminate it. Gateway behavior after EPP failure depends on the deployed gateway configuration.

The configuration surface

The default column is historical: router v0.9.0, commit 5f4e762f341a5196393ce79f8a57c3e1900c4a6b. Experimental rows describe separate later benchmark builds, not features tested by the older campaign. Tested examples retain their published values; current-source behavior is documented separately below.

Settingv0.9 default or feature stateTested examplesWhat it changes
Queue, priority, and fairness
priorityBands[].priorityconfigured per band100 / 0 / −10Sets which priority band is considered first.
fairnessPolicyRef (per band)global-strictround-robinDecides which tenant gets the next turn inside a priority band.
orderingPolicyRef (per band)fcfsfcfsOrders requests inside one tenant's queue.
maxBytes (per band)1 GB each; effective capunset → 1 GB × 3 bandsBounds queued request data in each band.
maxBytes (global)unset10 GiB; did not bindBounds queued request data across all bands.
maxRequestsunlimitedunsetBounds the number of queued requests.
defaultRequestTTL0 · client context only60sLimits how long a request may wait in the Endpoint Picker.
expiryCleanupInterval1s1sSets how often expired queued requests are removed.
flowGCTimeout · priorityBandGCTimeoutidle flows and bands collecteddefaultsRemoves inactive flow-control state.
Saturation detector and per-replica headroom
flowControl.saturationDetector.pluginRefutilization-detectorutilization and concurrencySelects how model-replica pressure is measured. Both detector types produce the pool-saturation signal.
utilization-detector.queueDepthThreshold5 requests2 / 5 / 8 requestsSets the ideal waiting-queue depth for each model replica.
utilization-detector.kvCacheUtilThreshold0.80 · 80%0.50–0.90Sets the ideal KV-cache use for each model replica.
concurrency-detector.concurrencyModerequestsrequests and tokensChooses whether pressure is counted as in-flight requests or estimated in-flight tokens.
concurrency-detector.maxConcurrency100 requests48–160 requestsSets ideal in-flight request capacity per model replica when mode is requests.
concurrency-detector.maxTokenConcurrency1,000,000 tokens20,000–75,000 tokensSets ideal in-flight token capacity per model replica when mode is tokens.
headroom on either detector0.0 · 0%0–25%Raises the selected detector's per-replica filtering boundary above its ideal limit. Priority-specific dispatch ceilings use the usage-limit policy.
Priority dispatch ceilings · experimental beyond the v0.9 static policy
flowControl.usageLimitPolicyPluginRefstatic policy at 1.0static and experimental priority holdbackSelects the pool-admission policy that decides whether a priority may dispatch.
static-usage-limit-policy.threshold1.0; full ceiling1.0 in v0.9 scenariosSets one pool-saturation ceiling for every priority.
priority-holdback-policy.minCeiling / maxCeilingexperimental build0.50 / 1.0Holds lower-priority dispatch at lower total-pool saturation; it does not allocate private GPU memory.
priority-holdback-policy.shape / domainexperimental buildlinear / rankControls how ceilings are distributed across the configured priorities.
Batch eviction and retry · experimental
flowControl.enableEvictionexperimental buildtrue in the experimental eviction-and-retry testsLets the Endpoint Picker stop negative-priority work already running in vLLM when higher-priority work is blocked.
Eviction eligibilityrequest priority must be below 0batch priority −10Limits reclamation to work explicitly marked as lower-priority and sheddable.
Retry ownerexternal client or processorAsync ProcessorThe client handles the HTTP 429, starts a new request, and prevents duplicate final results.
Eviction pacing is derived from the selected saturation detector. Confirmation grace and timeout are internal implementation values.

Optional: current-build reference

Source checked September 16, 2026 at router bb2113e4ecd79b049c7322164794ca9ea30b8cbb. These controls describe that source snapshot. Availability does not mean an older benchmark exercised them, and a deployment must be checked against its own image and effective configuration.

Two endpoints: which priority may dispatch, and where?

Selected request-count example: endpoints A and B each have maxConcurrency: 100; nominal pool capacity is 200 accounted requests. With headroom: 0.1, each endpoint's filter threshold is 110. Priority holdback assigns the high band ceiling 1.0 and low band 0.8.

LoadsPoolGateEndpoint
A 109 + B 61170 / 200 = 0.85High may advance; low waitsBoth below 110
A 110 + B 60170 / 200 = 0.85High may advance; low waitsA excluded; B remains
A 80 + B 80160 / 200 = 0.80Low waits at its ceilingBoth remain

A lower-priority ceiling tests total observed pool saturation, not a private queue size or GPU-memory allocation. Headroom leaves the denominator at 200 and creates no physical capacity. The filter uses a strict below-threshold check and returns the original candidates if all would be excluded. Passing these checks permits scheduling; it does not guarantee immediate engine execution.

The loader injects the selected detector into scheduling profiles when it implements filtering. Its absence from handwritten profile YAML alone does not establish inactivity.

What each detector measures

Requests: total accounted requests divided by endpoint count × request limit. Useful for call concurrency; a short call and a long call count equally. Tokens: accounted tokens divided by endpoint count × token limit. Useful for estimated work; it is not a direct GPU utilization or KV-memory measurement. Hybrid: average the larger normalized request/token pressure at each endpoint. Utilization: average the larger waiting-queue/KV pressure at each endpoint, subject to the configured staleness policy.

The accounting unit is an endpoint, not automatically a GPU, pod, or data-parallel rank. Current utilization defaults are waiting threshold 5, KV threshold 0.8, stale after 200 ms, and stale endpoints treated as saturated. The earlier illustrated threshold 4 is a selected example.

Token accounting depends on the in-flight producer and tokenizer. Cached-prefix discounts depend on the selected prefix producer. Optional output estimates use request metadata and an estimate cap; that cap does not limit the response length. With output estimates disabled, first response chunks release prompt-token accounting even while decoding continues. A busy decoder can therefore have zero accounted prompt tokens.

Ceilings, fairness, and deadlines answer different questions

Ceiling: may this priority dispatch at this pressure? Static applies one threshold. Priority holdback spreads ceilings between minCeiling and maxCeiling. Soft-reflective ceilings respond to saturation, priority rank, and cycle state; they do not read measured latency misses.

Fairness: which flow gets a turn within a band? Round-robin is the selected teaching example. Current defaults are global-strict fairness, FCFS ordering, and a static ceiling. Ordering: which request comes first under that policy? FCFS uses enqueue time; EDF uses enqueue time plus queue TTL; slo-deadline-ordering-policy uses received time plus TTFT metadata. Ordering does not guarantee a deadline will be met.

Prefill/decode stages and multiple Endpoint Pickers

The controller applies one selected detector configuration separately to prefill and decode endpoint partitions, then uses the higher stage pressure. Prefill 0.65 and decode 0.90 produce effective pressure 0.90. Combined endpoints contribute to both stages. Do not average all stage endpoints into one undifferentiated fleet.

Model replicas and Endpoint Picker processes are different scaling choices. Cross-picker accounting requires synchronization wiring and does not establish globally shared fairness queues. Recalibrate when hardware, parallelism, workload, or stage ratios change.

Identity, queue outcomes, and cancellation

Objectives model-a-high and model-b-high can both use priority 100. Two InferenceObjective objects named high cannot coexist in the same namespace just because their pool references differ. Missing objectives or priorities fall back to 0. Fairness identity uses an explicit header, then a produced agent identity if available, then the default ID. The gateway must propagate trusted identity; these fields are not authentication.

Dispatch waiting, queue-capacity rejection, queue expiry, client cancellation, and in-flight eviction are separate outcomes. Queue budgets limit router memory or waiting time, not GPU capacity. At this source pin, capacity or TTL expiry against a contended pool maps to 429; an empty pool maps to 503. The historical v0.9 table above retains its TTL 503 mapping. Client disconnects may prevent delivery of any HTTP response.

Priority changes future dispatch order. Optional eviction can terminate eligible lower-priority work already dispatched; the selected sheddable filter accepts priorities below 0. Cancellation must propagate through the gateway to the engine, and stream termination does not prove immediate resource recovery. The client owns retries and duplicate-result handling. For an agent task containing many inference calls and tool pauses, cancelling one call does not automatically discard the entire task; retry logic must account for tool side effects.

Two checks: priority ceiling and endpoint headroom

The priority ceiling decides whether queued work may advance at the observed pool saturation. Headroom changes the selected detector's endpoint filtering threshold when a destination is chosen. Both utilization and concurrency detectors support headroom; neither creates physical capacity.

Two flow diagrams showing reserved capacity comparing pool saturation with a priority ceiling and saturation-detector headroom filtering each model replica against its safety boundary. Utilization uses queue and KV-cache pressure; concurrency uses in-flight requests or tokens.
Priority holdback sets dispatch ceilings; detector headroom changes endpoint eligibility. The selected filter fails open if it would exclude every candidate.
Version and configuration details

Stable v0.9: the static usage-limit policy provides one pool-wide ceiling for all priorities. The default threshold is 1.0. Headroom defaults to 0.0 for both saturation detectors. The experimental priority-holdback policy adds priority-specific ceilings.

Batch eviction interrupts lower-priority work for retry

In the linked eviction experiment, blocked higher-priority demand triggered eviction of eligible negative-priority work. Envoy reset the upstream stream and returned HTTP 429; the Async Processor retried as a new request. This optional path is separate from ordinary dispatch priority.

Actor-aligned sequence showing the Endpoint Picker instructing Envoy to return HTTP 429, Envoy resetting the upstream vLLM stream and sending the response, and the Async Processor retrying the batch job.
The Endpoint Picker creates the eviction instruction. Envoy sends the HTTP response and resets the upstream stream. The Async Processor owns retry.
Benchmark evidence. All 38 observed evictions completed through retry in the one-model-replica tests. All 57 observed evictions completed through retry across two model replicas. Every retried batch job produced one final result. The three-run TTFT scaling comparison for two model replicas was inconclusive.

Configuration patterns

Combine queue limits, admission policy, detector settings, and eviction according to the workload requirement.

Multi-tenant isolation · benchmarked

bandsround-robin + fcfs
detectorutilization · 4 / 0.80
queueTTL 60s · ~1 GB per band

Effect Round-robin gives each active tenant dispatch turns within the priority band.

Tradeoff Requests consume Endpoint Picker queue memory while they wait.

SLO-driven interactive

bandsglobal-strict + slodeadline
detectorconcurrency · maxConcurrency
queueshort TTL

Effect An in-flight request limit controls how much work reaches vLLM. SLO deadlines can order queued requests.

Tradeoff Lower in-flight limits can reduce throughput. Clients must provide the SLO header.

Batch-heavy consolidation

bandsper-band maxBytes / maxRequests
detectorutilization
queuelong TTL on the batch band

Effect Per-band limits bound queued batch request count or bytes.

Tradeoff For a nonempty contended pool, the Endpoint Picker returns HTTP 429 when the batch band reaches its limit.

Protect realtime before batch dispatches · benchmarked

priorityrealtime above batch
admissionpriority holdback ceilings
reservelower priorities stop sooner

Effect Lower-priority work stops dispatching at a lower total-pool saturation, allowing higher-priority dispatch below its own ceiling.

Tradeoff Batch may use less GPU capacity during quiet periods. Calibrate the ceilings for each model and hardware configuration.

Recover capacity after batch starts · benchmarked

eligibilitynegative-priority batch
evictionImmediateResponse → stream reset
retryclient-owned + idempotent

Effect Eviction stops eligible lower-priority generation when realtime demand needs capacity occupied by batch.

Tradeoff Retry recomputes interrupted tokens. Engine metrics may report the released capacity after a delay. The retry owner must prevent duplicate final results.

Max throughput, minimum machinery

bandsglobal-strict + fcfs · defaults
detectorutilization · defaults
ceilingusage-limit 1.0 · shared ceiling

Effect First-come, first-served ordering preserves request arrival order with minimal policy work.

Tradeoff A tenant that sends more requests receives a larger share of dispatches under global-strict fairness.

Priority holdback controls future dispatch. Optional eviction attempts to terminate eligible work already dispatched; engine resource recovery is not instantaneous. Detector headroom changes endpoint eligibility.

Operating flow control

Configuration to tune

Band priorities in InferenceObjective map business tiers to dispatch order. Per-band limits bound how much low-priority work can queue.

Request TTL limits queue time.

The detector threshold defines the pressure signal; the applicable priority ceiling decides when dispatch pauses. maxRequests and maxBytes cap the Endpoint Picker queue.

What to watch

pool_saturation reports the detector's pool saturation value. Dispatch pauses when the value reaches the applicable ceiling.

queue_size reports queued requests per tenant and priority band. Autoscaling systems can use sustained queue growth as a demand signal.

How to verify flow control is active

Verify Endpoint Picker routing, Endpoint Picker queueing, and vLLM pressure during the same saturated interval.

CheckWherePass when
Priority routingEndpoint Picker metricsThe policy-queue metric includes the expected priority bands and tenant identifiers
Flow control engagedEndpoint Picker metricsPool saturation reaches the configured boundary and the policy queue rises above zero requests
Engine pressurevLLM metricsRunning, waiting, KV-cache, and preemption metrics explain what the model workers experienced

Sweep each hardware, model, and workload combination to find its observed latency/throughput knee. Treat it as a calibration candidate, then choose detector limits and priority ceilings for the operating goal. A measured knee is not an automatically selected universal boundary.

The published benchmark evidence pages include configuration sweeps, production traffic, scale tests, and claim boundaries.