All evidence sections
Evidence packages and reports for the benchmark runs.
Plan an experiment
Benchmark decision map
Choose what to test and which measurements to compare.
Red Hat AI Inference 3.5
Upstream llm-d router v0.9.0
Admission tuning and mixed workloads
Request and token limits, priority, fairness, cache routing, scaling, and stability.
Experimental llm-d router build
Hold back and batch eviction
Request limits, priority holdback, realtime latency, and recovery of interrupted Batch requests.
Red Hat AI Inference 3.4
Benchmark-Walkthrough
The benchmark walkthrough: operating-point sweeps, priority traffic, and Batch isolation.