Serving Goodput
The rate of LLM requests completed within their latency SLO (not raw throughput), the metric a team actually optimizes when swapping models for a faster, cheaper agent pipeline.
grounded in: Trend theme 'Production LLM migration for speed and cost: Teams are actively swapping models to get faster, cheaper agent pipelines' — goodput is the SLO-bound metric behind that speed/cost tradeoff.
Connected concepts
Inference economics, Continuous batching, Adaptive inference serving, Prefill/Decode Disaggregation, Inter-Token Latency
Explore it live in the knowledge graph →