raghu@dark-factory :~/kb/serving-goodput $ cat

Serving Goodput

The rate of LLM requests completed within their latency SLO (not raw throughput), the metric a team actually optimizes when swapping models for a faster, cheaper agent pipeline.

grounded in: Trend theme 'Production LLM migration for speed and cost: Teams are actively swapping models to get faster, cheaper agent pipelines' — goodput is the SLO-bound metric behind that speed/cost tradeoff.

Connected concepts

Inference economics, Continuous batching, Adaptive inference serving, Prefill/Decode Disaggregation, Inter-Token Latency

Explore it live in the knowledge graph →