Inter-Token Latency
The per-token generation delay (time-per-output-token) during decode, distinct from prefill's time-to-first-token, and the dominant driver of an agent pipeline's end-to-end speed.
grounded in: Trend 'Production LLM migration for speed and cost' / NOTABLE 'GPT-5.6 agent migration: moved to a newer model for 2.2x speed' — the '2.2x speed' agents chase is decode throughput, not just TTFT.
Connected concepts
Time to First Token, Inference economics, Model migration, Continuous batching, KV Cache, Speculative decoding
Explore it live in the knowledge graph →