Time to First Token
A serving-latency metric measuring the delay from request arrival to the first generated token, dominated by prompt prefill and a primary driver of perceived agent responsiveness.
grounded in: Latest AI/tech trends, HN notable 2026-07-13 'GPT-5.6 agent migration: A production AI agent moved to a newer model for 2.2x speed and 27% lower cost' under the 'Agent migration for speed and cost win
Connected concepts
Inference economics, Chunked prefill, Continuous batching, Prefill/Decode Disaggregation, Prefix/KV-cache reuse
Explore it live in the knowledge graph →