raghu@dark-factory :~/kb/time-to-first-token $ cat

Time to First Token

A serving-latency metric measuring the delay from request arrival to the first generated token, dominated by prompt prefill and a primary driver of perceived agent responsiveness.

grounded in: Latest AI/tech trends, HN notable 2026-07-13 'GPT-5.6 agent migration: A production AI agent moved to a newer model for 2.2x speed and 27% lower cost' under the 'Agent migration for speed and cost win

Connected concepts

Inference economics, Chunked prefill, Continuous batching, Prefill/Decode Disaggregation, Prefix/KV-cache reuse

Explore it live in the knowledge graph →