Inference economics
The unit-cost analysis of serving model inference (cost per token, GPU capex and utilization) that drives choices like local-first serving and tiered routing.
grounded in: HN trend theme 'AI infrastructure economics' (growing scrutiny of inference cost), grounded in the system's tiered-router, local-first vLLM stack built to control serving cost
Connected concepts
Tiered LLM router, Distributed inference, Local-first inference, vLLM / PagedAttention, Quantization
Explore it live in the knowledge graph →