raghu@dark-factory :~/kb/inference-economics $ cat

Inference economics

The unit-cost analysis of serving model inference (cost per token, GPU capex and utilization) that drives choices like local-first serving and tiered routing.

grounded in: HN trend theme 'AI infrastructure economics' (growing scrutiny of inference cost), grounded in the system's tiered-router, local-first vLLM stack built to control serving cost

Connected concepts

Tiered LLM router, Distributed inference, Local-first inference, vLLM / PagedAttention, Quantization

Explore it live in the knowledge graph →