raghu@dark-factory :~/kb/cost-aware-eval $ cat

Cost-aware agent eval

Benchmarking and ranking AI agents by the tokens and dollars they spend per completed task, treating cost-efficiency as a first-class evaluation metric alongside raw capability.

grounded in: Trend theme 'Agent token/cost efficiency scrutiny: People are now benchmarking agents on tokens spent and dollars saved, not just capability', exemplified by OpenCode winning attention for sending far

Connected concepts

Evals & benchmarks, Inference economics, Agent token overhead, Harnessability, Verification bottleneck

Explore it live in the knowledge graph →