Efficiency frontier
Global Index versus measured cost per 1,000 evaluation items (metered tokens × live pricing), on a log scale. Models on the frontier are not dominated on both axes; every other model is strictly worse on quality and cost than some frontier point. By design this is published as a set — never collapsed into a single blended “value” number. The chart updates itself as the live benchmark lands each model.
Pareto frontier (7)dominated (54) 95% CIlive · 61 models
Prefer numbers? The same data lives in the leaderboard table and the public API. Methodology: cost is the mean measured cost per item across a model's scored domains × 1,000; whiskers are the 95% CI of the Global Index; the frontier is computed on (score ↑, cost ↓) dominance.