Efficiency frontier

Global Index versus measured cost per 1,000 evaluation items (metered tokens × live pricing), on a log scale. Models on the frontier are not dominated on both axes; every other model is strictly worse on quality and cost than some frontier point. By design this is published as a set — never collapsed into a single blended “value” number. The chart updates itself as the live benchmark lands each model.

Pareto frontier (7)dominated (54) 95% CIlive · 61 models
1002003004005006007008009001000$0.001$0.002$0.005$0.01$0.02$0.05$0.10$0.20$0.50$1.0$2.0$5.0$10$20$50$100cost per 1,000 items (USD, log scale) →Global Index →granite-4.0-h-micronova-lite-v1granite-4.1-8bphi-4deepseek-v4-flash-0731ring-2.6-1ttrinity-large-thinking

Prefer numbers? The same data lives in the leaderboard table and the public API. Methodology: cost is the mean measured cost per item across a model's scored domains × 1,000; whiskers are the 95% CI of the Global Index; the frontier is computed on (score ↑, cost ↓) dominance.