The discriminative LLM index

IRT-scored, contamination-resistant rankings on dynamically generated items — with confidence intervals, consistency, calibration, and an efficiency frontier. No saturated benchmarks, no leaked test sets. Results stream in live as each model finishes its evaluation.

Demo data — these scores are deterministic placeholders, not a real evaluation. They will be replaced by the first published IRT fit run.

Global Index

index v0.1.0 · run cmsfsx39 · 2026-08-05 08:04 UTC

Efficiency frontier

Score vs. cost per 1k items. Frontier models are not dominated on both axes — never collapsed into a single blended number.

ModelGlobal IndexCost / 1k itemsPareto
Anthropic: Claude Sonnet 4.6655$0.850frontier
Qwen: Qwen3.8 Max494$1.58dominated
DeepSeek: DeepSeek V4 Flash 0731584$2.82dominated
Google: Gemini 3.6 Flash697$6.93frontier
Meta: Llama Guard 4 12B551$7.96dominated
OpenAI: GPT-5.6 Terra Pro570$8.43dominated
Google: Gemini 3.5 Flash Lite604$8.60dominated
Anthropic: Claude Sonnet 5590$8.96dominated
OpenAI: GPT-5.6 Terra576$10.67dominated
Qwen: Qwen3.7 Plus621$12.56dominated
Meta: Llama 4 Scout507$15.23dominated
SpaceXAI: Grok 4.20647$15.61dominated
DeepSeek: DeepSeek V4 Pro767$22.38frontier
Mistral: Mistral Medium 3.5570$25.87dominated
SpaceXAI: Grok 4.20 Multi-Agent710$27.58dominated
Mistral: Mistral Small 4667$29.37dominated