The discriminative LLM index
IRT-scored, contamination-resistant rankings on dynamically generated items — with confidence intervals, consistency, calibration, and an efficiency frontier. No saturated benchmarks, no leaked test sets. Results stream in live as each model finishes its evaluation.
Demo data — these scores are deterministic placeholders, not a real evaluation. They will be replaced by the first published IRT fit run.
Global Index
index v0.1.0 · run cmsfsx39 · 2026-08-05 08:04 UTC| # | Model | Provider | Score (95% CI) | |
|---|---|---|---|---|
| 1 | DeepSeek: DeepSeek V4 Pro | deepseek | 767 [732–802] | |
| 2 | SpaceXAI: Grok 4.20 Multi-Agent | x-ai | 710 [675–745] | |
| 3 | Google: Gemini 3.6 Flash | 697 [662–732] | ||
| 4 | Mistral: Mistral Small 4 | mistralai | 667 [632–702] | |
| 5 | Anthropic: Claude Sonnet 4.6 | anthropic | 655 [620–690] | |
| 6 | SpaceXAI: Grok 4.20 | x-ai | 647 [612–682] | |
| 7 | Qwen: Qwen3.7 Plus | qwen | 621 [586–656] | |
| 8 | Google: Gemini 3.5 Flash Lite | 604 [569–639] | ||
| 9 | Anthropic: Claude Sonnet 5 | anthropic | 590 [555–625] | |
| 10 | DeepSeek: DeepSeek V4 Flash 0731 | deepseek | 584 [549–619] | |
| 11 | OpenAI: GPT-5.6 Terra | openai | 576 [541–611] | |
| 12 | Mistral: Mistral Medium 3.5 | mistralai | 570 [535–605] | |
| 13 | OpenAI: GPT-5.6 Terra Pro | openai | 570 [535–605] | |
| 14 | Meta: Llama Guard 4 12B | meta-llama | 551 [516–586] | |
| 15 | Meta: Llama 4 Scout | meta-llama | 507 [472–542] | |
| 16 | Qwen: Qwen3.8 Max | qwen | 494 [459–529] |
- #1DeepSeek: DeepSeek V4 Prodeepseek767 [732–802]
- 710 [675–745]
- 697 [662–732]
- #4Mistral: Mistral Small 4mistralai667 [632–702]
- #5Anthropic: Claude Sonnet 4.6anthropic655 [620–690]
- 647 [612–682]
- 621 [586–656]
- 604 [569–639]
- #9Anthropic: Claude Sonnet 5anthropic590 [555–625]
- 584 [549–619]
- #11OpenAI: GPT-5.6 Terraopenai576 [541–611]
- #12Mistral: Mistral Medium 3.5mistralai570 [535–605]
- 570 [535–605]
- #14Meta: Llama Guard 4 12Bmeta-llama551 [516–586]
- #15Meta: Llama 4 Scoutmeta-llama507 [472–542]
- 494 [459–529]
Efficiency frontier
Score vs. cost per 1k items. Frontier models are not dominated on both axes — never collapsed into a single blended number.
| Model | Global Index | Cost / 1k items | Pareto |
|---|---|---|---|
| Anthropic: Claude Sonnet 4.6 | 655 | $0.850 | frontier |
| Qwen: Qwen3.8 Max | 494 | $1.58 | dominated |
| DeepSeek: DeepSeek V4 Flash 0731 | 584 | $2.82 | dominated |
| Google: Gemini 3.6 Flash | 697 | $6.93 | frontier |
| Meta: Llama Guard 4 12B | 551 | $7.96 | dominated |
| OpenAI: GPT-5.6 Terra Pro | 570 | $8.43 | dominated |
| Google: Gemini 3.5 Flash Lite | 604 | $8.60 | dominated |
| Anthropic: Claude Sonnet 5 | 590 | $8.96 | dominated |
| OpenAI: GPT-5.6 Terra | 576 | $10.67 | dominated |
| Qwen: Qwen3.7 Plus | 621 | $12.56 | dominated |
| Meta: Llama 4 Scout | 507 | $15.23 | dominated |
| SpaceXAI: Grok 4.20 | 647 | $15.61 | dominated |
| DeepSeek: DeepSeek V4 Pro | 767 | $22.38 | frontier |
| Mistral: Mistral Medium 3.5 | 570 | $25.87 | dominated |
| SpaceXAI: Grok 4.20 Multi-Agent | 710 | $27.58 | dominated |
| Mistral: Mistral Small 4 | 667 | $29.37 | dominated |