Model comparison
All evaluated models side by side — capability profile across the five task axes and per-task scores.
- Claude Opus 4.8 (high) *
- GLM 5.2
- Fable 5 (high)
- GPT-5.5
- Claude Sonnet 5 (high)
- Cursor Composer 2.5
- Kimi K2.7 Code
- Grok Build 0.1
- Claude Sonnet 4.6 (high)
- Gemini 3.1 Pro
- Kimi K2.5
Scores by task
| Dimension | Claude Opus 4.8 (high) * | GLM 5.2 | Fable 5 (high) | GPT-5.5 | Claude Sonnet 5 (high) | Cursor Composer 2.5 | Kimi K2.7 Code | Grok Build 0.1 | Claude Sonnet 4.6 (high) | Gemini 3.1 Pro | Kimi K2.5 |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Global index | 97 | 95.6 | 94.2 | 93.6 | 92.9 | 92.4 | 88.8 | 80.9 | 78 | 71 | 63.1 |
| Elo (Bradley–Terry) | 1932 | 1715 | 1736 | 1532 | 1550 | 1490 | 1413 | 1322 | 1456 | 1139 | 1214 |
| Reliability | 71.4% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 60% |
| Efficiency | 71.2 | 74.1 | 41.1 | 92.1 | 33.8 | 79.7 | 30.9 | 81.9 | 31.6 | 90.9 | 93.1 |
| Tasks scored | 5/5 | 5/5 | 5/5 | 5/5 | 5/5 | 5/5 | 5/5 | 5/5 | 5/5 | 5/5 | 3/5 |
| Premium Storefront | 95 | 92 | 85 | 88 | 81 | 81 | 82 | 76 | 88 | 73 | 84 |
| Client-side WebGPU Product Q&A | 98 | 99 | 91 | 97 | 94 | 96 | 81 | 88 | 64 | 52 | ✗ |
| Microsoft Dynamics 365 Order Integration | 98 | 94 | 97 | 91 | 98 | 94 | 93 | 93 | 95 | 91 | 63 |
| In-browser Fashion Fit Estimation | 97 | 98 | 100 | 96 | 95 | 96 | 91 | 52 | 48 | 43 | 43 |
| Autonomous Buying Agent | 98 | 96 | 98 | 96 | 96 | 96 | 96 | 96 | 96 | 96 | ✗ |