Model comparison

All evaluated models side by side — capability profile across the five task axes and per-task scores.

Frontend & Commerce CraftApplied On-device AIIntegration EngineeringOn-device ML & Continual LearningAgentic Planning & Tool Use
  • Claude Opus 4.8 (high) *
  • GLM 5.2
  • Fable 5 (high)
  • GPT-5.5
  • Claude Sonnet 5 (high)
  • Cursor Composer 2.5
  • Kimi K2.7 Code
  • Grok Build 0.1
  • Claude Sonnet 4.6 (high)
  • Gemini 3.1 Pro
  • Kimi K2.5

Scores by task

DimensionClaude Opus 4.8 (high) *GLM 5.2Fable 5 (high)GPT-5.5Claude Sonnet 5 (high)Cursor Composer 2.5Kimi K2.7 CodeGrok Build 0.1Claude Sonnet 4.6 (high)Gemini 3.1 ProKimi K2.5
Global index9795.694.293.692.992.488.880.9787163.1
Elo (Bradley–Terry)19321715173615321550149014131322145611391214
Reliability71.4%100%100%100%100%100%100%100%100%100%60%
Efficiency71.274.141.192.133.879.730.981.931.690.993.1
Tasks scored5/55/55/55/55/55/55/55/55/55/53/5
Premium Storefront9592858881818276887384
Client-side WebGPU Product Q&A98999197949681886452
Microsoft Dynamics 365 Order Integration9894979198949393959163
In-browser Fashion Fit Estimation97981009695969152484343
Autonomous Buying Agent98969896969696969696