← Back to overview

Rank #1 · 5/5 tasks scored

Claude Opus 4.8 (high) *

* Manual baseline: hand-built in-IDE by Claude Opus 4.8 from the frozen v3 prompt because the automated SDK storefront runs hit RESOURCE_EXHAUSTED. Not a metered one-shot (time/tokens/cost n/a, not comparable); scored by the identical evaluator. — ORBE is a floor-complete premium PDP with raw WebGL configuration, persistent cart/pricing, a five-step checkout, rule-based agentic assistant with undo, and rich JSON-LD, plus thoughtful extras like room-fit guidance, engraving, bundles, and loyalty gamification. Visual craft and microcopy are strong and cohesive, but system typography, SVG placeholders, static stock, and locale formatting without UI translation keep it below truly exceptional tier.

  • Elo 1932
  • Reliability 71.4% (5/7 runs)
  • Efficiency 71.2/100
97
Global capability index (weighted mean of scored tasks)

Capability profile

Five task axes (each normalized to its 0–100 task score).

Frontend & Commerce CraftApplied On-device AIIntegration EngineeringOn-device ML & Continual LearningAgentic Planning & Tool Use

Per-task scores

TaskScoreBaseExcel.JudgeRobust.TierElo
Premium Storefront95451920.810Full2187
Client-side WebGPU Product Q&A984519.92310Full1945
Microsoft Dynamics 365 Order Integration984518.823.810Full2063
In-browser Fashion Fit Estimation974517.724.210Full1829
Autonomous Buying Agent9845182510Full1888

Score composition per task: 45 base + 20 excellence + 25 judge + 10 robustness. "n/a" = no data for that component in this run (weights renormalized). Failed runs are reliability events, not 0-scores.

Task A — Premium Storefront

94.8base 45/45 + excellence 19/20 + judge 20.8/25 + robustness 10/10 · tier full · agent 73/100

Score breakdown

  • Functional20 / 20
  • Visual Design18 / 20
  • UX18 / 20
  • Engineering17.5 / 20
  • AI Quality18 / 20

Engineering in detail

  • Structure / maintainability / readability8 / 10
  • Performance3.5 / 4
  • Accessibility3 / 3
  • Error-freeness3 / 3

Lighthouse

88Performance
100Accessibility
100Best Practices
100SEO

CLS 0.007 · LCP 1588ms · Page weight 97 KB · axe 0 (crit 0)

Live preview

The actual storefront generated by the model — interactive.Open in new tab ↗

Screenshots

Claude Opus 4.8 (high) * — Desktop · Light
Desktop · Light
Claude Opus 4.8 (high) * — Desktop · Dark
Desktop · Dark
Claude Opus 4.8 (high) * — Mobile
Mobile

Verified interactions

Behavior actually driven in the browser (not just present in the DOM).

  • Add to cart workspass
  • Dark mode togglespass
  • Variant changes price/gallerypass
  • Search filters the catalogpass
  • AI assistant performs an actionpass
  • Cart persists across reloadpass

Structured data & SEO

✓ Product✓ Offer✓ AggregateRating✓ BreadcrumbList
  • ✓ Meta description
  • ✓ Canonical URL
  • ✓ Open Graph image
  • 100% of images have alt text (14/14)

Tokens & cost

Token usage is reported by the agent run. Cost is an estimate (tokens × configured rates); shows “—” until rates are set.

Total tokens
Input tokens
Output tokens
Est. cost
Cost / 100 pts

Runtime metrics

Run time
Tool calls
Time to first tool
1.6sTime to first render
0Runtime errors

Deductions (−1)

  • −1 Copy-paste template

Feature matrix (25/25)

  • 3D configurator (WebGL)present
  • Product gallerypresent
  • Image zoompresent
  • Color variantspresent
  • Size selectionpresent
  • Live stockpresent
  • Pricepresent
  • Discountpresent
  • Buy boxpresent
  • Sticky behaviorpresent
  • Reviewspresent
  • Cross-sellingpresent
  • Search / filterpresent
  • Mobile navigationpresent
  • Wishlistpresent
  • Cartpresent
  • Multi-step checkoutpresent
  • Currency / locale switchpresent
  • AI assistantpresent
  • Dark modepresent
  • Animationspresent
  • Accessibility (basics)present

Per-task results

Each task this model also ran, with the same depth as Task A where the task allows it: static-app tasks show screenshots and a live preview; backend / agent tasks show the harness probe breakdown and key metrics.

Client-side WebGPU Product Q&A

97.9

Applied On-device AI · static app

  • Tier Full
  • Agent 91/100
  • Auto 75/75
  • Contract ✓
  • Base 45/45
  • Excellence 19.9/20
  • Judge 23/25
  • Robustness 10/10

A trust-first dual-layer design (deterministic catalog retrieval with optional verified LLM paraphrase) delivers visible token streaming, per-answer citations, and strong grounding guards. Minor gaps: no cancel during generation and the backend label can read WebGPU even when answers are served from the catalog fallback.

Client-side WebGPU Product Q&A on Client-side WebGPU Product Q&A — Desktop · Light
Desktop · Light
Client-side WebGPU Product Q&A on Client-side WebGPU Product Q&A — Desktop · Dark
Desktop · Dark
Client-side WebGPU Product Q&A on Client-side WebGPU Product Q&A — Mobile
Mobile
Live preview — interactive app

Loads the actual app generated by this model.Open in a new tab ↗

Probe breakdown — automated 75/75
  • Harness hook present & well-formed__ask returns { answer, sources? }6/6passed
  • In-scope factual answers grounded in catalog6/6 in-scope factual correct16/16passed
  • Multi-fact reasoning answers2/2 multi-fact correct8/8passed
  • Refuses out-of-scope questions2/2 out-of-scope refused7/7passed
  • Refuses adversarial / fabrication bait2/2 adversarial refused7/7passed
  • Answers cite their catalog source8/8 in-scope answers cited a source5/5passed
  • On-device model initialization tierreported tier=webgpu, engine=web-llm@0.2.84 · Qwen2.5-0.5B-Instruct-q4f32_1-MLC, navigator.gpu=true10/10passed
  • Latency within budget (TTFT + tokens/sec)ttftMs=1 (budget 30000), tokensPerSec=420004/4passed
  • No off-allowlist traffic after loadno off-allowlist hosts6/6passed
  • Robust to hostile input (empty / very long / rapid-fire)empty=true longInput=true rapidFire=true6/6passed

Key metrics

1Time to first token (ms)
42000Tokens / sec
3159Model load (ms)
17Network requests
0Off-allowlist requests
0Console errors
12Q&A total
12Q&A passed
Compare all models on this task →

Microsoft Dynamics 365 Order Integration

97.6

Integration Engineering · backend / agent task

  • Tier Full
  • Agent 88/100
  • Auto 80/80
  • Contract ✓
  • Base 45/45
  • Excellence 18.8/20
  • Judge 23.8/25
  • Robustness 10/10

The deliverable implements a declarative, side-effect-free mapping DSL with complete entity coverage and a well-isolated transport client featuring retry, jitter, Retry-After handling, and idempotency keys. Validation returns field-level paths with specific reasons, and structured JSON logging redacts secrets while recording sync outcomes and retries.

Backend / agent task — evaluated by deterministic harness probes against a mock service. There is no visual preview for this submission.

Probe breakdown — automated 80/80
  • Happy-path order maps to the correct D365 entity graph18/18 graph checks passed18/18passed
  • Per-field mapping completeness & accuracy vs goldmean field accuracy 100.0% over 4 orders20/20passed
  • Idempotent on retry (one key => exactly one sales order)salesorders=1 (want 1), lines=2 (want 2), replayFlag=true12/12passed
  • Recovers from injected faults (429/500/reset) with retry + backoffgraph 100%, faultsServed=4, soPosts=3, accPosts=212/12passed
  • Customer upsert: lookup-or-create without duplicatesreusedNoDup=true, salesorderRefsSeed=true, newCreated=true8/8passed
  • Structured 4xx on malformed input with no partial writes3/3 rejected with structured 4xx, partialWrites=false8/8passed
  • No credentials hardcoded or loggedleakInSource=false, leakInLogs=false, readsProcessEnv=true2/2passed

Key metrics

118API calls
4Retries
5p50 latency (ms)
Compare all models on this task →

In-browser Fashion Fit Estimation

96.9

On-device ML & Continual Learning · static app

  • Tier Full
  • Agent 92/100
  • Auto 70/70
  • Contract ✓
  • Base 45/45
  • Excellence 17.7/20
  • Judge 24.2/25
  • Robustness 10/10

Outstanding probabilistic UX with labeled confidence, per-metric ranges, size-distribution bars, and explicit flags for low-confidence or borderline inputs. Privacy and opt-in consent are credible and repeated; chest-driven sizing and return-adjustment explanations are clear, with a measured holdout before/after learning demo.

In-browser Fashion Fit Estimation on In-browser Fashion Fit Estimation — Desktop · Light
Desktop · Light
In-browser Fashion Fit Estimation on In-browser Fashion Fit Estimation — Desktop · Dark
Desktop · Dark
In-browser Fashion Fit Estimation on In-browser Fashion Fit Estimation — Mobile
Mobile
Live preview — interactive app

Loads the actual app generated by this model.Open in a new tab ↗

Probe breakdown — automated 70/70
  • Loads & produces well-formed metricswell-formed { measurements, size, confidence }8/8passed
  • Real on-device model + runtime loadedreal model (external=true, backend=webgpu)6/6passed
  • Measurement accuracy vs gold (MAE)meanMAE=1.22cm per-metric={"chest":2.2,"waist":1.57,"hip":1.58,"inseam":0.28,"shoulder":0.45}16/16passed
  • Recommended size top-1 accuracytop1=8/810/10passed
  • Recommended size within ±1within1=8/86/6passed
  • Online learning lowers holdout error (CORE)holdout err 1 -> 014/14passed
  • Graceful no-person / bad-image / non-imagenoperson:no_person✓ bad:bad_image✓ notimage:non_image✓4/4passed
  • No image egress / on-device onlyoffAllowlist=0 bigUploadsAfterEstimate=06/6passed

Key metrics

1598Model load (ms)
103Estimate latency (ms)
1.22Mean abs. error (cm)
2.2MAE chest (cm)
1.57MAE waist (cm)
1.58MAE hip (cm)
0.28MAE inseam (cm)
0.45MAE shoulder (cm)
1Size top-1
1Holdout error (before)
0Holdout error (after)
0Off-allowlist requests
Compare all models on this task →

Autonomous Buying Agent

98

Agentic Planning & Tool Use · backend / agent task

  • Tier Full
  • Agent 91/100
  • Auto 85/85
  • Contract ✓
  • Base 45/45
  • Excellence 18/20
  • Judge 25/25
  • Robustness 10/10

The agent documents a clear six-step strategy up front, evaluates every candidate with best-coupon and shipping totals, and reports impossibility with concrete breakdowns instead of placing orders. It goes beyond the brief with retry/backoff, stock-change fallbacks, coupon caching, and a post-checkout budget safety check.

Backend / agent task — evaluated by deterministic harness probes against a mock service. There is no visual preview for this submission.

Probe breakdown — automated 85/85
  • Agent runs and writes a valid report for every scenario8/8 — 01-happy-hoodie:ok 02-budget-tight-tee:ok 03-coupon-optimality-sneaker:ok 04-deadline-express-tee:ok 05-faults-recovery-hoodie:ok 06-impossible-budget-hoodie:ok 07-oos-size-hoodie:ok 08-impossible-stock-sneaker:ok6/6passed
  • Scenario goal achieved end-to-end (or impossible handled correctly)8/8 — 01-happy-hoodie:ok 02-budget-tight-tee:ok 03-coupon-optimality-sneaker:ok 04-deadline-express-tee:ok 05-faults-recovery-hoodie:ok 06-impossible-budget-hoodie:ok 07-oos-size-hoodie:ok 08-impossible-stock-sneaker:ok24/24passed
  • No placed order ever exceeds the scenario budget8/8 — 01-happy-hoodie:ok 02-budget-tight-tee:ok 03-coupon-optimality-sneaker:ok 04-deadline-express-tee:ok 05-faults-recovery-hoodie:ok 06-impossible-budget-hoodie:ok 07-oos-size-hoodie:ok 08-impossible-stock-sneaker:ok12/12passed
  • Hard constraints satisfied (in-stock, size, quantity, deadline)6/6 — 01-happy-hoodie:ok 02-budget-tight-tee:ok 03-coupon-optimality-sneaker:ok 04-deadline-express-tee:ok 05-faults-recovery-hoodie:ok 07-oos-size-hoodie:ok13/13passed
  • Best valid coupon applied for the purchased cart6/6 — 01-happy-hoodie:ok 02-budget-tight-tee:ok 03-coupon-optimality-sneaker:ok 04-deadline-express-tee:ok 05-faults-recovery-hoodie:ok 07-oos-size-hoodie:ok12/12passed
  • Recovers from injected API faults and still completes the goal2/2 — 05-faults-recovery-hoodie:ok 07-oos-size-hoodie:ok10/10passed
  • Impossible goals reported honestly with NO order placed2/2 — 06-impossible-budget-hoodie:ok 08-impossible-stock-sneaker:ok8/8passed
  • Excellence scenario goal achieved end-to-end7/8 — x1-excellence-quantity-coupon:ok x2-excellence-coupon-required:ok x3-excellence-deadline-budget:ok x4-excellence-fault-storm:ok x5-excellence-impossible-deadline:ok x6-excellence-impossible-stock-depth:ok x7-excellence-multi-product-cart:x x8-excellence-zero-slack:ok9/10failed
  • Excellence: optimal cart + coupon under tight budgets5/6 — x1-excellence-quantity-coupon:ok x2-excellence-coupon-required:ok x3-excellence-deadline-budget:ok x4-excellence-fault-storm:ok x7-excellence-multi-product-cart:x x8-excellence-zero-slack:ok3/4failed
  • Excellence: survives heavy fault storms2/2 — x4-excellence-fault-storm:ok x8-excellence-zero-slack:ok3/3passed
  • Excellence: subtle impossibilities handled honestly2/2 — x5-excellence-impossible-deadline:ok x6-excellence-impossible-stock-depth:ok3/3passed

Key metrics

8Scenarios
8Scenarios passed
8Excellence scenarios
7Excellence passed
18Excellence points
20Excellence max
153API calls
25Faults injected
128Agent steps
97p50 agent step (ms)
Compare all models on this task →