Home
ShopOS — Practical Model Routing

Opus where it matters. Flash everywhere else.
Built for retail scale.

Retail runs at scale. Small routing decisions compound across millions of sessions and SKUs.

Keep Opus on pricing, fraud, nuanced recommendations, and final review. Route bounded execution to Flash or Flash-Lite. The result preserves the best model on the hard 15% while outperforming an Opus + Sonnet stack on cost and latency.

Build
SDLC Coding
Run
Shopping Assistant
Run
Visual Search
Extend
Merch Agent
Keep for judgment

Claude Opus 4.8

Pricing, fraud, architecture

$5 / $25 /1M

Complex (15%)

Gemini 3.5 Flash

Features and multimodal

$1.50 / $9 /1M

Mid (30%)

Gemini 3.1 Flash-Lite

Lookups and scaffolding

$0.25 / $1.50 /1M

Routine (55%)
Baseline

Claude Sonnet 4.6

Conventional executor tier

$3 / $15 /1M

Baseline

Gemini 3.1 Pro scores 92 on the Artificial Analysis (AA) Intelligence Index vs Claude Opus 4.8 at 89. AA Intelligence Index scores as of Q2 2026. See artificialanalysis.ai for current rankings.

Scenario 1Build Phase

AI-Assisted Development

Three-tier routing across a real e-commerce codebase.

55% of tasks are routine, 30% are mid-complexity, and 15% are hard. Opus stays on the hard 15%; Flash models replace Sonnet on the bounded 85%.

Comparison Baseline

Team Configuration

Team Size
12 devs
5 devs50 devs
Sprint Count
13 sprints
1 sprints26 sprints
Tasks/dev/sprint: 40
Total tasks: 6,240

Opus + Gemini

$4.74K

over 13 sprints · 12 developers

All Opus / Sprint

$863.40

Opus + Sonnet / Sprint

$618.84

Opus + Gemini / Sprint

$364.86

Controlled comparison: both mixed routes reserve Opus for search relevance, fraud logic, and agent architecture. Opus + Gemini costs $4.74K versus $8.04K for Opus + Sonnet.

Task Board — 20 tasks

Hover for routing rationale

Design search-relevance & ranking architectureClaude Opus
Build pricing/promotions rules engineClaude Opus
Handle payment & fraud edge casesClaude Opus
Design multi-tenant data isolationClaude Opus
Plan merchandising agent graphClaude Opus
Build cart/checkout flowFlash
Integrate payment + tax providersFlash
Refactor catalog serviceFlash
Wire RAG ingestion for product docsFlash
Scaffold PLP/PDP React componentsFlash-Lite
CRUD for SKU & inventoryFlash-Lite
Write supplier-feed CSV/XML importersFlash-Lite
Generate OpenAPI docs from route handlersFlash-Lite
Unit tests for cart/checkoutFlash-Lite
Image-alt & SEO meta generationFlash-Lite
Seed catalog fixturesFlash-Lite
Write E2E test flows for checkoutFlash-Lite
Add input validation to all form fieldsFlash-Lite
Generate i18n string keys for storefrontFlash-Lite
Add error boundary wrappers to product routesFlash-Lite
5 Opus4 Flash11 Flash-Lite

Total Cost Comparison (13 sprints)

Opus + Gemini vs Opus + Sonnet — 13 sprints

$3.30K

41% reduction · same Opus allocation, cheaper execution tier

Scenario 2 — Run Phase

In-App Shopping Assistant

Compare single-model deployments with two tiered stacks. Routine and FAQ categories reach parity from Flash-Lite up; complex gifting and styling questions still route to Opus.

Ask your shopping assistant

Select a question

Select a question above to see the model ladder

Monthly Query Volume

Adjust for your retailer's traffic

15.0Mqueries/mo
1M50M

Compare tiered routing vs.

Monthly Cost — Full Ladder

Tiered (green) vs each single-model option · selected baseline in amber

Opus + Gemini Baseline

Annual Savings vs Opus + Sonnet

$1.8M/yr

62% reduction15.0M queries/mo
Opus + Gemini: $91K/mo · Opus + Sonnet: $243K/mo

Opus + Gemini routing mix

Flash-Lite 60% — FAQ/order-statusFlash 28% — catalog lookupsOpus 12% — gifting/styling

The complex 12% stays on Opus. Flash and Flash-Lite replace Sonnet on the routine 88%.

Scenario 3 — Run Phase

Visual Search

Shopper photographs an outfit, room, or product — the app returns matching items. Four pipelines compare tiered and single-model choices.

Sub-second visual search is a conversion event — every 100 ms of latency correlates with lower add-to-cart rates.

Shopper Input (select to demo)

Matched Products — Summer outfit — shopper photo

Dresses

Bloom Floral Midi Dress

96%
Footwear

Classic Canvas Sneakers (White)

91%
Bags

Woven Raffia Tote Bag

88%
Dresses

Sundrop Linen Shirt Dress

82%

Pipeline Race — Tiered + Single Model

All Opus3,500 ms
Opus + Sonnet2,700 ms
Opus + Flash ★ Recommended2,200 ms
All Flash700 ms
0 ms3,500 ms

Extend · Scenario 4

Nightly Merchandising Agent

Runs only on SKUs needing attention — new items, stockouts, price moves. Compare one single-model deployment with three tiered routes. 16 executor steps per SKU dominate total cost — Opus keeps the pricing call while the executor tier changes.

Strategy:Per SKU: $0.4093
Agent Execution TraceClick node for details
Planner
$0.0750
Executors × 16
$0.2568 total
Reviewer
$0.0775
Claude Opus 4.8
Claude Sonnet 4.6
Gemini 3.5 Flash
Gemini 3.1 Flash-Lite
Opus plan+review, Flash executors — recommended

Per-Run Cost (per SKU)

All 4 strategies compared · selected in green

Scale Simulator

SKUs processed nightly — only those needing attention.

SKUs / night5,000
50025K50K

All Opus

$1.7M

/yr

Opus + Sonnet

$1.1M

/yr

Opus + Flash

$747K

/yr

Opus + Flash-Lite

$356K

/yr

Opus + Flash vs Opus + Sonnet (annual)

$364K

33% reduction · 5,000 SKUs × 365 nights

Opus orchestrates. Gemini Flash executes.

16 executor steps per SKU — inventory pulls, demand reads, copy generation, feed validation — are pattern-following work. Routing them to Flash cuts per-SKU cost from $0.6085 (Opus + Sonnet) to $0.4093 (Opus + Flash). At 5,000 SKUs/night that's $364K/yr saved.

Optimal Frontier

Where model economics become business outcomes

Six routing strategies applied to the same ShopOS workload. Cost is calculated from existing token volumes; quality, latency, governance, and value remain inspectable.

The feasible frontier excludes routes that ship quality failures. Workload-aware routing preserves quality while avoiding frontier-model spend everywhere.
Cost × Quality
Existing defaults: six-month build plus one year of assistant, visual search, and agent activity
Raw Pareto frontier
Includes the absolute cheapest point even when it fails quality.
Quality-feasible frontier
Only strategies whose final requests all clear the quality bar.
E and F overlap on raw economics
Switch to Governance or Business Value to see GEAP separate.
AQuality risk
Cheapest
$367K
Quality: 4.0
B
Opus only
$6.7M
Quality: 5.0
Dominated by E
C
Claude stack
$4.4M
Quality: 5.0
Dominated by E
DFrontier
Gemini family
$1.4M
Quality: 4.6
EFrontier
Hybrid router
$1.8M
Quality: 5.0
FFrontier
GEAP
$1.8M
Quality: 5.0
The optimal answer is not one model.

Cheapest everywhere is quality-infeasible. Best everywhere wastes spend. Workload-aware routing reaches the efficient frontier; GEAP pushes it further when governance and business value enter the decision.

Lifecycle Summary

ShopOS savings across Build → Run → Extend, at your current slider settings.

Total Annual Savings
$0

Cost by Lifecycle Phase

Per-Scenario Savings

Build
Scroll to activate
Assistant
Scroll to activate
Visual Search
Scroll to activate
Merch Agent
Scroll to activate

Seller: validate these assumptions per retailer

  • Monthly query volume (assistant): 15M/mo is a large retailer — adjust the S2 slider.
  • SKUs processed nightly: 5K is mid-market. Catalogue size and change velocity drive this.
  • Visual search interactions: 8M/yr assumes active mobile app with camera-search feature.
  • Task mix (S1): 55/30/15 routine/mid/complex — adjust for platform maturity and greenfield vs. brownfield work.