Opus where it matters. Flash everywhere else.
Built for retail scale.
Keep Opus on pricing, fraud, nuanced recommendations, and final review. Route bounded execution to Flash or Flash-Lite. The result preserves the best model on the hard 15% while outperforming an Opus + Sonnet stack on cost and latency.
Claude Opus 4.8
Pricing, fraud, architecture
$5 / $25 /1M
Gemini 3.5 Flash
Features and multimodal
$1.50 / $9 /1M
Gemini 3.1 Flash-Lite
Lookups and scaffolding
$0.25 / $1.50 /1M
Claude Sonnet 4.6
Conventional executor tier
$3 / $15 /1M
Gemini 3.1 Pro scores 92 on the Artificial Analysis (AA) Intelligence Index vs Claude Opus 4.8 at 89. AA Intelligence Index scores as of Q2 2026. See artificialanalysis.ai for current rankings.
AI-Assisted Development
Three-tier routing across a real e-commerce codebase.
55% of tasks are routine, 30% are mid-complexity, and 15% are hard. Opus stays on the hard 15%; Flash models replace Sonnet on the bounded 85%.
Comparison Baseline
Team Configuration
Opus + Gemini
$4.74K
over 13 sprints · 12 developers
All Opus / Sprint
$863.40
Opus + Sonnet / Sprint
$618.84
Opus + Gemini / Sprint
$364.86
Task Board — 20 tasks
Hover for routing rationale
Total Cost Comparison (13 sprints)
Opus + Gemini vs Opus + Sonnet — 13 sprints
$3.30K
41% reduction · same Opus allocation, cheaper execution tier
In-App Shopping Assistant
Compare single-model deployments with two tiered stacks. Routine and FAQ categories reach parity from Flash-Lite up; complex gifting and styling questions still route to Opus.
Select a question
Select a question above to see the model ladder
Monthly Query Volume
Adjust for your retailer's traffic
Compare tiered routing vs.
Monthly Cost — Full Ladder
Tiered (green) vs each single-model option · selected baseline in amber
Annual Savings vs Opus + Sonnet
$1.8M/yr
Opus + Gemini routing mix
The complex 12% stays on Opus. Flash and Flash-Lite replace Sonnet on the routine 88%.
Visual Search
Shopper photographs an outfit, room, or product — the app returns matching items. Four pipelines compare tiered and single-model choices.
Shopper Input (select to demo)
Matched Products — Summer outfit — shopper photo
Bloom Floral Midi Dress
Classic Canvas Sneakers (White)
Woven Raffia Tote Bag
Sundrop Linen Shirt Dress
Pipeline Race — Tiered + Single Model
Extend · Scenario 4
Nightly Merchandising Agent
Runs only on SKUs needing attention — new items, stockouts, price moves. Compare one single-model deployment with three tiered routes. 16 executor steps per SKU dominate total cost — Opus keeps the pricing call while the executor tier changes.
Per-Run Cost (per SKU)
All 4 strategies compared · selected in green
Scale Simulator
SKUs processed nightly — only those needing attention.
All Opus
$1.7M
/yr
Opus + Sonnet
$1.1M
/yr
Opus + Flash
$747K
/yr
Opus + Flash-Lite
$356K
/yr
Opus + Flash vs Opus + Sonnet (annual)
$364K
33% reduction · 5,000 SKUs × 365 nights
Opus orchestrates. Gemini Flash executes.
16 executor steps per SKU — inventory pulls, demand reads, copy generation, feed validation — are pattern-following work. Routing them to Flash cuts per-SKU cost from $0.6085 (Opus + Sonnet) to $0.4093 (Opus + Flash). At 5,000 SKUs/night that's $364K/yr saved.
Where model economics become business outcomes
Six routing strategies applied to the same ShopOS workload. Cost is calculated from existing token volumes; quality, latency, governance, and value remain inspectable.
Cheapest everywhere is quality-infeasible. Best everywhere wastes spend. Workload-aware routing reaches the efficient frontier; GEAP pushes it further when governance and business value enter the decision.
Lifecycle Summary
ShopOS savings across Build → Run → Extend, at your current slider settings.
Cost by Lifecycle Phase
Per-Scenario Savings
Seller: validate these assumptions per retailer
- Monthly query volume (assistant): 15M/mo is a large retailer — adjust the S2 slider.
- SKUs processed nightly: 5K is mid-market. Catalogue size and change velocity drive this.
- Visual search interactions: 8M/yr assumes active mobile app with camera-search feature.
- Task mix (S1): 55/30/15 routine/mid/complex — adjust for platform maturity and greenfield vs. brownfield work.