Home
WealthAI — Practical Model Routing

Keep Opus. Route the rest.
Across the whole lifecycle.

Opus is the strongest choice for the work that genuinely needs deep judgment. The mistake is paying for that level on every task. Pair Opus with Gemini Flash and Flash-Lite for routine execution, and the stack is cheaper and faster than pairing Opus with Sonnet.

Build
AI-Assisted Dev
Run
In-App Intelligence
Run
Multimodal & Latency
Extend
Agentic Orchestration
Recommended anchor
Keep for the hard work

Claude Opus 4.8

Architecture, judgment, final review

$5 / $25 per 1M

Replace Sonnet here

Gemini 3.5 Flash

Fast execution and multimodal work

$1.50 / $9 per 1M

Cheapest bounded work

Gemini 3.1 Flash-Lite

High-volume lookups and transforms

$0.25 / $1.50 per 1M

Baseline
Comparison baseline

Claude Sonnet 4.6

The conventional mid-tier pairing

$3 / $15 per 1M

Gemini 3.1 Pro scores 92 on the Artificial Analysis (AA) Intelligence Index vs Claude Opus 4.8 at 89. AA Intelligence Index scores as of Q2 2026. See artificialanalysis.ai for current rankings.

Scenario 1Build Phase

AI-Assisted Software Development

Your engineers don't need frontier intelligence for every task.

75% of coding tasks are routine — tests, scaffolds, docs, CRUD. Keep Opus on architecture and hard debugging. Route the routine 75% to Flash instead of Sonnet.

Routing Strategy

Team Configuration

Team Size
12 devs
5 devs50 devs
Sprint Count
13 sprints
1 sprints26 sprints
Tasks per dev per sprint50
Total tasks over period7,800

Opus + Gemini Cost

$0.00

over 13 sprints · 12 developers

You save

$0.00

0% less

Opus + Sonnet / Sprint

$256.50

Sonnet routine + Opus complex

Opus + Gemini / Sprint

$202.50

Flash routine + Opus complex

Task Board — 20 tasks

Hover for routing rationale

Design RAG architecture for fund-prospectus searchOpus
Write compliance rules engine for SEC 17a-4Opus
Debug portfolio dashboard state race conditionOpus
Design multi-account aggregation schemaOpus
Plan agent orchestration graph for rebalancingOpus
Generate unit tests for 40 API endpointsFlash
Scaffold 15 React components from Figma specFlash
Write CRUD for client profilesFlash
Generate OpenAPI docs from route handlersFlash
Write CSV statement import transformsFlash
Add input validation to all form fieldsFlash
Write seed/fixture data for 12 account typesFlash
Create Storybook stories for UI componentsFlash
Write SQL migration for positions tableFlash
Add error boundary wrappers to all routesFlash
Generate i18n string keys for settings pagesFlash
Write Dockerfile and docker-compose for local devFlash
Add TypeScript strict-mode fixes to utils/Flash
Write E2E test flows for onboarding wizardFlash
Convert REST endpoints to tRPC proceduresFlash
15 Flash·5 Opus

Task Mix

Routine (15) → Flash
Complex (5) → Opus

Cost Comparison

Opus + SonnetOpus + GeminiAll Opus

Model Pricing (per 1M tokens)

Claude Opus 4.8$5 in / $25 out
Gemini 3.5 Flash$1.5 in / $9 out
Claude Sonnet 4.6$3 in / $15 out

Bottom line: A 12-person team saves $702.00 over 13 sprints vs Opus + Sonnet. Opus handles the same complex work in both routes; Flash replaces Sonnet only on routine tasks.

Scenario 2 — Run Phase

In-App Intelligence

Same question. One answer costs 10–20× less. Can you tell?

Ask your portfolio

Select a question

Select a question above to compare model responses

Monthly Query Volume

Adjust to match your expected in-app traffic

5.0Mqueries/mo
100K10M

Monthly COGS Comparison

Opus + Sonnet

$0.000000

$778K/yr

Opus + Gemini

$0.000000

$294K/yr

Tiered routing mix

Flash-Lite 60%Flash 28%Opus 12%
Annual Savings
$0.000000/yr
62.1% reductionvs. Opus + Sonnet at 5.0M queries/mo

Monthly: $40K saved  ·  Opus + Gemini: $25K/mo vs Opus + Sonnet: $65K/mo

Honesty note: The same 12% of complex planning questions route to Opus in both architectures. The difference is the routine 88%: Sonnet in the baseline, Flash-Lite and Flash in the recommended route.

Scenario 3 — Run Phase

Multimodal & Latency

Client photographs a statement and asks “what changed?” — keeping Opus for financial reasoning while comparing Sonnet and Flash on extraction.

Incoming Statements

Pipeline Race — Text Mode

All Opus3,800 ms target
Opus OCR
Opus Analysis
Opus Structured
Opus + Sonnet3,200 ms target
Sonnet Vision
Opus Analysis
Opus Structured
Opus + Flash2,700 ms target
Flash Vision
Opus Analysis
Opus Structured
0 ms3,800 ms

Extend · Scenario 4

Agentic Orchestration

Every night an agent reviews 50,000 portfolios. It decomposes each rebalancing job into 18 steps—only the planner and reviewer run on Opus. The 16 executor steps run on Gemini Flash.

Strategy:
Per run: $0.3925
Agent Execution TraceClick any node for details
Planner
$0.0750
Executors × 16
$0.2400 total
Reviewer
$0.0775
Claude Opus 4.8
Gemini 3.5 Flash (Fast)
Opus plans and reviews; Flash executes

Per-Run Cost Breakdown

Executor volume dominates total cost — this is where tiering shines.

Scale Simulator

Drag to set nightly portfolio volume. Costs scale linearly.

Portfolios / night50,000
1K25K50K75K100K

Nightly (Opus + Sonnet)

$29K

Nightly (Opus + Flash)

$20K

Annual Totals50,000 × 365 nights

Opus + Sonnet

$10.7M

Opus + Flash

$7.2M

Annual Savings

$3.5M

33% reduction

Opus orchestrates. Gemini Flash executes.

Only 2 of 18 agent steps require top-tier reasoning — the planner that decomposes the rebalancing task and the compliance reviewer that signs off. Opus handles both in either architecture. The 16 executor steps (data retrieval, calculations, formatting) route to Gemini Flash instead of Sonnet. Total per-run cost drops from $0.5845 (Opus + Sonnet) to $0.3925. At 50K portfolios nightly, that's $3.5M/yr saved.

Optimal Frontier

Where model economics become business outcomes

Six routing strategies applied to the same WealthAI workload. Cost is calculated from existing token volumes; quality, latency, governance, and value remain inspectable.

The feasible frontier excludes routes that ship quality failures. Workload-aware routing preserves quality while avoiding frontier-model spend everywhere.
Cost × Quality
Existing defaults: six-month build plus one year of run and agent activity
Raw Pareto frontier
Includes the absolute cheapest point even when it fails quality.
Quality-feasible frontier
Only strategies whose final requests all clear the quality bar.
E and F overlap on raw economics
Switch to Governance or Business Value to see GEAP separate.
AQuality risk
Cheapest
$960K
Quality: 4.5
B
Opus only
$17.3M
Quality: 5.0
Dominated by E
C
Claude stack
$11.6M
Quality: 5.0
Dominated by E
DFrontier
Gemini family
$3.2M
Quality: 4.8
EFrontier
Hybrid router
$3.9M
Quality: 5.0
FFrontier
GEAP
$3.9M
Quality: 5.0
The optimal answer is not one model.

Cheapest everywhere is quality-infeasible. Best everywhere wastes spend. Workload-aware routing reaches the efficient frontier; GEAP pushes it further when governance and business value enter the decision.

Lifecycle Summary

Accumulated savings across Build → Run → Extend, driven by your current slider configurations.

Total Annual Savings
$0

Cost by Lifecycle Phase

Total Annual Comparison

Build (SDLC)
Scroll to activate
Run: In-App
Scroll to activate
Run: Multimodal
Scroll to activate
Extend: Agent
Scroll to activate