Your AI estate is 54% efficient — $2.4M of annualized spend sits behind the Pareto frontier. The same workloads, at the same quality floors and latency SLAs, could run for $228K/mo instead of $426K/mo.
Sequential attribution per workload: capacity → caching → batching → model substitution. Routing and prompt-reduction upside (app-layer work) is listed separately under Opportunities.
| Recommendation | Workload | Savings/mo | Annualized | Tier |
|---|---|---|---|---|
Route easy traffic to DeepSeek V3.2 100% of requests on GPT-5.1 → ~60% routed to DeepSeek V3.2, escalation on low confidence | Enterprise Knowledge RAG Data Platform | $33.3K | $400K | ◐Near Frontier |
Migrate to GPT-5.1 Claude Sonnet 4.5 · Anthropic → GPT-5.1 · OpenAI | Engineering Code Assistant Engineering | $29.5K | $354K | ◔Optimization Opportunity |
Migrate to DeepSeek V3.2 GPT-4o · Azure → DeepSeek V3.2 · Azure | Global Support Copilot Customer Support | $26.2K | $314K | ○Severely Inefficient |
Trim prompt & retrieved context 6.8K input tokens/request, 18% context utilization → ~5.3K tokens/request (dedupe boilerplate, rerank retrieval) | Enterprise Knowledge RAG Data Platform | $14.3K | $172K | ◐Near Frontier |
Migrate to GPT-5.1 Claude Sonnet 4.5 · Anthropic → GPT-5.1 · OpenAI | Agentic Procurement Assistant Supply Chain | $10.4K | $124K | ◔Optimization Opportunity |
| Workload | Model · Cloud | Tokens/mo | Cache hit | Batch | TTFT | $/1M eff. | $/task | Cost/mo | 30d trend | Gap/mo | Pareto tier |
|---|---|---|---|---|---|---|---|---|---|---|---|
Enterprise Knowledge RAG Data Platform · RAG / grounding | Gemini 2.5 Pro Google Cloud · us-central1 | 87B 82B in · 5.0B out | 48% | 0% | 680ms | $1.10 | $0.0089 | $95.2K | $23.8K | ◐Near Frontier | |
Engineering Code Assistant Engineering · Code generation | Claude Sonnet 4.5 Anthropic · us-east-1 | 24B 20B in · 3.5B out | 68% | 0% | 640ms | $2.83 | $0.0190 | $67.3K | $31.2K | ◔Optimization Opportunity | |
Agentic Procurement Assistant Supply Chain · Agentic workflow | Claude Sonnet 4.5 Anthropic · us-east-1 | 18B 17B in · 1.2B out | 59% | 0% | 640ms | $2.02 | $0.1013 | $36.9K | $20.6K | ◔Optimization Opportunity | |
Global Support Copilot Customer Support · Conversational assistant | GPT-4o Azure · eastus2 | 11B 9.4B in · 1.4B out | 24% | 0% | 460ms | $3.19 | $0.0109 | $34.2K | $31.4K | ○Severely Inefficient | |
Multilingual Translation Hub Operations · Translation | Gemini 2.5 Flash Google Cloud · europe-west1 | 28B 14B in · 14B out | 2% | 50% | 340ms | $0.84 | $0.0014 | $23.5K | $2.4K | ●Frontier Efficient | |
Fraud Alert Narratives Risk & Compliance · Analysis & drafting | GPT-5.1 Azure · eastus2 | 9.3B 8.6B in · 702M out | 26% | 0% | 560ms | $2.39 | $0.0087 | $22.3K | $10.2K | ◔Optimization Opportunity | |
Contract Intelligence Legal · Document analysis | Claude Opus 4.5 Anthropic · us-east-1 | 2.8B 2.5B in · 288M out | 4% | 10% | 980ms | $6.55 | $0.1086 | $18.4K | $14.8K | ○Severely Inefficient | |
Data Quality Rules Copilot Data Platform · Code generation | GPT-5.1 Azure · eastus2 | 3.7B 2.8B in · 864M out | 21% | 0% | 560ms | $3.98 | $0.0150 | $14.6K | $5.6K | ◔Optimization Opportunity |