Pricing by model
Sonnet 3.5: $3/$15. Opus: $15/$75. Haiku: $0.25/$1.25. Per 1M tokens. Output is 5× input — long generations dominate cost.
Prompt caching: the 90% discount
Cache long system prompts and reference docs. Cached reads are 90% cheaper than fresh input. On a 50K-token system prompt across 1K requests/day, savings are >$1,000/mo.
Batch API
50% off everything, async with 24h SLA. Perfect for bulk classification, periodic summarization, and offline RAG ingestion.
When Opus is worth it
Almost never for production. Sonnet 3.5 matches or exceeds Opus on most benchmarks at 1/5 the cost. Reserve Opus for genuinely hard reasoning.
Frequently asked questions
- Is caching automatic?
- No — explicit cache breakpoints. Easy to add; TokenAtlas estimates savings.
- Bedrock vs direct API price?
- Bedrock matches direct list price; volume discounts vary.
- When to use Haiku?
- Classification, routing, simple Q&A, structured extraction.

TokenAtlas