Current per-token rates
GPT-4o: $2.50/1M input, $10/1M output. GPT-4 Turbo: $10/$30. GPT-4o-mini: $0.15/$0.60. All rates billed per token, no minimums.
Cached input pricing
GPT-4o cached input is billed at 50% of standard ($1.25/1M). Cache hits trigger automatically on identical prefixes ≥1024 tokens.
Batch API discount
Async jobs submitted via Batch API run at 50% off both input and output. Completes within 24h. Ideal for evals, embeddings backfills, and bulk summarization.
Worked cost example
A support agent: 4K tokens in, 600 out per call, 100K calls/month. GPT-4o ≈ $1,600. GPT-4o-mini ≈ $96. Tier choice changes the bill by 16×.
When each tier fits
GPT-4o: reasoning, tool use, accuracy-critical paths. Turbo: legacy compatibility only. Mini: chat, classification, extraction, anything where mini benchmarks within 5% of flagship.

TokenAtlas