Cost-plus pricing
Mark up raw AI cost 5–10× for managed services. A $200/mo AI bill becomes a $1,500/mo retainer line item — justified by the integration, prompts and monitoring you provide.
Value-based pricing
Don't bill on tokens. Bill on outcomes: leads generated, hours saved, content shipped. Token cost becomes a margin lever, not a price ceiling.
Client transparency reports
Monthly: tokens consumed, cost, model mix, savings vs baseline. TokenAtlas exports white-label PDF reports on Pro.
Worked example for a retainer client
A content client at 40 pieces/month, 500 input / 900 output tokens each on Claude 3.5 Sonnet, is 20K input and 36K output tokens — modelled cost of 0.02M/1M x $3 + 0.036M/1M x $15 ≈ $0.06 + $0.54 = $0.60/month in raw inference. The retainer markup sits entirely on top of that figure once you decide the multiple.
Cost drivers ranked across a client roster
Across a roster, the number of deliverables per client per month ranks first as the volume driver, and output length per deliverable ranks second. Model choice ranks third — it is the lever you set once per client rather than one that fluctuates month to month, which is why it belongs in the scope-of-work rather than in the monthly variance.
When a cheaper model is a false economy for client work
Switching a client's deliverables to a cheaper model lowers the modelled inference cost, but if the output needs more editing time before it meets the brief, the saved token cost is smaller than the added labor cost you don't bill separately. Model the swap only where quality has been checked to hold, not purely from the rate-card difference.
A second scenario: a ten-client roster
Ten clients at the same 40-piece/month, 500-in/900-out shape on Sonnet each cost roughly $0.60/month in raw inference (as modelled on this page), so the roster's raw AI cost is about $6/month — trivial next to labor and account management cost, and the arithmetic that shows a retainer markup is priced on service, not on token volume, for content workloads at this scale.
Assumptions checklist before setting a retainer price
— Deliverables per client per month and average output length per deliverable — Model assigned per client or per deliverable type — Editing/review hours per deliverable, since this is the labor cost the token estimate excludes — Target markup multiple and whether it's applied per-client or blended across the roster
What this page does not do
This page models raw inference cost per client from volumes you supply; it does not generate an invoice, track actual token consumption per account, or calculate labor hours. Use the modelled inference figure as one input into a retainer price alongside your own time-tracking and account-management cost, which this page does not attempt to estimate.
Frequently asked questions
- Should I pass through AI cost at-cost?
- Only for enterprise clients who demand it. SMB clients want fixed-price.
- How do I prove ROI?
- Baseline against the manual workflow's labor cost. AI almost always wins 5–50×.
- Can I white-label TokenAtlas?
- White-label reports are on the Team plan.

TokenAtlas