Skip to main content
TokenAtlas

For agencies

Agencies — AI budgeting per client engagement with TokenAtlas

AI budgeting per client engagement without spreadsheets, surprise bills, or guessing which model to ship.

The problem

Agencies consistently report the same pain: AI costs are growing faster than revenue, the engineering team can't explain the bill to finance, and there's no clear answer when leadership asks "what does AI cost us per active user?"

What changes with TokenAtlas

Agencies price work before they deliver it, so the AI cost has to be known at proposal time. TokenAtlas models a client engagement from its token volumes — documents processed, prompts per user, expected monthly runs — and returns a per-engagement cost you can put into a scope of work with the assumptions written down.

The typical workflow

Build one scenario per engagement rather than one for the agency. Enter the workload the client is actually buying, model it on the model you intend to ship, then add a headroom case at two or three times the volume for the overrun clause. Save the scenario against the account so the estimate can be re-run when the scope changes.

What you can model

Modelling runs off volumes you enter, not connected client accounts — which matters when you do not control the client's provider billing. What you get is a defensible cost floor per engagement, the margin implied by your quoted retainer, and a clear view of which deliverable in the scope carries most of the inference cost.

Why it matters for agencies

Fixed-fee AI work fails on unmodelled token volume, not on delivery. A per-engagement cost model, re-run whenever scope moves, keeps the retainer priced against arithmetic instead of a guess, and gives the client a rate-card explanation when volumes change mid-project.

Worked example: a document-review engagement

A client engagement processing 500 documents/month at 6,000 input / 400 output tokens each is 3M input and 0.2M output tokens monthly. On Claude 3.5 Sonnet: 3M/1M x $3 = $9, plus 0.2M/1M x $15 = $3, for $12/month in modelled inference cost — the figure to build the scope-of-work margin on top of, before adding your delivery time.

Cost drivers ranked for client work

Document count or call volume ranks first since it is usually set by the client's business, not by you. Document length ranks second and is worth confirming at scoping time rather than assuming. Output length ranks third and is the one lever you can often control by asking for shorter structured outputs instead of free-form prose, which lowers cost without changing the deliverable's usefulness.

Headroom scenarios for the overrun clause

Model the quoted volume, then re-run the same scenario at 2x and 3x volume to see what an overrun actually costs before you write the clause. If a client's volume triples and the modelled cost only grows by a proportionally small amount because most of the cost sits in a fixed per-document prompt, that is useful to know before, not after, the engagement runs hot.

Frequently asked questions

Is TokenAtlas appropriate for our team size?
Yes — pricing scales from solo founders on the free tier to enterprise teams on custom contracts.
How long does setup take?
There is nothing to connect. You model a workload from token volumes you already know — prompt size, response size and call volume — so a first estimate takes a few minutes.
Do you support our LLM provider?
Published rate cards for OpenAI, Anthropic, Google, Mistral, Groq, DeepSeek, Cohere and Perplexity are modelled in the catalogue, so you can price the same workload on any of them.

Related