TokenAtlas

Embedding Models Pricing

Embeddings are cheap individually and expensive at scale. RAG pipelines with ongoing re-embedding can run $1K+/month before you notice.

Current per-token rates

OpenAI text-embedding-3-small: $0.02/1M. text-embedding-3-large: $0.13/1M. Cohere embed-v3: $0.10/1M. Voyage 3-large: $0.18/1M. Gemini text-embedding-004: $0.025/1M.

Cost per million vectors

At an average chunk of 500 tokens, 1M vectors ≈ 500M tokens. With OpenAI small, that's ~$10. With Voyage large, ~$90. Pick by recall need, not vibes.

Re-embedding cost

Most teams forget that schema changes, chunking changes, and model upgrades require full re-embed. Budget 2–4× initial embedding cost over a year.

Storage and query cost

Vector DB (Pinecone, Weaviate, pgvector) cost often exceeds embedding cost. A 10M-vector index runs $200–$800/mo across hosted providers.

Picking the right embedder

For most RAG: OpenAI 3-small or Gemini 004 — cheapest with 95th-percentile recall. For high-stakes or domain-specific: Voyage 3-large.

Related