initial full embed
monthly churn
year 1 total
one full re-index

Annual cost by monthly churn

Churn is the driver most people miss. A corpus that turns over 20% a month costs far more to keep current than the one-time embed suggests — before any model migration.

Monthly churnDocs re-embedded / moYear 1 total

The cost that arrives after launch

Retrieval-augmented generation projects almost always get costed the same way: count the documents, multiply by tokens and the embedding price, and there is your number. That number is real, but it is only the cost of embedding the corpus once. A knowledge base is not a museum — support articles get rewritten, product docs get updated, new content lands every week. Every one of those changes has to be re-embedded, and if a tenth of your corpus changes each month you are quietly paying to re-embed a tenth of everything, month after month. On a large corpus that recurring bill can rival the initial one within a year, and it never shows up in the launch estimate.

Then there is the migration cliff. Embedding vectors from different models are not interchangeable, so the day you move to a newer or cheaper embedding model, you have to re-embed the entire corpus at once — a full re-index. That single event can cost more than a whole year of ordinary churn, which is exactly why teams put off model upgrades they should make. Budgeting for it up front turns a scary surprise into a line item. To keep the ongoing number down, only re-embed what genuinely changed by tracking document hashes, chunk finely so a small edit does not re-embed a whole file, and batch the jobs for any discount. Size the rest of the pipeline with the RAG cost calculator, the embeddings cost calculator and the vector database cost calculator. Prices are editable reference estimates — confirm the current rate in your provider's docs.

Host your project:DigitalOcean — $200 free ↗Hostinger VPS
Crypto ExchangesPaymentsEmail & SMSMarket DataMaps