Free business calculator

Estimate your RAG cost and realistic savings

Answer plain questions about customers, employees, or documents. The estimator converts business volume into workload, separates context optimization from model switching, and shows both effects.

  • Calculated in your browser
  • No registration
  • Can start from your real bill

01 / Estimator

Start with the situation, not token counts

Quick mode estimates from business volume. If you have traces and exact numbers, open Advanced mode below.

Where are you now?
011. Describe volume in business terms
022. Choose models and known spend
033. Set a realistic RAG optimization range
10%60%

Bounded to 10–60%. System/history/retrieval shrink. User query, model output, and fixed infrastructure do not.

Advanced mode: I know tokens and prices

Exact calculation from traces and invoice

The engineering mode remains here: two scenarios, prompt caching, embeddings, and vector/reranker cost.

Traffic and pricing
AA. Current system
BB. After optimization
Savings / month$0
Savings / year$0
MetricAB
Input / request00
Retrieval / request00
Cached / request00
LLM input / month$0$0
LLM output / month$0$0
Embeddings / month$0$0
Vector and reranker / month$0$0
Total / month$0$0

Presets checked July 30, 2026 against official provider pages. Region, batch, long context, discounts, and marketplaces can change the rate.

02 / Methodology

No magic behind the slider

Quick mode turns business volume into request count and uses a disclosed token envelope. It estimates variable run cost, not project delivery cost.

01

Business volume

volume → requests / month

Customers × conversations × answers; employees × questions × days; or documents × questions × passes.

02

RAG optimization

(system + history + retrieval) × (1 − rate)

The slider reduces system, history, and retrieval by 10–60%. Query and output stay unchanged.

03

If RAG is already live

forecast × invoice / modeled baseline

The actual model/API bill normalizes the workload model. The baseline therefore matches the invoice, not our guess.

04

Model switching

optimized workload × target rates + fixed infra

The same optimized workload is repriced on the selected model. Fixed infrastructure carries over unchanged.

Excluded from the quick estimate

Engineering, observability, OCR, web search, tool calls, source-document storage, taxes, and one-time migrations. Exact embeddings, prompt cache, and infrastructure changes remain available in Advanced mode.

Frequently asked questions

Where do token counts come from if I do not enter them?

From the visible answer-depth assumption: brief, working, or deep. Once traces exist, replace that assumption with exact inputs in Advanced mode.

Why does a 60% slider not cut my entire bill by 60%?

Because only controllable context is optimized. Output and fixed infrastructure do not disappear when top-k or conversation history shrinks.

Can I simply switch to DeepSeek, GLM, or Kimi?

Only after use-case evals. Price shows economic potential, not tool-calling compatibility, response format, latency, safety, or answer quality.

What if my current bill contains several models?

Choose the primary model and enter the complete model/API bill as the baseline. Use Advanced mode or a trace audit for precise routing-level analysis.

Next step

Is the bill high or the estimate still too rough?

We inspect traces, retrieval, routing, prompt caching, and evals, then identify savings that can be verified without blindly degrading answers.

    Protected by reCAPTCHA. The Google Privacy Policy and Terms of Service apply.