Free business calculator
Estimate your RAG cost and realistic savings
Answer plain questions about customers, employees, or documents. The estimator converts business volume into workload, separates context optimization from model switching, and shows both effects.
- Calculated in your browser
- No registration
- Can start from your real bill
01 / Estimator
Start with the situation, not token counts
Quick mode estimates from business volume. If you have traces and exact numbers, open Advanced mode below.
02 / Methodology
No magic behind the slider
Quick mode turns business volume into request count and uses a disclosed token envelope. It estimates variable run cost, not project delivery cost.
Business volume
volume → requests / monthCustomers × conversations × answers; employees × questions × days; or documents × questions × passes.
RAG optimization
(system + history + retrieval) × (1 − rate)The slider reduces system, history, and retrieval by 10–60%. Query and output stay unchanged.
If RAG is already live
forecast × invoice / modeled baselineThe actual model/API bill normalizes the workload model. The baseline therefore matches the invoice, not our guess.
Model switching
optimized workload × target rates + fixed infraThe same optimized workload is repriced on the selected model. Fixed infrastructure carries over unchanged.
Excluded from the quick estimate
Engineering, observability, OCR, web search, tool calls, source-document storage, taxes, and one-time migrations. Exact embeddings, prompt cache, and infrastructure changes remain available in Advanced mode.
Frequently asked questions
Where do token counts come from if I do not enter them?
From the visible answer-depth assumption: brief, working, or deep. Once traces exist, replace that assumption with exact inputs in Advanced mode.
Why does a 60% slider not cut my entire bill by 60%?
Because only controllable context is optimized. Output and fixed infrastructure do not disappear when top-k or conversation history shrinks.
Can I simply switch to DeepSeek, GLM, or Kimi?
Only after use-case evals. Price shows economic potential, not tool-calling compatibility, response format, latency, safety, or answer quality.
What if my current bill contains several models?
Choose the primary model and enter the complete model/API bill as the baseline. Use Advanced mode or a trace audit for precise routing-level analysis.
Next step
Is the bill high or the estimate still too rough?
We inspect traces, retrieval, routing, prompt caching, and evals, then identify savings that can be verified without blindly degrading answers.