Engineering

RAG Inference Cost Calculator: Cost Per Token, Broken Down (2026)

Back to BlogWritten by Published Sep 29, 2026
RAG Inference Cost CalculatorRAG Cost Per TokenRAG Cost Per QueryEmbedding Cost RAGReranking Cost Per QueryRAG Pipeline PricingCost Per TokenGPU Cloud
RAG Inference Cost Calculator: Cost Per Token, Broken Down (2026)

Most "how much does RAG cost" posts quote one number: a flat development-cost range for building the application, usually somewhere in five to six figures of engineering time. That number answers a different question than the one a team running RAG in production actually needs answered: what does each query cost once the system is live and serving traffic. This post is a RAG inference cost calculator for that question: it prices embedding, retrieval, reranking, and generation separately, because they don't behave the same way.

TL;DR: What Does the RAG Inference Cost Calculator Show, Stage by Stage?

A RAG query has four cost centers: embedding, retrieval, reranking, and generation. Three are cheap. One isn't.

  • Embedding is near-free: Voyage's voyage-3.5-lite encodes a query for $0.02/M tokens, about $0.0000002 for a 12-token query.
  • Retrieval and reranking round to nothing: Pinecone's serverless minimum is 0.25 read units per query; reranking 20 candidates with rerank-2.5 ($0.05/M tokens) adds roughly $0.0005.
  • Generation is the real bill: RAG typically fetches 3 to 8x more context tokens than the model needs, so the same question costs far more than a plain chat call.
  • Spheron runs the self-hosted side: H100 SXM5 starts at $2.64/hr on-demand as of 01 Oct 2026. See the GPU cost optimization playbook before sizing generation.

Why RAG Cost Per Token Isn't the Same Number as Chat Cost Per Token

A plain chat completion has one cost-per-token number, because there's one call: a system prompt, a user message, an answer. RAG adds a retrieval hop in front of generation, and that hop doesn't just add its own cost, it inflates the cost of the step after it.

Here's the mechanism. A standard RAG configuration retrieving the top 5 chunks at 500 tokens each pulls in 2,500 tokens of context per query. Add an 800-token system prompt and a 100-token question, and total input reaches 3,400 tokens. That's before the model writes a single output token. A chat call answering the same question with no retrieval might run 100 to 200 input tokens total. RAG isn't 20% more expensive than chat on a comparable question, it's often 15 to 20x more expensive on the input side alone, and the multiplier comes entirely from retrieval, not from anything the generation model itself is doing differently.

This is also why a single blended "RAG cost per token" figure is close to meaningless without stating your top-k and chunk size. Two teams running the same generation model at the same price per token will see wildly different bills if one runs top-k=3 at 300-token chunks and the other runs top-k=10 at 800-token chunks. The general cost-per-token methodology (cluster cost divided by throughput) carries over from any inference workload, our GPU FinOps playbook covers that base math in more depth, but RAG's added variable is what goes into the prompt before that formula ever runs.

Embedding Cost: Batch Indexing vs Real-Time Query Encoding

Direct answer: Embedding shows up twice in a RAG pipeline, once when you index your corpus (a large, mostly one-time or periodic cost) and once when you encode each incoming query (a tiny, per-request cost). Conflating the two is the most common mistake in RAG cost estimates.

Query-time embedding is cheap because a query is short. Voyage AI's embedding models price at $0.06 per million tokens for voyage-3.5 and $0.02 per million tokens for voyage-3.5-lite, per Voyage's pricing documentation. A 12-token query encoded with voyage-3.5-lite costs $0.0000002, a number so small it doesn't move a per-query budget at any realistic volume. Cohere's embed-v4 prices higher, at $0.12 per million input text tokens (EmbeddingCost.com's Cohere pricing breakdown), and even at that rate a single query still costs a fraction of a thousandth of a cent.

Batch indexing is where embedding cost actually matters, and it's a corpus-size problem, not a query-volume problem. Indexing 100 million tokens of documents at voyage-3.5's $0.06/M rate costs $6, a one-time (or per-reindex) charge that has nothing to do with how many queries you serve afterward. Self-hosting the embedding step changes this math further: running BGE-M3 or Qwen3-Embedding on Text Embeddings Inference on a rented GPU pushes the marginal cost toward the GPU-hour rate divided by throughput rather than a per-token API fee, which our self-hosted embeddings and rerankers guide works through with real cost-per-million-token figures against the managed APIs. Embedding and reranking are memory-bandwidth-bound workloads rather than compute-bound ones: the bottleneck is how fast the GPU can read vectors and model weights out of memory, not raw FLOPs. On Spheron, an A100 instance is sized for exactly that job without paying for compute headroom the embedding stage never uses.

Retrieval and Reranking: The Hop Everyone Forgets to Price

Direct answer: Retrieval and reranking are usually the cheapest two stages in a RAG pipeline on a per-query basis, and that's exactly why they get skipped in cost estimates. Skipping them isn't wrong on the dollar amount, but it hides where the real cost gets created, since an oversized retrieval and rerank stage is what inflates the generation prompt that follows it.

Pinecone's serverless pricing charges by read units for queries: 1 read unit per 1GB of namespace scanned, with a minimum of 0.25 read units per query, and separately, write units for indexing (1 write unit per 1KB, minimum 5), plus per-GB monthly storage, according to Pinecone's own cost documentation. Pinecone's Standard plan prices those read units at $16 to $18 per million depending on cloud provider and region, per Pinecone's pricing page. A query against a namespace under 1GB, which covers most single-tenant RAG deployments, bills that 0.25-read-unit minimum: at $16 to $18 per million read units, that's roughly $0.000004 to $0.0000045 per query. It doesn't get meaningfully cheaper than that; it's a rounding error at any query volume a single team is likely to run.

Reranking costs more than retrieval but is still small relative to generation. Voyage's rerank-2.5 prices at $0.05 per million processed tokens, where processed tokens equal query tokens multiplied by document count, plus document tokens, per Voyage's pricing documentation (rerank-2.5-lite runs $0.02 per million). A candidate pool of 20 documents at roughly 500 tokens each, reranked against a 12-token query, processes 12 x 20 + 10,000 = 10,240 tokens, costing about $0.00051 at the rerank-2.5 rate. That's real money at billions of queries, but at the query volumes most production RAG systems actually run, it's a fraction of a cent that never shows up as the line item worth optimizing first.

The reason this stage matters for cost anyway isn't its own price tag, it's what it controls downstream. Retrieval accuracy in RAG systems tends to peak around a top-k of 3 and decrease at higher retrieval depths, according to Compresr's guide to RAG over-retrieval costs, which directly contradicts the instinct to retrieve more chunks "to be safe." Most production RAG pipelines fetch 3 to 8x more context tokens than the model needs to answer the query, with an estimated 60 to 80% of retrieved tokens redundant or irrelevant, per the same guide. "Most RAG applications retrieve 3-5x more context than the model meaningfully uses," is how The Token Company puts it. Every one of those extra tokens is billed at generation prices, not retrieval or rerank prices, which is why the fix belongs in this section even though the dollar savings shows up in the next one.

Generation Cost With Retrieved Context: Why Your Prompt Is Bigger Than You Think

This is the stage that actually determines your bill. Using the standard configuration from the sections above, top-k=5 at 500-token chunks plus an 800-token system prompt and a 100-token question, generation input lands at roughly 3,400 tokens before a single output token gets written. A 300-token answer brings the full call to 3,700 tokens. Run that same question with no retrieval at all and a typical chat call might total 300 to 500 tokens end to end. Retrieval didn't just add a step, it multiplied the cost of the step that follows it by roughly 8 to 12x on this configuration alone.

For a managed generation API, that means the input-token price matters more in a RAG pipeline than it would in a plain chat product, because input volume is structurally larger. For self-hosted generation, the same effect shows up as more prefill work per request, which is exactly what continuous batching exists to absorb efficiently rather than let sit idle between decode steps; our continuous batching explainer covers why that scheduling change is the single biggest lever on generation-stage cost per token, independent of hardware. Model choice matters here too: a RAG-specific generation model needs to handle long, citation-heavy context reliably rather than just generate fluent text, which is the case our Cohere Command A self-hosting guide makes for its 256K context window and enterprise RAG tool-calling support over a generic chat model of similar size. Whichever model runs generation, it's the stage worth benchmarking first, on an H100 instance sized for the model you've chosen, since it's the only one of the four stages where cost scales meaningfully with how retrieval-heavy your prompt design actually is.

The RAG Inference Cost Calculator: Formula and Worked Example

The four-stage cost per query is:

Cost per query = Embed(query tokens) + Retrieve(minimum charge) + Rerank(query tokens x candidate count + candidate tokens) + Generate(input tokens + output tokens)

Where embedding and reranking prices come from your chosen provider's per-million-token rate, retrieval is your vector database's per-query charge, and generation is either a managed API's input/output token price or, for self-hosted generation, GPU-hours divided by tokens-per-second throughput.

Plugging in the standard configuration used throughout this post (top-k=5, 500-token chunks, 800-token system prompt, 100-token question, 20-document rerank candidate pool, 300-token output) against the unit prices already cited:

StageTokens processedRate usedCost per query
Embedding (query)12 tokens$0.02/M (voyage-3.5-lite)~$0.0000002
Retrieval1 query, sub-1GB namespace0.25 RU min at $16-18/M RU (Pinecone)~$0.000004-0.0000045
Reranking10,240 processed tokens$0.05/M (rerank-2.5)~$0.00051
Generation3,400 input + 300 output = 3,700 tokensmodel/GPU-dependent, see worked table belowdominant cost

Embedding, retrieval, and reranking together add up to roughly $0.0005 per query, under a tenth of a cent, regardless of which generation model or hardware you pick. Generation is the one line that swings the total by an order of magnitude or more depending on model size, quantization, and whether you're paying a managed API's per-token rate or amortizing a rented GPU's hourly rate across throughput. That's the number worth spending your optimization time on, and it's the one the worked table below prices out on real infrastructure.

Full Stack Cost Per Query on Spheron: A Worked Table

Stage-by-stage unit prices are useful for planning, but a full pipeline benchmark shows how they add up in practice. Our Perplexity Sonar API pricing comparison ran exactly that: a self-hosted RAG stack, Llama 3.3 70B for generation, a BGE-M3 reranker, and an embedding model, running on a single on-demand H100 PCIe instance, cost about $1,447 a month at 100,000 queries a month, working out to $0.0145 per query as of 07 May 2026.

That $0.0145 per query is almost entirely GPU-hours spent on the generation model. The embedding and rerank stages add a rounding error on top of it, consistent with the unit-price math above; the H100 instance is paid for by the hour whether it's running generation, reranking, or sitting between requests, so the per-query figure is really "GPU cost divided by query volume" with generation dominating how many GPU-seconds each query actually consumes.

ComponentApproximate share of the $0.0145/query totalWhy
Embedding (query-time)Effectively $0Query tokens are a handful; cost is sub-thousandth-of-a-cent regardless of GPU or API
RetrievalEffectively $0A vector index lookup, whether self-hosted or Pinecone-style, bills a near-zero minimum per query
RerankingSmall, roughly a tenth of a cent or lessScales with candidate pool size, not query volume
GenerationThe remaining, dominant share70B-class model, prefill inflated by retrieved context, decode for the full output

That per-query figure moves with query volume because the H100's monthly cost is close to fixed; running fewer than 100,000 queries a month against the same instance raises the effective cost per query, and running more lowers it, until you need a second GPU. GPU pricing on Spheron changes with availability, on-demand rates span $0.96/hr on the lighter end up to well beyond that for top-tier accelerators, so treat the $0.0145 figure as a shape to model against, not a number to copy into a budget without checking current rates first.

Pricing fluctuates based on GPU availability. Spheron rates above are live as of 01 Oct 2026; other providers reflect their most recent published rates and may have changed. Check current GPU pricing → for live rates.

For teams that don't want to run any GPU infrastructure at all, a managed embedding and rerank API (Voyage, Cohere) with a managed or API-based generation model removes the ops burden entirely and has zero fixed monthly floor, the right call at low query volume where a 24/7 GPU instance would sit mostly idle. Spheron's pricing page doesn't publish RAG-specific numbers; the stack-level math above is this post's own worked example, not a published Spheron benchmark for this exact configuration.

How to Cut RAG Cost Per Query Without Cutting Answer Quality

Since generation carries the bill, the highest-leverage cuts are the ones that shrink what gets fed into generation, not the ones that shave fractions of a cent off embedding or retrieval.

  • Cut top-k before anything else. Retrieval accuracy peaks around top-k=3 and declines at higher retrieval depths, so dropping from top-k=10 to top-k=3 can improve answer quality while cutting the retrieved-context share of your prompt by more than half.
  • Fix your chunk size with real benchmark data, not intuition. Recursive character splitting at 512 tokens with 50 to 100 token overlap was the benchmark-validated default in the largest real-document RAG chunking test of 2026, scoring 69% accuracy and outperforming pricier alternatives, per PremAI's 2026 chunking benchmark guide. Smaller, over-fragmented chunks from semantic chunking cost more in retrieval count and scored worse; that's a case where the more expensive approach was also the less accurate one.
  • Let the reranker do the filtering, not the generation model. Retrieve a wider recall set cheaply, rerank it down to your final top-k, and only pass the reranked set to generation. The rerank stage costs a fraction of a cent; skipping it and passing a large unranked candidate set straight to the generation model costs the full generation-token price on every token you didn't need.
  • Batch generation requests instead of serving them one at a time. Continuous batching keeps the GPU filled between decode steps rather than idling on the slowest request in a batch, which is what turns a rented GPU's hourly rate into a lower effective cost per token without touching model quality.
  • Cache repeated and near-duplicate queries. A semantic cache absorbing even a moderate share of repeat traffic pushes effective query capacity higher on the same fixed GPU cost, the same lever that pushed the $1,447/month benchmark above toward serving more than 100,000 queries on one instance.

None of these change what the model says. They change how much of the prompt retrieval hands it and how efficiently the GPU processes that prompt, which is where the actual cost lives.


Once you know which stage is driving your bill, size hardware for that stage specifically instead of one GPU for the whole pipeline. Rent an L40S for the memory-bandwidth-bound embedding and rerank stages, or match generation to the model size you've chosen.

Get started on Spheron →

FAQ / 05

Frequently Asked Questions

Per query, most of the cost sits in one stage: generation. Embedding a query costs a fraction of a cent (Voyage's voyage-3.5-lite prices at $0.02 per million tokens), retrieval on a serverless vector database bills a fixed minimum per query that rounds to near zero, and reranking a candidate pool typically adds under a tenth of a cent. Generation is what varies, because retrieval inflates the prompt: a standard top-k=5 setup with 500-token chunks adds roughly 2,500 tokens of context to every call, on top of the system prompt and question. A published full-stack benchmark (embedding, reranking, and a 70B model) on one on-demand H100 came out to $0.0145 per query at 100,000 queries a month. Development cost is a separate number entirely and isn't part of this per-query math.

It depends almost entirely on your generation model, your top-k, and your chunk size, not on the embedding or retrieval stages. Use the formula in this post: (query tokens x embed price) + (retrieval minimum) + (rerank candidate tokens x rerank price) + (total generation input tokens x input price + output tokens x output price, or GPU-hours divided by throughput for self-hosted). For a standard top-k=5 config with 500-token chunks, generation input alone runs to roughly 3,400 tokens before the model writes a single output token.

Generation, by a wide margin, on a per-query basis. Embedding a single query is a few tokens processed through a model priced in fractions of a cent per million tokens. Generation processes the full retrieved context plus the system prompt plus the output, which is typically thousands of tokens per call. Batch indexing (embedding your entire corpus up front) can be a meaningful one-time or recurring cost at large corpus sizes, but real-time query-time embedding is not.

With a hosted reranker like Voyage's rerank-2.5 at $0.05 per million processed tokens, reranking a candidate pool of 20 documents at roughly 500 tokens each, plus the query repeated against each candidate, comes out to a little over $0.0005 per query, well under a tenth of a cent. It scales with candidate pool size, so an unusually large rerank pool (hundreds of candidates) can push this higher, but it rarely approaches what generation costs on the same query.

Cut top-k before you cut anything else. Retrieval accuracy in RAG systems tends to peak around a top-k of 3 and decline at higher retrieval depths, so pulling fewer, better-ranked chunks often improves answers while shrinking the generation prompt that's driving most of your cost. Past that, cache repeated or near-duplicate queries, use a reranker to cut a large recall set down before it reaches the generation model, and batch generation requests so GPU idle time stops padding your cost per token.

Try It Yourself

Try It on Real GPUs

The GPUs behind these guides are the ones you can rent here: H100s, H200s, B200s, and more, billed per minute after a 20-minute minimum runtime, with no contracts. Pick one and you are live in under two minutes.

Deploy Time
< 2 min
Uptime SLA
99.9%
GPU Models
10+
Billing
Per-Min