Every number on your quote is one of four things.
Something you told us. Arithmetic we ran. A figure a named third party published. Or an assumption we made, badged as one. This page shows which is which, and how to check us.
1.2M requests / month. You typed it, so it carries a grey dot.
you set thisReserved units = ceil(peak ÷ published rate). Printed below.
ƒ formulaEpoch AI, GPQA Diamond, 78.1. Named, dated, unmodified.
‡ citedCache hit rate 72%, range 58–84%. We defaulted it. It widens the corridor.
† we defaulted thisA preset decides what you are walked through. It does not decide what exists.
Picking "RAG support copilot" fills six underlying dimensions with defaults. Every one of them stays yours to override without leaving the preset. The quote is computed from the dimensions, never from the label.
- A grey dot beside a number means you set it. A red dot means we defaulted it and you have not confirmed it yet.
- Hybrids are normal. A coding agent presents as a code assistant and behaves as an agent on calls per task. Override that one dimension. Keep the preset.
- Nothing you type leaves your browser until you choose to save. The free quote runs the same way.
There is exactly one calculator.
The free quick pass and the full Studio pass run the same engine. At the same inputs they produce the same likely figure. Only the corridor width differs, because only the amount you told us differs.
reserved units = ceil( peak(effective) ÷ published rate ) → raised to the vendor minimum
gpu $/Mtok(out) = $/gpu·hr ÷ ( max decode tok/hr × utilization )
preamble tokens = (scaffold + schema tokens) × turns × requests
units[i] = units[i−1] × pass rate[i−1] → chains attenuate only
quote = Σ phase cost + one badged overhead row → nothing hidden
// every rate in these formulas is a published list price or a rate you negotiated. The utilization and peak factors are ours. They are badged as assumptions in register 04, and their sources are cited, but their values are not printed here.
Attributed, dated, unmodified.
When a number on screen is someone else's measurement, we show whose, from when, and the exact model version. We never roll third-party scores into a composite of our own.
Quality · Epoch AI
Named benchmarks, not indices. A blank cell means the model was not run, not that it scored zero. 92 of 162 cells are filled, and we show that.
CC BY 4.0 · read 2026-08-28Throughput · three sources, three jobs
vLLM CI is the lower bound, a stock stack in week one. InferenceMAX is the default, a nightly run. MLPerf is the upper bound, a tuned submission.
Prices · the vendors' own pages
Every rate shows the page's own date and the date we read it. An undated vendor page is a weaker citation, and you can see which you have.
‡ published · read on// as of September 2026, only 3 of 19 providers publish a pricing API. The other 16 are read from their pages and carry a freshness badge: unchecked, current, aging, stale. Unchecked is neutral, not red. A never-checked rate is not a decayed one.
Every default carries six fields. You can read all of them.
A default is not a guess we hid inside the engine. It is a value with a range, a one-line basis, a source class, a confidence, and a citation you can follow.
| parameter | retrieved_k |
| value | 6 chunks · range 3 – 10 |
| basis | Common RAG top-k. |
| source | engineering default |
| confidence | medium |
| editable | Yes. Change it in Studio and the dot turns grey, the range collapses to your value, and the corridor narrows. |
Not yet sourced
Three workload types have sizing defaults with no external citation. We say so on the card instead of borrowing a basis. Content generation. Document multimodal. Semantic search.
Where the traffic anchors come from
Token sizes for conversation and code workloads are anchored to Microsoft's public Azure LLM inference traces, released under CC-BY. Daily load shapes follow the trace characterizations in the Splitwise and DynamoLLM papers. Request volumes are founder estimates until real quotes replace them, and each of those rows says so. Agentic call counts follow the vendors' own agent guides.
| default | anchor |
|---|---|
| conversation tokens in / out | azure trace |
| code completion tokens in / out | azure trace |
| daily load shape by workload | splitwise · dynamollm |
| requests per day | founder estimate |
| agentic calls per task | anthropic agents guide · agentbench |
An honest range beats a confident guess.
Every quote ships with a corridor. Its width comes from how many of the inputs are yours and how many are ours. Answer more, and it tightens. Declaring confidence does nothing.
- The likely figure is the deterministic run at the values on screen. It is never a median of the simulation.
- Best is the 10th percentile and worst the 90th, from a seeded simulation. The corridor never moves on reload.
- A value you authored gets a narrow band. A value we defaulted gets the wide band for that parameter class. Mode has no say.
- The contributors list ranks which defaults widen it most. Answering those first is the fastest way to tighten it.
The audit pack never re-prices anything.
A pack that recomputed the quote from the same inputs would agree with a broken engine. Ours cannot. It sums the lines the engine emitted and compares them with a total from a different code path. Then it prints the difference on the cover.
Corridors are never paywalled
Every tier produces the same likely figure and a real corridor. Paid tiers add depth, memory and the market view, not honesty.
Sample sizes are always shown
No quartile without its n. Benchmarks report the resolution the data supports and nothing finer.
Open questions ship as stated assumptions
Every red dot becomes a row in the assumption ledger. Blocking rows hold the save. Nothing is deleted, only resolved.
- A "recommended" or "best value" badge.
- A composite quality-per-dollar score we authored.
- A similarity or substitution percentage between models.
- A default ordering that is not a third party's published index or plain arithmetic on your inputs.
- We do not call the benchmark pool differentially private. It is anonymized, floored, and dominance-checked. Those are different claims.
- We do not call the audit pack anonymous. It carries the name you typed into it.
- We do not call the minimum sample size a guarantee. It is a floor, described as a floor.
Both dates, always.
The date a page states and the date we read it are different claims. A page dated two years before we read it is a weaker citation than one dated last week, and you should see which you have.
| publisher · title | supports | dates |
|---|---|---|
| Epoch AI · AI Benchmarking Hub | Quality benchmarks, per named benchmark and exact model version. CC BY 4.0. | snapshot 2026-08-28 |
| InferenceMAX (SemiAnalysis) · nightly open benchmarks | Default GPU throughput curves. vLLM CI perf dashboard is the lower bound and MLPerf Inference (MLCommons) the upper. | snapshot 2026-07-15 |
| Microsoft · Right-size your PTU deployment and save big | Reserved capacity need not cover the peak. Microsoft's own illustration routes 10% of traffic to pay-as-you-go, and names the peak-to-valley distance as the decision. | page dated 2024-02-12 · read 2026-09-02 |
| Microsoft · Manage traffic with spillover for provisioned deployments | Overflow from a provisioned deployment routes to a STANDARD deployment and is billed per token; provisioned capacity is served first. | page dated 2026-06-18 · read 2026-09-02 |
| Microsoft · Best Practice Guidance for PTU | The spillover pattern, batch scheduling into idle provisioned hours, and semantic caching are the three cost levers Microsoft names for provisioned deployments. | page dated 2024-05-28 · read 2026-09-02 |
| Microsoft · Managing Traffic Jams with Azure OpenAI PTU Spillover | Spillover carries no surcharge, and its risk is that the shared tier may itself be congested — an overflow share buys availability, not latency. | page dated 2025-03-20 · read 2026-09-02 |
| Google Cloud · Calculate Provisioned Throughput requirements | Reserved capacity is sized in input-equivalent tokens, with published burndown rates converting output and cached tokens. This is the formula our unit counts use. | page dated 2026-09-02 · read 2026-09-02 |
| Amazon Web Services · Increase model invocation capacity with Provisioned Throughput in Amazon Bedrock | Per-Model-Unit throughput is not published — AWS directs you to your account manager — which is why a Bedrock reservation cannot be sized here. | page undated · read 2026-09-02 |
| Anthropic · Prompt caching | Cache write tokens cost 1.25 times the base input price (5-minute) or 2 times (1-hour); cache read tokens cost 0.1 times, with Claude Fable 5.1 and Claude Mythos 5.1 at 0.025 times. | page undated · read 2026-09-12 |
| Anthropic · Claude Developer Platform documentation | RAG, multi-turn and agentic worked examples whose prompt structures and token counts the sizing defaults were derived from. | page undated · read 2026-09-12 |
| Andreessen Horowitz · Welcome to LLMflation – LLM inference cost is going down fast | For an LLM of equivalent performance, inference cost has been falling about 10x a year, which is why a quoted rate is a point in time and not a forecast. | page dated 2024-11-12 · read 2026-09-12 |
| Andreessen Horowitz · How 100 Enterprise CIOs Are Building and Buying Gen AI in 2025 | A survey of 100 CIOs across 15 industries (as of 2025-05-08) on adoption by use case, used only as a proxy for engineering-org size in the code-assistant default. | page dated 2025-06-10 · read 2026-09-12 |
| GitHub · How GitHub Copilot is getting better at understanding your code | Completion models process about 6,000 characters of context; fill-in-the-middle raised acceptance 10% and neighboring tabs a further 5%. Sizes the code-assistant input default. | page dated 2023-05-17 · read 2026-09-12 |
| Microsoft · Enable priority processing for Microsoft Foundry Models | Priority Processing is a pay-as-you-go service tier; requests over 272k prompt tokens on the 5.6 models are downgraded to standard and billed at standard; the flex fallback is retired on 2026-09-25. | page dated 2026-07-13 · read 2026-09-12 |
| Databricks · Foundation Model Serving | Regional processing applies a stated 10% uplift to DBU rates on labelled models; priority pay-per-token and provisioned throughput are listed beside standard pay-per-token. | page undated · read 2026-09-12 |
| Snowflake · Snowflake editions | The editions are Standard, Enterprise, Business Critical and Virtual Private Snowflake. Edition governs Platform Credit rates, not the per-token AI Credit charge. | page undated · read 2026-09-12 |