A tier you can buy is a fact about the provider.
Every provider sells inference in its own set of tiers. This matrix says which tiers each one sells, in the provider's own words, with the page it comes from and the date we read it. It does not price anything. Per-model prices for each tier live in the catalogue and a tier the provider does not publish for a model is not offered, not estimated.
Read in a browser
A person opened the cited page on the date shown and the cell says what the page says. Both the page and the date are on the cell.
Imported, not re-read
Carried from our 2026-08 research snapshot. Marked † until the weekly check re-reads the page or files a change.
Not checked
Nobody has looked yet. Shown as a dashed cell, never as "No": not checked and not offered are different facts.
| Provider | StandardPay per token, on demand | PriorityPer-token premium for tighter latency | BatchAsync queue, hours | FlexOnline, yields to standard traffic | Provisioned throughputReserved capacity per unit-hour | Scale tierCommitted throughput per unit-hour | Prompt cachingCache write and read rates | In-geo / data residencyPriced region or zone pinning | Dedicated / reservedSingle-tenant capacity | Raw GPU rentalHardware, not a model API | Fine-tuned hostingServing your tuned weights |
|---|---|---|---|---|---|---|---|---|---|---|---|
| hyperscalers | |||||||||||
| Microsoft Azure | |||||||||||
| Google Vertex AI | |||||||||||
| AWS Bedrock | |||||||||||
| Alibaba Model Studio | |||||||||||
| frontier labs | |||||||||||
| Anthropic | |||||||||||
| OpenAI | |||||||||||
| Google AI Studio | |||||||||||
| Mistral | |||||||||||
| Cohere | |||||||||||
| DeepSeek | |||||||||||
| xAI | |||||||||||
| Moonshot AI | |||||||||||
| Z.ai | |||||||||||
| data clouds | |||||||||||
| Databricks | |||||||||||
| Snowflake Cortex | |||||||||||
| neoclouds | |||||||||||
| Together AI | |||||||||||
| Fireworks AI | |||||||||||
| Groq | |||||||||||
| Lambda Labs | |||||||||||
Select any cell. The detail shows what the provider calls the tier, the page it comes from, and the date we read it.
Same word, different arithmetic.
- Priority always means a per-token premium paid as you go. Scale tieralways means committed capacity billed per unit-hour. A provider's own name for a tier is kept in the cell note; the column is chosen by how it bills.
- Prompt caching is a pair of rates, write and read, and where a provider sells several tiers each tier carries its own pair. One cell here means the provider publishes them at all.
- In-geocovers data-zone and regional pinning. Where the provider lists a price it is a price; where it states an uplift in a footnote it is recorded as the provider's rule with the page cited. No constant multiplier is assumed across providers or over time.
- Dedicated and raw GPU rental are listed separately because neoclouds in particular blur the line between a managed model API and renting the hardware under it.
A weekly check re-reads each cited page and either moves the cell's date or files a watchlist row. A cell's date is the claim; a cell without one is labelled as such. Rows are the sellers in our catalogue. Sellers we do not yet list are not rows.
// per-model prices for each tier: model explorer · how prices are read and dated: #benchmarks