// how it works · workload assumptions & provenance

A preset is six dials, not a category.

The six workload types you pick from are archetypes over six underlying cost dimensions. Studio keys the estimate off the dimensions and treats the type as a preset, so a hybrid that fits no bin degrades gracefully: you override one dial and keep the preset. This page shows the dials, the defaults each preset sets, and for every default whether it rests on a published source, a derivation, or an estimate we badged as one.

// 01 the six dimensions#dimensions

Each dimension gates a pricing mechanism the providers themselves made structural.

dimensionwhat it determinespricing mechanism it gates
D1 · Latency toleranceWhether work can move off the synchronous path.Batch eligibility; provisioned-throughput economics.
D2 · Context growthWhether input cost is constant, linear, or super-linear over a session or task.Context management as a lever; long-context pricing tiers.
D3 · Calls per taskThe multiplier between tasks (what you count) and requests (what the invoice counts).Estimate variance; retry cost; where spend guardrails belong.
D4 · Output : input ratioWhich price, input or the dearer output, dominates the bill.Whether max_tokens discipline or context trimming moves the total.
D5 · Prefix stabilityThe fraction of the prompt that is identical across requests.Prompt caching, which stacks with everything else.
D6 · Non-inference infra shareHow much of the total lives outside the token bill.Vector stores, sandboxes, seats, compliance: the costs domain-cut data never sees.
// 02 workload × dimension#crosswalk

No two presets share a profile. That is what makes them presets.

The last column is what the dimensions imply about delivery, described rather than recommended. Studio's picker still shows only the tiers your chosen provider actually sells.

workloadD1 · Latency toleranceD2 · Context growthD3 · Calls per taskD4 · Output : input ratioD5 · Prefix stabilityD6 · Non-inference infra sharedelivery modes that fall out
RAG · support copilotSecondsUser-facing, synchronousFlat, assembledContext rebuilt per request from retrieval; 8k to 32k in1 to 3Main call, embed, optional rerankabout 1 : 20200 to 800 out against 8k to 32k inMedium to highSystem prompt stable; retrieved chunks churnHighVector store, embeddings, corpus storageOn-demand; caching on the prefix; provisioned only at sustained high utilization
Chatbot/GeneralSecondsStreaming expectedAccumulates per turnSuper-linear; 15k to 40k by turn 20 unmanaged1 per turnSessions of 5 to 30 turnsabout 1 : 10, fallingRatio worsens as history growsHighPersona and guardrails identical across sessionsLow to mediumSession store, chat UIOn-demand; caching first; provisioned only at very high sustained concurrency
Agentic workflowMixedInteractive trigger, async executionAccumulates per stepScratchpad and tool outputs until the task completes5 to 30+The multiplicative dimension; high varianceVariableReasoning and structured output inflate both sidesHighTool schemas and instructions fixed across stepsMedium to highSandboxes, orchestration, external APIs, tracingOn-demand or async queue; caching on the schema prefix; spend guardrails
Batch summarizationHours (up to 24h)No user in the loopFlat, per documentIndependent jobs; 8k to 50k in each1 per documentVolume from document count, not call fan-outabout 1 : 40200 to 500 out; the most predictable of the sixVery highInstructions identical; only the content slot variesLowStorage and pipeline orchestrationBatch by default; caching stacks; provisioned is not warranted
Code assistantSub-second to asyncCompletions strict; review and test generation tolerantLarge, semi-stableRepository snapshot reused across many requests1 (completion) to 30+ (agent mode)Spans the widest range of any typeabout 1 : 1520 to 300 token completions; reviews longerHighest of any type90%+ cache hits on active codebasesDistinct axisSeats against raw API; the break-even is a seat questionOn-demand; repository-context caching is the lever; async for PR review
Classification / extractionHours (mostly)On-demand only for live routingFlat, per itemSchema prompt constant, content varies1 per itemPlus validation retries, tracked as a KPIabout 1 : 50+10 to 200 token structured outputnear 100%Cache hit rate approaches the ceilingLowPipeline only; the strongest fine-tune candidateBatch plus caching; a fine-tuned small model at volume
// hybrids the cases that walk in the door

Coding agenthighest

Presents as Code assistant. Behaves as Agentic on D3 (1 to 20+ calls per task) and D1 (sub-second to minutes).

A code-assistant preset under-estimates by the full call multiplier. Override D3 and keep the preset.

RAG-backed chatmost common

Presents as Chat assistant. Behaves as Chat on D2 (per-turn growth) plus RAG on D6 (vector store, embeddings) and per-request chunk injection.

Neither preset alone captures both. Chat misses the infrastructure lines; RAG misses the history growth.

Extraction inside an agentcontained

Presents as Agentic. Behaves as A classification step within D3.

Harmless on the orchestrator's context. If it fans out to a cheap fine-tuned model, the blended per-task cost has two model prices.

Long-form generationno bin exists

Presents as Batch summarization. Behaves as Output-dominated (D4 inverts to about 5 : 1), batch-tolerant, high prefix stability.

The nearest preset carries input-heavy assumptions and puts the optimization on the wrong side. Content generation exists as its own preset for this reason.

// 03 sources & grounding#grounding

The defaults, graded by what stands behind them.

The value in each cell is what Studio prefills today, read from the same module Studio reads. The grade and the note say where that kind of number comes from. Cache-eligible share is the best grounded: it follows from the caching pages. Token shapes are derivable from first principles. Usage volumes are the weakest, with no public benchmark, and your own telemetry replaces them first.

Strong primary published dataPartial first principles or a proxyEstimated no external sourceN/A does not apply
unit economicsRAGChatAgenticBatchCodeClassification
Monthly volume (users or items)
Estimated
500users· engineering default
No public source. Sized as a mid-market contact center (about 500 active agents). Validate against headcount or session analytics.
Estimated
2,000users· engineering default
No public source. Sized for a mid-size internal deployment (about 2,000 employees). Validate against auth or active-session logs.
Estimated
150users· engineering default
No public source. Sized for a power-user persona (about 150 users). Validate against orchestration or product analytics.
PartialE
100,000docs/mo· engineering default
Pipeline workload, no MAU. Monthly document volume from a first-principles cadence: 5,000 documents a day over 20 working days.
PartialC
120users· engineering default
Engineering-org AI adoption rates consistent with about 120 active developers at a mid-size software company. A rough proxy, not a direct figure.
Estimated
500,000records/mo· engineering default
Pipeline workload, no MAU. About 25,000 records a day from typical CRM or ticket ingestion. No published workload-level benchmark.
Requests per user per month (or calls per item)
Estimated
400req/user· engineering default
No public source. From a work pattern: 20 working days times 20 queries a day per agent. Validate with contact-center analytics.
Estimated
60req/user· engineering default
No public source. From a session model: 3 sessions a week, 5 turns each, 4 weeks. Validate against session logs.
Estimated
40req/user· engineering default
No public source. From a work pattern: 2 agentic tasks a day over 20 days. Validate with workflow run history.
PartialE
Monthly volume expressed directly. First principles from a typical ingestion cadence. Validate against job submission logs.
PartialD
800req/user· engineering default
Copilot's fill-in-the-middle studies imply about 30 completions a day for active developers. Review frequency estimated separately. The best public proxy available.
PartialE
1.03calls· engineering default
Monthly record volume from a first-principles pipeline rate. No workload-level public benchmark. Validate against upstream event counts.
Input tokens per request
PartialEG
4,200tokens· engineering default
Cookbook RAG examples show the prompt structure. System prompt 800, three chunks 2,700, query 150, history 550. First-principles word-count math.
PartialEG
3,500tokens· engineering default
Multi-turn docs show history accumulation. Averaged across session depth at word count times 1.33. Grows super-linearly without compression.
PartialFG
3,200tokens· engineering default
Benchmark traces publish step-level context. Agentic docs show scratchpad accumulation. Averaged across early and late steps.
PartialE
8,500tokens· engineering default
An average business document of about 6,000 words times 1.33, plus a system prompt of about 500 tokens. Highly variable by document type.
PartialDE
1,400tokens· engineering default
Copilot's completion context is about 6,000 characters. Blended with review and explanation requests that carry a full diff.
PartialE
750tokens· engineering default
Schema definition of about 350 tokens, fixed, plus about 400 tokens of variable content. Much shorter than summarization.
Output tokens per request
PartialE
380tokens· engineering default
A support answer of about 280 words times 1.33. Validate by sampling response lengths from your logs.
PartialE
420tokens· engineering default
An assistant response of about 315 words times 1.33. Varies widely by use case; set max_tokens from real sessions.
PartialF
480tokens· engineering default
Benchmark traces publish reasoning-trace lengths per step. Structured tool calls add format overhead. Averaged across orchestrator steps.
PartialE
350tokens· engineering default
A structured summary of about 260 words (three bullets and a paragraph) times 1.33. Constrain with an explicit format and max_tokens.
PartialDE
165tokens· engineering default
Completions run 20 to 100 tokens; reviews and explanations are much longer. Blended at a 75 : 25 completion-to-review ratio.
PartialE
45tokens· engineering default
Structured JSON with two to five fields, about 45 tokens. Highly predictable with JSON mode.
Cache-eligible share of inputStudio's input is the cache-eligible share; the research graded the resulting hit rate. Same page, different quantity.
StrongA
55% of input· engine default (`defaultPerformance`)
The caching page states cache read at 0.1 times and write at 1.25 times the base input price. The share is derived from the system prompt's part of total input with proper prefix placement.
StrongA
55% of input· engine default (`defaultPerformance`)
Automatic caching for multi-turn conversations is documented. The share is the static prefix's part of total input at average session depth.
StrongA
50% of input· engine default (`defaultPerformance`)
Tool definitions and system instructions cache as a fixed prefix; the scratchpad rotates. Varies by task length, so a share rather than a fixed number.
StrongA
45% of input· engine default (`defaultPerformance`)
The schema prompt is identical across all documents in a job, so most of the input is cacheable. The most favorable caching profile of the six.
StrongA
70% of input· engine default (`defaultPerformance`)
A fixed repository-context prefix enables high hit rates. The pre-warming pattern is described on the page.
StrongA
75% of input· engine default (`defaultPerformance`)
Schema and label definitions are identical across records in a batch. Placing them before the per-record content is what makes the share high.
LLM calls or actions per request
N/A
One LLM call per user request: embed, retrieve, generate.
N/A
One LLM call per user turn.
PartialFG
10calls· engineering default
Benchmark traces publish LLM call counts per task; multi-agent docs show orchestration patterns. The default of 10 LLM and 8 tool calls is a midpoint.
N/A
One LLM call per document, or two to four for very long documents that need chunking passes.
N/A
One LLM call per completion or review request.
N/A
Effectively 1.03 LLM calls per record: 97% first-pass success plus a 3% retry rate.
business-unit conversionRAGChatAgenticBatchCodeClassification
Primary business unit
Estimated
Support tickets a month. No industry benchmark for contact-center query volume; ranges span orders of magnitude. Validate against your ticketing system.
Estimated
Conversations a month. Derived from MAU times sessions per user, both themselves ungrounded. Validate against session logs.
Estimated
Tasks a month. No published benchmark, and a task is workload-specific: a lookup or a multi-hour research run. Validate against run history.
PartialE
Documents a month, first principles from a typical ingestion cadence. Validate against scheduler or platform logs.
PartialC
Developers. The CIO survey gives context on org sizes and adoption; a proxy, not a per-developer token figure.
Estimated
Records a month. No published benchmark. Validate against upstream event counts: CRM, support platform, transaction database.
Primary conversion ratio
Estimated
Queries per ticket, default 3.5. No published benchmark; simple tickets resolve in one or two queries, escalations need five to ten. A practitioner estimate.
Estimated
Turns per conversation, default 5. Varies widely: Q&A about 2, research about 15, support bots about 4. Validate against session analytics.
PartialF
LLM calls per task, default 10. Benchmark tasks run 8 to 15; 10 is a centre estimate. Validate against orchestration logs.
PartialG
Calls per document, default 1. One request per document is the documented batch pattern; multi-pass chunking for very long documents raises it. A design parameter, not an estimate.
PartialD
Completions per developer per day, default 30. Consistent with published activation and acceptance patterns. Validate with IDE telemetry.
PartialAG
Effective calls per record, default 1.03, from JSON-mode reliability: 97% first-pass success. Track your own retry rate; above 5% signals a prompt or schema problem.
Secondary conversion ratio
N/A
One LLM call per query; the query-to-ticket ratio is the whole conversion.
N/A
One LLM call per turn; turns per conversation is the whole conversion.
PartialF
Tool calls per task, default 8. About 80% of LLM steps invoke a tool. Each tool result re-enters the context, compounding input cost beyond the call count.
N/A
One call per document. No secondary factor.
Estimated
Review requests per PR, default 5. No published benchmark; estimated from typical PR interactions: diff review, tests, docs, refactor, explanation.
N/A
One call per record; the retry factor sits in the primary ratio.
// source keys
AAnthropic · Prompt cachingCache read and write multipliers against the base input price; per-model minimum token thresholds; the pre-warming pattern.page undated · read 2026-09-12
BAndreessen Horowitz · Welcome to LLMflation – LLM inference cost is going down fastThe cost trajectory for equivalent-performance inference. Informs how long a quoted rate is likely to hold.page dated 2024-11-12 · read 2026-09-12
CAndreessen Horowitz · How 100 Enterprise CIOs Are Building and Buying Gen AI in 2025Adoption rates by use case and engineering-org deployment patterns. A proxy for org sizes, not a token figure.page dated 2025-06-10 · read 2026-09-12
DGitHub · How GitHub Copilot is getting better at understanding your codeFill-in-the-middle, the character budget a completion model sees, and the acceptance-rate deltas from more context.page dated 2023-05-17 · read 2026-09-12
EFirst-principles derivationWord count times 1.33 tokens per word; session-depth models; typical document-length distributions by type. Any practitioner can reproduce it.a method, not a page
FAgentBench, SWE-bench, WebArena task tracesPublished academic benchmarks whose task traces include LLM call counts, reasoning-trace lengths and tool-invocation patterns. Named as methods; individual papers are not cited to a customer.a method, not a page
GAnthropic · Claude Developer Platform documentationRAG, multi-agent and agentic workflow examples with token counts and prompt structure.page undated · read 2026-09-12
// every page above is also in the references table on the overview, with both dates

// the one default the overview discloses in full, with all six fields: #assumptions · what is described but never printed: #never