The Real Cost of Running an AI Agent in 2026
There is no universal "cost per agent" figure, and any article that gives you one is guessing. An agent's cost is the sum of several genuinely different meters — model tokens (which themselves grow as the agent's context accumulates across a session), tool and API calls it makes along the way, retries when a step fails, and increasingly, separate charges for agent-orchestration infrastructure some providers now bill on top of raw token usage. This guide doesn't model a universal number; it models the stack, with real verified rates for each layer, so you can build your own estimate for your own agent.
The five cost layers of a real agent
- Base model tokens — the underlying LLM cost, priced per input/output token like any API call.
- Context growth within a session — an agent's context typically grows turn over turn as it accumulates conversation history, tool results, and intermediate reasoning, meaning later turns in a long session cost more than early ones even at a flat per-token rate.
- Tool and function calls — web search, code execution, file search, and similar capabilities are frequently billed separately from token usage, per call.
- Retries and failed steps — a step that errors, times out, or produces an unusable result still consumed tokens (and possibly a billed tool call) before being discarded.
- Orchestration/session infrastructure — some providers now bill separately for the runtime that manages a multi-step agent session, on top of the tokens and tool calls within it.
For a provider-specific breakdown of the managed harness, sandbox charges and build-versus-buy trade-off, see the OpenAI Agents API cost guide.
Layer 1 and 2: base tokens and context growth, worked through
Take a single agent task: research a topic, then write a summary, across 10 conversational turns. Assume each turn's input grows by roughly 500 tokens (accumulated history and tool outputs) and each turn generates roughly 800 output tokens.
| Turn | Cumulative input tokens | Output tokens (this turn) |
|---|---|---|
| 1 | 500 | 800 |
| 5 | 2,500 | 800 |
| 10 | 5,000 | 800 |
CALCULATION — total input tokens across all 10 turns (a running sum, since each turn resends the accumulated context): 500 + 1,000 + 1,500 +... + 5,000 = 500 × (1+2+...+10) = 500 × 55 = 27,500 input tokens. Total output: 10 × 800 = 8,000 output tokens.
At GPT-5.6 Luna rates ($0.20/M input, $1.20/M output): (0.0275 × $0.20) + (0.008 × $1.20) = $0.0055 + $0.0096 = $0.0151 per session. At Claude Sonnet 5 rates ($3/M input, $15/M output): (0.0275 × $3.00) + (0.008 × $15.00) = $0.0825 + $0.12 = $0.2025 per session — roughly 13x Luna's cost for the identical turn structure.
The mechanism worth internalizing: because input resends accumulated history every turn, a 10-turn agent session bills roughly 5.5x more input tokens than a single 10-turn's-worth of new content would suggest (the 55 multiplier above vs. a flat 10x for non-accumulating turns) — context growth compounds cost faster than turn count alone implies, and this effect gets worse, not better, on longer sessions or agents that maintain very large accumulated context.
Layer 3: tool and function call costs
Tool calls are billed per invocation, separately from the tokens spent processing whatever the tool returns. Verified current rates from two major providers:
| Provider | Tool | Rate |
|---|---|---|
| Anthropic | Web search | $10 per 1,000 searches |
| Anthropic | Code execution | 50 free container-hours/day per org, then $0.05/container-hour |
| Anthropic | Managed Agents (session runtime) | $0.08/session-hour, plus standard token rates |
| xAI | Web search, X search, code execution | $5 per 1,000 calls |
| xAI | File search | $10 per 1,000 calls |
| xAI | Document collection search | $2.50 per 1,000 calls |
CALCULATION — an agent that makes 3 web searches and 1 code-execution call per session, run 500 times/month, on Anthropic's rates:
- Web search: 500 × 3 = 1,500 searches × ($10/1,000) = $15.00
- Code execution: assuming each call runs under a minute and stays within the free container-hour allowance for a low-volume deployment, this can be $0 — but at higher volume, once the 1,550 free monthly container-hours per organization are exhausted, additional usage bills at $0.05/hour, which adds up quickly for compute-heavy agent workloads run at scale.
Tool costs are easy to underestimate specifically because they scale with agent behavior (how many searches or tool invocations the agent decides to make per task) rather than with a fixed per-session rate — an agent that's been prompted or designed to "search thoroughly" can multiply this cost several times over relative to one that searches sparingly, for the same underlying task.
Layer 4: retries and failed steps
An agent step that errors out, hits a rate limit, produces output that fails a validation check, or simply needs to be redone doesn't refund the tokens or tool calls it already consumed. The cost impact compounds with session length, because a failure late in a long session means redoing (and re-billing) a step built on top of substantial accumulated context.
CALCULATION — using the 10-turn session from Layer 1/2 above (Claude Sonnet 5 rates, $0.2025/session at 0% failure), add a 15% step-failure rate where a failed step requires redoing that turn with its full accumulated context:
- Expected extra turns: 10 × 0.15 = 1.5 additional turns, at an average context depth (since failures can happen at any point in the session) — a reasonable approximation is the mid-session average input size, roughly 2,750 tokens, plus 800 output tokens per redone turn.
- Extra cost: 1.5 × [(0.00275 × $3.00) + (0.0008 × $15.00)] = 1.5 × [$0.00825 + $0.012] = 1.5 × $0.02025 ≈ $0.030
- New session cost: $0.2025 + $0.030 ≈ $0.233, roughly a 15% increase — tracking the failure rate closely, as expected, but the absolute dollar impact of that same 15% failure rate grows with session length and context depth, not just with the failure rate itself.
The longer and more context-heavy an agent's sessions get, the more expensive its failure rate becomes in absolute terms, even if the failure rate itself doesn't change — a dynamic that doesn't show up if you only track failure rate as a percentage without connecting it to session cost structure.
Layer 5: orchestration and session infrastructure
Beyond tokens and tool calls, at least one major provider now bills separately for the runtime that manages a multi-step agent session: Anthropic's Managed Agents feature, introduced in 2026, costs $0.08 per session-hour of active runtime, on top of standard token rates for the model calls within that session. This is a genuinely distinct cost layer from the four above — it's not a token cost or a tool-call cost, it's a charge for the orchestration infrastructure itself, and it scales with session duration (wall-clock time) rather than token volume or step count.
CALCULATION — an agent session running 20 minutes of active runtime, on top of the $0.233 base-token-and-retry cost calculated above (the separate $0.03 tool-call cost from Layer 3 is added alongside this in the full stack below, not included in the $0.233 figure itself): 20 minutes = 0.333 hours × $0.08 = $0.027 additional. On this specific example, orchestration is a relatively small add-on — but for agents that run long, mostly-idle sessions (waiting on external processes, polling, long-running background tasks), session-hour billing can become the dominant cost layer rather than a minor one, since it accrues on wall-clock time regardless of how much actual model or tool activity occurs within it.
Products differ meaningfully here, and this guide won't collapse that difference into a single figure. Not every agent framework or provider bills a separate orchestration fee — some fold all costs into token and tool-call billing alone. Check the specific documentation for whichever agent framework or provider you're using before assuming either model applies to your setup.
The layer this guide hasn't priced yet: memory and storage
Agents that persist state across sessions — remembering prior conversations, maintaining a knowledge base, or building up a working memory of a task over time — add an optional additional cost dimension, outside the five-layer session-level model used in the worked calculation above: embedding generation and vector storage.
Embedding costs are typically priced per token, separately from chat/completion tokens, and are usually far cheaper per token than generation — commonly a fraction of a cent per million tokens on budget-tier embedding models, though exact current rates vary by provider and should be checked directly rather than assumed, because embedding prices vary by provider; check the provider's current rate card before estimating.
Vector storage costs scale with the size of what's stored (number of vectors and their dimensionality) rather than with query volume directly, and are billed either as a managed vector-database subscription (typically a flat monthly tier based on storage volume) or as self-hosted infrastructure cost. An agent with a small, bounded memory (a single user's recent conversation history) adds negligible storage cost; an agent building a large shared knowledge base across many users or a long time horizon can accumulate a meaningful, ongoing storage bill independent of how often that memory is actually queried.
The honest scope limit here: unlike the token, tool-call, and orchestration rates verified above, memory and storage costs vary enormously by architecture choice (managed vector DB vs. self-hosted, embedding model choice, retention policy) in a way that resists a single representative worked example. The mechanism to budget for is directional: if your agent design includes persistent memory beyond a single session, add a storage-cost line that scales with how much you retain and for how long — not with how often the agent runs — as a genuinely separate budget item from the five session-level layers above.
Putting the full stack together
CALCULATION — the running example, all five layers combined, for one agent session:
| Layer | Cost |
|---|---|
| Base tokens (10-turn session, Sonnet 5 rates) | $0.2025 |
| Tool calls (3 web searches) | $0.03 (3 × $10/1,000) |
| Retries (15% step-failure rate) | $0.030 |
| Orchestration (20 min active runtime) | $0.027 |
| Total per session | ≈ $0.29 |
CALCULATION — scaled to 1,000 sessions/month: 1,000 × $0.29 = $290/month. Scaled to 10,000 sessions/month (a genuinely high-volume deployment): $2,900/month. Notice that in this worked example, base tokens are roughly 70% of total session cost, tool calls roughly 10%, retries roughly 10%, and orchestration roughly 9% — but that split is specific to this session's shape (turn count, tool-call frequency, failure rate, session duration) and will shift substantially for an agent with a different profile: a research-heavy agent making many more tool calls, or a long-running background agent where orchestration dominates, will have a very different layer breakdown from the one modeled here.
Why "cost per agent" isn't a meaningful universal number
The worked example above used one specific set of assumptions — turn count, context growth rate, tool-call frequency, failure rate, session duration, and model choice. Change any one of those and the total shifts substantially, often by an order of magnitude:
- Swapping Claude Sonnet 5 for GPT-5.6 Luna in the base-token layer alone drops that layer's cost by roughly 13x.
- Doubling the tool-call frequency (6 searches instead of 3) doubles that layer's cost directly.
- An agent with a 40% step-failure rate instead of 15% sees a proportionally much larger retry-cost layer, and that layer's absolute size depends on how deep into the session those failures happen.
- A long-running agent (hours of wall-clock time, sparse activity) shifts the balance heavily toward the orchestration layer, which a short, dense agent session barely touches.
Any single published "AI agents cost $X" figure necessarily picked one point in this multidimensional space and presented it as representative — which it usually isn't, for a different agent doing a different task on a different provider.
For a product-level comparison of managed agent experiences, see Manus vs ChatGPT Work vs Claude Agent.
For customer-support agents, compare Intercom Fin with Zendesk AI and Gorgias AI Agent billing. For voice agents, see Vapi vs Retell vs Bland.
How to model your own agent's real cost
- Estimate turn count and context growth per session — how many back-and-forth steps does a typical task take, and how much does context grow per turn (system prompt, accumulated history, tool outputs)?
- Count expected tool calls per session — web searches, code execution, file lookups — and their published per-call rate.
- Measure your actual failure/retry rate on a representative test batch, not an assumed number — this is the variable no vendor publishes and every deployment differs on.
- Check whether your provider bills separate orchestration/session time, and if so, estimate typical session wall-clock duration, including idle/waiting time.
- Multiply the per-session total by expected monthly volume, and re-run the estimate whenever you change model, add a tool, or materially change the agent's task complexity — any of those shifts the whole stack, not just one line item.
FAQ
What does it cost to run an AI agent per month? There's no single answer — see the worked example above, which lands at roughly $290/month for 1,000 sessions of one specific agent shape. A different agent, model, or usage volume will land at a meaningfully different number. Model your own stack using the five-layer framework above rather than adopting a published average.
Which cost layer is usually the biggest? Base model tokens are the largest layer in most agent designs, but this isn't universal — tool-call-heavy agents (frequent web search or file lookup) and long-running, mostly-idle agents (where session-hour orchestration billing accrues regardless of activity) can shift the balance substantially.
Does context growth really matter that much? Yes — in the worked example, a 10-turn session with accumulating context bills roughly 5.5x the input tokens that 10 turns of flat, non-accumulating content would. Long agent sessions with large accumulated context are a genuine and often underestimated cost driver.
How do I reduce agent costs without changing the model? The layers most within a builder's direct control are context growth (trimming or summarizing accumulated history rather than letting it grow unbounded), tool-call frequency (prompting or designing the agent to be more selective about when it invokes a tool), and failure rate (better step validation and error handling reduce the retry layer directly).
Is orchestration billing common across agent platforms? Not universally — it's a real, verified cost on at least one major platform (Anthropic's Managed Agents, at $0.08/session-hour), but not every agent framework or provider bills this layer separately. Check your specific provider's documentation rather than assuming either model applies.
The NotCheapAI verdict
An AI agent's real cost is a five-layer stack — base tokens, context growth, tool calls, retries, and sometimes orchestration billing — and collapsing that into one number always means picking assumptions that may not match your actual agent. The layers that most commonly get underestimated are context growth (which compounds faster than turn count suggests) and retry rate (whose absolute cost impact grows with session length, not just with the failure percentage itself). Model your own agent against this framework, using your own measured failure rate rather than a guess, before trusting any single published figure for what agents "cost."