OpenAI Agents API Pricing: What a Production AI Agent Really Costs in 2026
OpenAI opened the Agents API in public beta on September 10, 2026, giving every developer with an API key the same managed harness that runs Codex and ChatGPT for Work: session orchestration, automatic context compaction, tool and MCP server support, and subagent coordination, all behind one API call. The headline line from OpenAI's own announcement is unambiguous: "There are no additional fees for using the Agents API — you simply pay for the tokens and tools your agents use." That sentence is accurate, and it's also incomplete on its own, because "tools" turns out to include OpenAI's own hosted sandbox infrastructure, which bills separately at container rates that don't appear anywhere in the headline pricing conversation.
This article covers what the Agents API actually costs to run, and — because the real decision most teams face isn't "which price tier" but "managed or build it ourselves" — models that comparison directly, without pretending an engineering-hour estimate is a vendor price.
This is a companion to, not a replacement for, any broader "real cost of running an AI agent" content — this piece is specifically about OpenAI's Agents API architecture and the managed-vs-DIY decision it forces, not a general survey of agent economics across providers.
What's verified, directly from OpenAI's launch materials
- The Agents API entered public beta on September 10, 2026, available to all API developers.
- It exposes the Codex harness — the orchestration layer OpenAI built for its own coding agent and ChatGPT for Work — as a general-purpose API.
- No additional platform fee. Usage is billed through the models and tools a session consumes, confirmed in OpenAI's own announcement post and same-day API changelog entry.
- The API organizes around four primitives: an agent (model, instructions, tools, MCP servers), an environment (an optional sandbox for file access and command execution), a session (a durable instance persisting state across turns), and events/items (inputs sent and outputs returned).
- Developers choose where agent compute runs: an OpenAI-hosted sandbox, their own infrastructure (including inside a VPC), or one of nine named partner sandbox providers (Blaxel, Cloudflare, Daytona, DigitalOcean, E2B, Modal, Oracle, Runloop, and Vercel).
- OpenAI-hosted sandboxes bill at standard container rates, separately from model API charges: $0.03, $0.12, $0.48, and $1.92 per 20-minute session depending on the memory tier selected, with eligible container sessions billed by the minute subject to a 5-minute minimum per session. (Both statements are printed on OpenAI's own pricing page; which sessions the per-minute treatment actually applies to is not defined there — see the dedicated section below.)
- Zero Data Retention is explicitly not supported on the Agents API as of the public beta; data stays US-only.
- OpenAI's own launch example used GPT-6 Astra, priced at $10 per million input tokens and $50 per million output tokens — a genuinely expensive model, chosen presumably to showcase the harness on OpenAI's most capable reasoning model, not as a suggestion that every Agents API workload needs it.
- Early customers cited in the launch (Ciridae, Long Lake, WithCoverage, SafetyKit, Dwelly, Hypha, deepsense.ai, Nash.ai) reported specific results: Ciridae claimed a 4x latency reduction from built-in subagent support, and Hypha reported an 86% drop in failed responses after migrating to the managed harness. These are vendor-cited customer claims, not independently audited figures — treat them as directional evidence the harness solves real reliability problems, not as a guaranteed outcome for any given workload.
The real cost stack
| Component | Who bills it | What drives the cost |
|---|---|---|
| Model tokens | OpenAI, per the chosen model's standard rate | Model choice, prompt length, output length, number of turns/tool-call round-trips |
| Tools (function calls, MCP servers, web search, etc.) | Whatever the tool costs to call — some free, some metered by a third party | Which tools the agent invokes and how often |
| Hosted sandbox (if used) | OpenAI, at container rates | Memory tier selected; whether a session is billed as a flat 20-minute block or prorated per minute is genuinely ambiguous in OpenAI's own documentation — see below |
| Long-running context/compaction | Folded into token cost — compaction reduces the tokens re-sent on long sessions, so it's a cost reducer, not a separate line item | Session length; compaction matters more the longer a session runs |
| Subagents | Each subagent's own model/tool/token usage, billed the same way as the primary agent | How many subagents are spawned and how much work each does |
| Retries and failed tasks | Full token cost of every attempt, whether it succeeds or not | Task complexity, model reliability, how aggressively your harness retries |
| Engineering/maintenance | Not billed by OpenAI at all — this is your own team's time | Whether you're building on the managed harness (less) or DIY orchestration (more) |
Sandbox pricing in practice — and a genuine ambiguity in OpenAI's own documentation
OpenAI's official API pricing page states the container rates as "$0.03 / $0.12 / $0.48 / $1.92 per 20-minute session per container" (1GB/4GB/16GB/64GB memory tiers) — a flat rate per session, tabulated directly. The same page also carries a separate footnote: "Eligible container sessions will be billed by the minute, with a 5-minute minimum per session." These are two different billing mechanics on the same page, and OpenAI's documentation does not define which sessions qualify as "eligible" for the per-minute treatment versus which are billed as a flat 20-minute block regardless of actual duration. Independent technical analysis comparing dated snapshots of OpenAI's own pricing page (via the Wayback Machine, captures from September 9, 10, and 12, 2026) found this same ambiguity present since launch — it is not a temporary documentation error that has since been fixed, and OpenAI has not publicly clarified it as of this writing.
Because the official source is genuinely ambiguous, we are not going to invent a resolution. For the calculations in this article, we model sandbox cost using the flat, unambiguously-tabulated per-20-minute-session rate — i.e., any session up to 20 minutes bills at the full listed rate for its memory tier, and a session running longer bills for additional 20-minute blocks. This is the more clearly documented of the two mechanics and, in most realistic short-session scenarios, the higher-cost (more conservative) interpretation for budgeting purposes. If the per-minute-with-5-minute-minimum footnote turns out to apply to your specific sessions, your real cost for short tasks could be lower than what's modeled here — verify directly against your own account's billing before finalizing a large-scale budget.
| Memory tier | Rate per 20-minute session (flat-block basis used throughout this article) | Cost per hour (continuous use, i.e., 3 full 20-min blocks) | Cost per 100 sessions, each under 20 minutes (each bills as one full block under this interpretation) |
|---|---|---|---|
| Smallest | $0.03 | $0.09 | $3.00 (100 × $0.03) |
| Small-medium | $0.12 | $0.36 | $12.00 (100 × $0.12) |
| Medium-large | $0.48 | $1.44 | $48.00 (100 × $0.48) |
| Largest | $1.92 | $5.76 | $192.00 (100 × $1.92) |
Under this flat-block interpretation, a coding agent that finishes a task in 90 seconds is billed the same as one that runs the full 20 minutes, as long as both stay under that threshold — which is what makes the memory-tier choice, not session length, the main lever for controlling sandbox cost on short-running tasks. Setting environment.type to none for any task that doesn't actually need file access or shell commands avoids sandbox billing entirely regardless of which interpretation turns out to be correct, and remains the single cheapest optimization available on this API.
Managed harness vs. DIY orchestration
This is the decision the brief specifically asked us to model honestly — with engineering cost treated as a labeled assumption, never as a vendor price.
What the managed harness (Agents API) removes from your build: session persistence and state management across turns; context compaction logic (deciding what to summarize or drop as a conversation grows past the model's effective context); tool-call scheduling and retry logic; subagent lifecycle management (spawning, monitoring, terminating); and sandboxed execution environments, if you use OpenAI's hosted option or a partner sandbox rather than standing up your own.
What DIY orchestration requires you to build and maintain instead: all of the above, from scratch, using open-source or custom frameworks; your own compaction strategy (get this wrong and you either blow your context window or silently drop information the agent needed); your own retry and error-handling logic; your own observability/tracing layer (a real cost the managed harness also affects — Agents API concentrates execution inside OpenAI's infrastructure, so teams with existing internal tracing and replay tooling will need to reconstruct equivalent visibility using the API's streaming events surface, since that internal layer doesn't carry over automatically).
The honest cost comparison: the Agents API's direct dollar costs (tokens, tools, sandbox) are close to what a well-built DIY system would spend on the same underlying model calls and compute — OpenAI isn't charging a premium for the harness itself, per its own "no additional fee" claim. The real cost difference is engineering time: building and maintaining orchestration, compaction, and retry logic is a genuine, ongoing engineering cost that doesn't show up on any OpenAI invoice. We are not going to assign that a specific dollar figure here, because fully-loaded engineering cost varies enormously by team, market, and seniority, and any single number we picked would be an assumption dressed as a fact. What we can say directionally: the managed harness trades a recurring infrastructure/token cost (visible, metered, scales with usage) for what would otherwise be a recurring engineering-time cost (invisible on any vendor invoice, largely fixed regardless of usage volume, and front-loaded before the system ever runs its first production task).
Worked scenarios
(Labeled assumptions throughout — verify against your own workload before budgeting.)
Scenario 1: 1,000 jobs/month, small SaaS feature
Assumptions: a lightweight agent feature inside an existing SaaS product (e.g., auto-drafting a document from user input), GPT-5.6 Terra as the model ($2/$12 per million tokens), no sandbox needed (no file/shell access required), average job uses ~3,000 input tokens and ~800 output tokens across 2–3 tool calls to internal, free APIs.
| Component | Monthly cost |
|---|---|
| Model tokens (1,000 × (3,000 × $0.000002 + 800 × $0.000012)) | 1,000 × ($0.006 + $0.0096) = $15.60 |
| Sandbox | $0 (not used) |
| Tools | $0 (internal APIs) |
| Total | ~$16/month |
Cost per attempted job: ~$0.016. If 90% succeed on the first attempt and the other 10% require one retry, cost per successful job rises to roughly $0.0176 — a small effect at this failure rate, but one that compounds fast at higher failure rates (see the retry-sensitivity table below).
Scenario 2: Coding/research agent
Assumptions: a coding-assistant-style agent handling ~300 tasks/month, each requiring a hosted sandbox (medium tier, $0.48/20 min) for running code and reading/writing files, GPT-6 Astra as the model given the complexity ($10/$50 per million tokens), average task runs 15 minutes of sandbox time and consumes ~15,000 input / 4,000 output tokens across multiple tool-call round-trips.
| Component | Monthly cost |
|---|---|
| Model tokens (300 × (15,000 × $0.00001 + 4,000 × $0.00005)) | 300 × ($0.15 + $0.20) = $105.00 |
| Sandbox (300 tasks, 15 min each — under the flat-block convention used throughout this article, a 15-minute session bills as one full 20-minute block at the medium tier) | 300 × $0.48 = $144.00 |
| Total | ~$249/month |
Cost per attempted task: ~$0.83. This scenario is dominated almost equally by model tokens and sandbox time — a useful reminder that for compute-heavy agent tasks, the sandbox line is not a rounding error.
Scenario 3: High-volume workflow
Assumptions: 50,000 jobs/month, routed model selection (80% on GPT-5.6 Luna at $0.20/$1.20 for routine jobs, 20% escalated to GPT-5.6 Terra at $2/$12 for complex ones), no sandbox (pure API/tool-calling workload, e.g., data enrichment), average job ~2,000 input / 500 output tokens.
| Component | Monthly cost |
|---|---|
| Luna jobs (40,000 × (2,000 × $0.0000002 + 500 × $0.0000012)) | 40,000 × ($0.0004 + $0.0006) = $40.00 |
| Terra jobs (10,000 × (2,000 × $0.000002 + 500 × $0.000012)) | 10,000 × ($0.004 + $0.006) = $100.00 |
| Total | ~$140/month |
Cost per attempted job: ~$0.0028. This is the scenario where volume and model routing matter more than any other lever — at 50,000 jobs/month, a naive all-Terra approach would cost roughly $500/month instead of $140, a 3.6x difference purely from routing 80% of routine work to a cheaper model.
Retry and failure sensitivity
(Methodology note: the table below is a NotCheapAI sensitivity calculation, not a claim about any specific agent's real failure rate. It shows how cost-per-successful-job scales under different labeled failure assumptions, using Scenario 1's $0.0156/attempt baseline.)
| First-attempt success rate | Effective attempts per success | Cost per successful job |
|---|---|---|
| 95% | ~1.05 | ~$0.0164 |
| 85% | ~1.18 | ~$0.0184 |
| 70% | ~1.43 | ~$0.0223 |
| 50% | ~2.00 | ~$0.0312 |
The takeaway that generalizes beyond this specific scenario: failure rate matters more to your real unit economics than most model-selection decisions do. Moving from a 95% to a 70% success rate raises your cost-per-successful-outcome by roughly 36% (from ~$0.0164 to ~$0.0223 per success in this baseline); moving all the way down to a 50% success rate very nearly doubles it (to ~$0.0312, roughly a 90% increase). Even the smaller of those two swings is a bigger effect than switching from Luna to Terra pricing tiers in Scenario 3. Reliability engineering (better prompts, better tool design, the managed harness's own compaction/retry handling) is very often the highest-leverage cost lever available, ahead of raw token-price shopping.
Engineering-cost break-even: when does DIY start making sense?
We won't manufacture a specific dollar crossover point, since it depends entirely on your team's cost structure — but the shape of the decision is worth stating plainly:
- At low-to-moderate volume and complexity (roughly Scenarios 1 and 3 above), the managed harness's direct costs are small enough that the engineering time saved by not building compaction/retry/subagent logic yourself is very likely to dominate the decision. Building that infrastructure yourself for a $16–$140/month workload is close to never worth it.
- At high volume with a stable, well-understood workflow (think: millions of jobs/month, a single well-characterized task type), the calculus shifts, because at that scale the fixed engineering cost of a custom, highly-optimized orchestration layer gets amortized across enough volume that shaving even a small percentage off per-job token or sandbox cost can outweigh the ongoing convenience of the managed harness. This is the same logic that governs "build vs. buy" decisions for any infrastructure category, not something unique to OpenAI's Agents API.
- The volume threshold where this becomes "meaningful" is a function of your engineering cost, not a fixed number of jobs — a team with abundant senior infrastructure engineering capacity reaches that crossover at lower volume than a team without it. Anyone quoting you a specific universal job-count threshold is guessing.
BUILD & EARN: contribution margin for a founder building a paid agent product
If you're building a product that charges customers to use an agent built on this API, your contribution margin per customer job is:
Price charged per job/task − (model tokens + tools + sandbox + a reasonable allocation for retries/failures) = contribution margin per job
Using Scenario 2's coding-agent economics (~$0.83/attempted task, higher once you build in a realistic retry allowance — say $0.95–$1.10 per task including some failure buffer) as an example: a product charging $2.00 per completed task would run a contribution margin of roughly $0.90–$1.05 per task before any allocation of fixed costs (your own engineering time, support, marketing, the SaaS wrapper around the agent). That margin needs to cover all of those fixed costs before the business is actually profitable — a $1,000/month engineering or support cost requires roughly 950–1,100 completed tasks just to break even on that one line item, separate from customer acquisition cost. The core lesson, consistent with how NotCheapAI treats every unit-economics question: a positive per-job margin is necessary but nowhere near sufficient — it has to be large enough, at your realistic volume, to cover everything that isn't per-job variable cost.
Frequently asked questions
Does OpenAI charge extra just for using the Agents API? No — per OpenAI's own launch announcement, there is no additional fee for the API itself. You pay for the tokens and tools your agents consume, and separately for OpenAI-hosted sandbox time if you use one.
Is the hosted sandbox required? No. You can run agent compute in your own infrastructure (including inside a VPC) or on one of nine named partner sandbox providers instead of OpenAI's hosted option.
What's the cheapest way to reduce Agents API costs? Two levers stand out from the verified pricing: set environment.type to none for any task that doesn't need file/shell access (avoiding sandbox billing entirely), and route routine, low-complexity jobs to a cheaper model rather than defaulting every task to your most capable (and most expensive) model.
Is this the same thing as "the real cost of running an AI agent" that NotCheapAI has already covered? No — that broader piece covers agent economics across providers and architectures generally. This article is specifically about OpenAI's Agents API launch, its particular cost stack (sandbox billing, the no-platform-fee structure), and the managed-vs-DIY decision that stack forces.
Hidden costs the headline pricing doesn't surface
Data residency and Zero Data Retention gaps. The Agents API public beta explicitly does not support Zero Data Retention, and data stays US-only. For teams in regulated industries or non-US jurisdictions with data-residency requirements, this isn't a line item OpenAI bills for — it's a compliance constraint that may force either a delay in adopting the managed harness, additional legal/compliance review cost, or a DIY approach specifically to satisfy a requirement the managed product doesn't yet meet. This is a real cost of the "managed" choice that has nothing to do with tokens or sandbox minutes.
MCP server and custom tool maintenance. The Agents API supports Model Context Protocol servers and custom tools, which is a genuine capability advantage — but every tool or MCP server you connect is something your team built, is billed by whatever it costs to call (a third-party API, your own infrastructure), and needs to keep working as that underlying service changes. The managed harness removes orchestration overhead; it does not remove the ongoing maintenance burden of the tools themselves.
Subagent cost multiplication. A primary agent that spawns subagents multiplies your token exposure in a way that's easy to underestimate when scoping a workload: if a task spawns three subagents, each running its own multi-turn conversation with its own tool calls, your effective per-task cost isn't the primary agent's token count — it's the sum across every subagent it spawned, plus the primary agent's own coordination overhead. Ciridae's cited 4x latency reduction from subagent support is a genuine benefit, but latency improvement and cost increase can arrive together on the same feature; model subagent-heavy workflows with this multiplication explicitly rather than assuming subagents are "free" because the harness manages their lifecycle for you.
Observability reconstruction cost. As noted earlier, teams with existing internal tracing and replay tooling lose that layer when execution moves into OpenAI's managed infrastructure, unless they rebuild equivalent visibility using the streaming events surface. This is a one-time-ish engineering cost (rebuilding the tooling) plus a smaller ongoing cost (maintaining it against an API surface OpenAI controls and can change) — worth budgeting explicitly for any team currently relying on custom observability around a DIY orchestration layer.
What changes at true scale
The scenarios modeled above (1,000 to 50,000 jobs/month) sit in a range where the Agents API's direct costs are modest in absolute terms and the managed-vs-DIY decision is dominated by engineering time, as argued earlier. At meaningfully larger scale — millions of jobs per month, sustained over years — a few dynamics shift the calculus that are worth naming even without assigning them specific numbers: token costs at that volume become large enough in absolute dollar terms that even small percentage optimizations (a cheaper model for a genuinely routable subset of tasks, more aggressive caching, tighter prompts) start to represent real money regardless of engineering cost; sandbox costs, if used heavily, compound the same way, and — under whichever of the two billing mechanics discussed above actually applies to a given session — short-lived, high-frequency tasks are where that ambiguity has the largest cumulative dollar effect; and the "no additional fee" framing, while still literally true, matters less as a competitive advantage once a team's engineering capacity is large enough to build and maintain sophisticated DIY orchestration anyway — the managed harness's main value proposition (removing that engineering burden) is worth the most to teams that don't already have deep infrastructure engineering capacity, and worth progressively less to teams that do.