NotCheapMAKE EVERY CREDIT COUNT

GPT-6 Luna vs Claude Sonnet 5 vs Gemini 3.8 Flash: Real API Cost in 2026

At the rates published on September 26, 2026, GPT-6 Luna is the lowest-cost model in this comparison by raw token price: $0.10 per million input tokens and $0.50 per million output tokens. Claude Sonnet 5 costs $2/$10 per million; Anthropic made its introductory rate permanent on August 10, so the previously announced $3/$15 price does not take effect. Gemini 3.8 Flash is temporarily priced at $0.75/$3.75 through December 31, 2026, with standard rates of $1.50/$7.50 scheduled from January 1, 2027.

These figures are list token rates, not a quality or total-cost ranking. Long prompts, cached input, batch processing, tool calls, retries and the number of attempts required to complete a task can change a production bill. (OpenAI GPT-6 release notes; Anthropic Claude Sonnet 5; Gemini 3.8 Flash pricing)

Current price table

API modelInput / 1M tokensOutput / 1M tokensCurrent pricing note
GPT-6 Luna$0.10$0.50Standard rate published September 22, 2026
Claude Sonnet 5$2.00$10.00Introductory price made permanent August 10, 2026
Gemini 3.8 Flash$0.75$3.75Introductory price through December 31, 2026; then $1.50/$7.50

OpenAI's September 22 changelog lists GPT-6 Luna's cached-input rate at $0.01 per million tokens. The model price page also documents higher pricing for prompts above 272K input tokens. Anthropic and Google offer their own caching, batch and context rules; do not apply one provider's discount to another provider's bill.

Workload examples

The formula is input tokens ÷ 1,000,000 × input rate + output tokens ÷ 1,000,000 × output rate. These examples use uncached standard token rates and exclude tools and taxes.

Long-context assistant

Assume 1,000 sessions per month, each with 100,000 input tokens and 5,000 output tokens. Total volume is 100M input and 5M output tokens.

ModelInput chargeOutput chargeMonthly total
GPT-6 Luna$10.00$2.50$12.50
Claude Sonnet 5$200.00$50.00$250.00
Gemini 3.8 Flash$75.00$18.75$93.75

For this input-heavy example, GPT-6 Luna's raw token bill is 20 times lower than Claude Sonnet 5's and 7.5 times lower than Gemini 3.8 Flash's. That ratio says nothing about whether the models complete the same tasks with the same number of tokens or retries.

Document summarization

For 500 documents at 20,000 input and 2,000 output tokens each, volume is 10M input and 1M output tokens.

ModelInput chargeOutput chargeMonthly total
GPT-6 Luna$1.00$0.50$1.50
Claude Sonnet 5$20.00$10.00$30.00
Gemini 3.8 Flash$7.50$3.75$11.25

Output-heavy generation

For 200 jobs at 2,000 input and 6,000 output tokens each, volume is 0.4M input and 1.2M output tokens.

ModelInput chargeOutput chargeMonthly total
GPT-6 Luna$0.04$0.60$0.64
Claude Sonnet 5$0.80$12.00$12.80
Gemini 3.8 Flash$0.30$4.50$4.80

Gemini's advertised introductory rate is less than half of its scheduled standard price in this example: at $1.50/$7.50, the same 0.4M/1.2M workload would cost $9.60 rather than $4.80. Budget the post-promotion amount for a service expected to run beyond 2026.

The cost of long prompts

Context length can matter as much as the posted rate. OpenAI's GPT-6 Luna documentation applies a higher rate to prompts exceeding 272K input tokens. Gemini 3.8 Flash supports long-running and complex tasks, but reasoning and tool iterations can consume additional tokens. For Claude, apply Anthropic's current context and cache rules for the selected API request. The safe estimate is based on observed production token counts rather than maximum context advertised on a product page.

Caching changes the effective rate only for eligible repeated input. For example, an application that resends a large system prompt may see a materially different bill if cache hits are recognized, while unique user content remains at the regular input rate. Measure cached and uncached tokens separately in provider usage reports.

When the cheapest model costs more

Raw tokens are only one component of cost. A model with a lower list price can become more expensive per successful task if it requires more retries, longer reasoning traces, extra tool calls, or human correction. Conversely, a higher-priced model may save money if it completes the workflow in one pass. Track cost per accepted task, not just cost per million tokens:

effective cost per accepted task = total API spend ÷ number of accepted results

Include search, code-execution or other tool charges as separate line items. If comparing providers for production, run the same representative tasks, capture input/output and cached tokens, record retries, and use the acceptance criteria your application actually needs.

For OpenAI's managed agent orchestration and its associated costs, see our OpenAI Agents API cost guide.

Which one should you choose?

  • GPT-6 Luna: strongest raw-price choice for high-volume, cost-sensitive workloads when its task quality and tool behavior fit.
  • Claude Sonnet 5: price it at $2/$10. Do not budget at the abandoned $3/$15 post-promotion rate.
  • Gemini 3.8 Flash: account for the temporary $0.75/$3.75 offer and the scheduled $1.50/$7.50 rate from January 2027.

For any model, estimate using your actual input/output mix and expected completion rate. The cheapest input meter does not automatically produce the cheapest finished result.

For a broader ranking across budget APIs, see Cheapest AI API in 2026. If your comparison extends from metered APIs to a consumer assistant, our Mistral Le Chat and Vibe cost guide covers that separate subscription decision.

Frequently asked questions

Is Claude Sonnet 5 still $2/$10 per million tokens?

Yes. Anthropic's August 10 update says the $2 input and $10 output introductory rates are permanent; the previously announced $3/$15 standard rates do not apply.

How long does Gemini 3.8 Flash introductory pricing last?

Google states $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026. The published standard price begins January 1, 2027 at $1.50/$7.50.

What is GPT-6 Luna's API price?

OpenAI's September 22 release notes list $0.10 input, $0.01 cached input and $0.50 output per million tokens for standard prompts up to the documented context threshold.

Sources