NotCheapMAKE EVERY CREDIT COUNT

OpenAI Fine-Tuning vs Together AI vs Fireworks AI: Real Fine-Tuning Cost per 100 Million Training Tokens in 2026

The short answer: this is no longer a genuine three-way comparison. OpenAI began winding down its self-serve fine-tuning API on May 7, 2026, and as of late September 2026 the only model still listed on OpenAI's standard fine-tuning platform is o4-mini, billed at a flat $100 per hour of wall-clock training time with no published per-token training rate at all — a fundamentally different pricing mechanism from the per-million-token rates this article's title implies, and one report puts the platform's full closure at January 2027. On the two providers still actively selling per-token fine-tuning, 100 million training tokens on an open-weight model under 16B parameters costs about $48 on Together AI (LoRA) or $120 (full fine-tuning), and $50 on Fireworks AI (LoRA) or $100 (full fine-tuning). How that compares to OpenAI's historical, now-discontinued rates depends on which OpenAI model and rate is used: roughly 1.6 to 10 times cheaper against the smaller and internally disputed historical rates (GPT-4o-mini, GPT-4.1-mini), but roughly 50 times cheaper on LoRA and 21 to 25 times cheaper on full fine-tuning against the $25-per-million-token historical rate for GPT-4o and GPT-4.1.

What changed: OpenAI's fine-tuning wind-down

OpenAI's self-serve supervised fine-tuning platform, which previously offered GPT-4o, GPT-4.1, and GPT-4.1-mini as fine-tunable base models, began winding down on May 7, 2026. By the time of the most recent sources reviewed, GPT-4.1 and GPT-4.1-mini no longer appeared on OpenAI's live fine-tuning pricing page at all, and o4-mini reinforcement fine-tuning (RFT) is the only offering still standing, priced at $100 per hour, prorated to the second, with no per-token training rate published. One industry tracker's article title states OpenAI's fine-tuning platform is closing entirely in January 2027. This article treats OpenAI's historical per-token training rates below as historical reference points for a discontinued product, not as a current, biddable option — a genuine reader budgeting fine-tuning today should not expect to fine-tune GPT-4o or GPT-4.1 on OpenAI's platform at all.

OpenAI's historical per-token rates (REPORTED, prior to wind-down, for reference only). GPT-4o: $25.00 per million training tokens, inference afterward at $3.75/$15.00 input/output per million (a markup over base GPT-4o rates). GPT-4.1: $25.00 per million training tokens in one source, or GPT-4.1-mini at either $5.00 or $0.80 per million across two conflicting sources — a 6x discrepancy this article cannot resolve from the sources reviewed. GPT-4o-mini: $3.00 per million. None of these rates can currently be paid for a new fine-tuning job on OpenAI's platform for the models in question, per the wind-down described above.

What the two live providers bill

Together AI (OFFICIAL, Together AI fine-tuning documentation). Priced per million training tokens by base model size, LoRA (adapter training, the default and cheapest method) versus full fine-tuning (updating every weight): LoRA, up to 16B parameters: $0.48/M. LoRA, 17B–69B: $1.50/M. LoRA, 70B–100B: $2.90/M. Full fine-tuning runs roughly 2.5x the LoRA rate at the same size tier (about $1.20/M at the smallest tier), and DPO (direct preference optimization) costs about 1.1x its matching SFT rate at every tier — the larger cost jump in Together's pricing is the training method (LoRA versus full), not the training objective (SFT versus DPO). No monthly subscription; pure usage-based billing with no minimum spend on standard jobs.

Fireworks AI (OFFICIAL, Fireworks AI pricing documentation, checked September 2026). Also priced per million training tokens by base model size: LoRA SFT, up to 16B: $0.50/M. LoRA SFT, 16.1B–80B: $3.00/M. LoRA SFT, 80B–300B: $6.00/M. LoRA SFT, over 300B: $10.00/M. LoRA DPO is 2x the matching LoRA SFT rate; full-parameter SFT is also 2x LoRA SFT; full DPO is 4x LoRA SFT (2x on top of full SFT's own 2x). Fireworks explicitly advertises no inference-price premium for serving a fine-tuned or LoRA-adapted model: a tuned model runs at the same per-token serverless rate as its base model, unlike OpenAI's historical markup on fine-tuned inference.

Cost per volume, smallest model tier (up to 16B parameters)

Formula: cost = training tokens (millions) × rate per million, one epoch. Real training runs typically use 2–4 epochs, multiplying this figure directly.

Training tokensTogether AI, LoRATogether AI, full fine-tuningFireworks AI, LoRA SFTFireworks AI, full SFT
10,000,000$4.80$12.00$5.00$10.00
100,000,000$48.00$120.00$50.00$100.00
1,000,000,000$480.00$1,200.00$500.00$1,000.00

At 3 training epochs — a common default — 100 million training tokens costs $144 on Together LoRA or $150 on Fireworks LoRA SFT, and both providers land within about 4% of each other at the LoRA tier despite being independently priced. Full fine-tuning costs roughly 2–2.5x the LoRA rate on both platforms, a consistent pattern across vendors: the training method is the larger cost lever than which of these two vendors you pick.

What OpenAI's historical rate would have cost, for scale

Formula: cost = training tokens (millions) × historical rate, at 100 million tokens, one epoch.

Model (discontinued or discontinuing)Historical training cost, 100M tokens
GPT-4o-mini$300
GPT-4.1-mini (lower reported figure)$80
GPT-4.1-mini (higher reported figure)$500
GPT-4o$2,500
GPT-4.1$2,500

Even at the lower end of OpenAI's disputed historical range (GPT-4.1-mini at $0.80/M), its smallest fine-tunable model cost roughly 1.7 times Together's LoRA rate for a comparably-sized open model, rising to about 10 times at the higher end of that same disputed range ($5.00/M); at the GPT-4o/GPT-4.1 tier ($25/M), OpenAI's historical rate was about 50 times the open-model LoRA rate and about 21 to 25 times the open-model full-fine-tuning rate. This gap, not any remaining live comparison, is likely the actual reason the market moved toward Together and Fireworks for fine-tuning work in 2026.

The epoch multiplier

Formula: effective training tokens = dataset tokens × epochs. A 100,000-token dataset trained for 3 epochs is billed as 300,000 training tokens, not 100,000 — a detail that catches teams off guard when comparing a quoted per-token rate to their raw dataset size rather than to what actually gets billed. For a reasoning-trace or multi-turn dataset, Fireworks' own guidance further multiplies effective tokens by roughly half the average number of turns per example, since longer conversational traces consume more training compute per epoch than a single input-output pair.

Inference after tuning: the cost that outlasts the training bill

Both Together and Fireworks serve fine-tuned models at or near base-model inference rates going forward (Fireworks explicitly states no premium; Together's own catalog reflects the same pattern for hosted LoRA adapters), while OpenAI's historical fine-tuned inference carried a real markup over base rates (GPT-4o's fine-tuned inference at $3.75/$15.00 input/output versus its base rate). A model fine-tuned once but served millions of times will accumulate far more in inference cost than training cost on any of these platforms; a provider with a fine-tuned-inference markup compounds that gap indefinitely.

Sensitivity

  1. Model size tier. Crossing from the ≤16B tier to the 17B–69B (Together) or 16.1B–80B (Fireworks) tier roughly triples the per-token training rate on both platforms.
  2. LoRA versus full fine-tuning. A consistent 2–2.5x cost multiplier for updating every weight instead of training adapters, on both live providers.
  3. Epoch count. Directly and linearly multiplies the training bill; a 5-epoch run costs 5 times a 1-epoch run at the same dataset size.
  4. Which OpenAI historical rate is real. The $5.00 versus $0.80 per-million conflict for GPT-4.1-mini is unresolved in the sources reviewed and moot in practice given the wind-down, but illustrates how quickly discontinued-product pricing data degrades in reliability.

Budgeting traps

  • Assuming OpenAI fine-tuning is still a live option for GPT-4o or GPT-4.1. It is winding down; only o4-mini reinforcement fine-tuning remains, at an hourly rate with no per-token equivalent.
  • Quoting a per-token training rate against raw dataset size instead of dataset size times epochs. This routinely understates the real bill by 2–5x.
  • Ignoring fine-tuned inference markup. A provider with a training-time discount but an inference-time premium can cost more over the model's serving life than a provider priced the other way around.
  • Comparing LoRA and full fine-tuning rates as if they produce equivalent results. They are different training methods with different cost and, per vendor positioning, different quality/adaptability tradeoffs, not the same underlying job at two prices.

What to ask before you buy

If evaluating OpenAI, ask directly whether the specific base model you want is still available for self-serve fine-tuning given the ongoing wind-down, rather than budgeting off a historical rate. Ask Together or Fireworks for their current rate at your specific model size tier and confirm whether your dataset's effective token count (after epochs and, for conversational data, turn-count adjustments) matches what you're mentally budgeting.