StepFun AI Pricing 2026: What Does Step Actually Cost?
Verification date: September 19, 2026. Prices, plan rules and promotions change.
Direct answer
StepFun sells two current sparse Mixture-of-Experts models that each activate about 11B of roughly 196B parameters per token. Its own price list, in US dollars per million tokens:
| Model | Input (cache miss) | Input (cache hit) | Output |
|---|---|---|---|
step-3.5-flash | $0.10 | $0.02 | $0.30 |
step-3.7-flash | $0.20 | $0.04 | $1.15 |
The prices are listed on StepFun's pricing page.
- A heavy month is cheap. 400M input + 100M output tokens costs $70 on Step 3.5 Flash and $195 on Step 3.7 Flash (CALCULATION).
- Output is where the money goes. Output costs 3× input on 3.5 Flash and 5.75× input on 3.7 Flash. Prompt caching bills cached input at 20% of the normal price, but it cannot discount your output.
- Both models have open weights under Apache 2.0, including Step 3.7 Flash.
- Step Plan credits are not tokens, and plan totals cannot be turned into a known token allowance. StepFun documents roughly $1 ≈ 7M Credits, but not how each model's tokens convert to Credits. We show only a nominal, implied credit value for each tier, and stop short of a token-level break-even.
- Self-hosting rarely saves money at these prices. In our modeled scenario, a rented 2-GPU setup (not a StepFun-documented configuration) would need to move about 18–39B tokens a month to match the API bill. The best reasons to self-host are control and data, not cost.
- Cheap tokens are not the same as equal quality. Activating 11B parameters explains the low price. It does not prove the model matches pricier ones on your work.
1. Current models and what StepFun says about them
| Model ID | Role (StepFun's wording) | Parameters | Context | Input | Weights |
|---|---|---|---|---|---|
step-3.7-flash | Flagship multimodal reasoning model | 198B total (196B language backbone + 1.8B vision encoder), about 11B active | 256K | Text, image, video | Open, Apache 2.0 |
step-3.5-flash | High-speed reasoning model for agent and coding tasks | 196.81B total, about 11B active | 256K (262,144) | Text | Open, Apache 2.0 |
step-3.5-flash-2603 | Version tuned for high-frequency agent scenarios; same price as 3.5 Flash | Not stated in the cited public documentation | 256K (per Artificial Analysis) | Text | Not specified by the vendor |
StepFun's documentation and Hugging Face model cards list the model IDs, roles, parameters, context and weights. The 2603 context figure is reported by Artificial Analysis. Step 3.5 Flash launched January 29, 2026; Step 3.7 Flash launched in late May 2026.
Two API details matter for budgeting. Step 3.7 Flash exposes three reasoning-effort settings (low, medium, high). And StepFun's own model card for 3.5 Flash lists "token efficiency" as a known limitation: it currently needs longer generation trajectories than Gemini 3.0 Pro to reach comparable quality. Longer trajectories mean more output tokens, which is your expensive line item.
Rate limits. The free tier (cumulative top-up $0) allows 5 concurrent requests, 10 requests per minute and 5M tokens per minute. Limits step up with cumulative top-ups: $15 unlocks 100 concurrent / 1,000 RPM / 20M TPM, $70 unlocks 200 / 5,000 / 30M, with higher thresholds up to $1,500.
Images and video. StepFun bills images at a flat 400 tokens each. At Step 3.7 Flash's $0.20 input rate, 1,000 images cost 0.4M tokens × $0.20 = $0.08 (CALCULATION). The image API has a detail setting (low by default, or high); public documentation does not specify whether it changes billing. Video is supported (MP4, under 128 MB, up to 5 minutes) but no video-specific billing rule is published. Both are.
2. Pay-as-you-go economics
Formula (prices per 1M tokens, volumes in millions):
Cost = (Input × HitShare × P_hit) + (Input × (1 − HitShare) × P_miss) + (Output × P_out)
StepFun bills tokens used × model unit price, drawing on free credit first and then your paid balance. It is a prepaid balance: you top up before you consume, and payments are generally non-refundable (StepFun's Terms of Service, effective August 7, 2026). Cache pricing is applied only where noted below.
Real workloads
| Workload | Formula (Step 3.5 Flash) | 3.5 Flash | 3.7 Flash |
|---|---|---|---|
| A: 1M in + 250K out | 1 × 0.10 + 0.25 × 0.30 | $0.175 | $0.4875 |
| B: 10M in + 3M out | 10 × 0.10 + 3 × 0.30 = 1.00 + 0.90 | $1.90 | $5.45 |
| C: 400M in + 100M out (monthly) | 400 × 0.10 + 100 × 0.30 = 40 + 30 | $70.00 | $195.00 |
| D: 100M in + 20M out, 80% cache hits | 80 × 0.02 + 20 × 0.10 + 20 × 0.30 = 1.60 + 2.00 + 6.00 | $9.60 | $30.20 |
| D with no caching | 100 × 0.10 + 20 × 0.30 | $16.00 | $43.00 |
All CALCULATION on prices. The 80% hit share is a MODELING ASSUMPTION. The 3.7 Flash column uses $0.20 / $0.04 / $1.15 (for example D: 80 × 0.04 + 20 × 0.20 + 20 × 1.15 = 3.20 + 4.00 + 23.00).
For scale: workload C averages about 190 tokens per second, around the clock.
3. Caching economics
StepFun's prompt-cache page says caching is automatic for prompts over 256 tokens and supported on step-3.7-flash, step-3.5-flash and step-3.5-flash-2603. Cached tokens are billed at 20% of the model's normal token price, and there is no extra fee for caching. It can cover the whole message array (system prompt, history, tool calls and results), images, and video. The caveats:
- Hits are not guaranteed. Eviction is least-recently-used, and entries are dropped faster at peak traffic. StepFun says explicitly that it cannot force every request to hit.
- The prefix must stay identical. Put static content first and dynamic content last.
- You can see hits.
usage.cached_tokensin the response shows how many tokens were served from cache.
What the discount is worth. Workload D under different hit rates (CALCULATION):
| Cache-hit share | 3.5 Flash | 3.7 Flash |
|---|---|---|
| 0% | $16.00 | $43.00 |
| 50% | $12.00 | $35.00 |
| 80% | $9.60 | $30.20 |
| 95% | $8.40 | $27.80 |
Even at 95% hits, output is still $6.00 of the 3.5 Flash bill and $23.00 of the 3.7 Flash bill. Caching helps most on long, repetitive prompts and agent loops, and least on output-heavy work. Watch cached_tokens for your own workload before you bank on any hit rate.
4. Step Plan: credits, tiers and what we can and cannot calculate
Step Plan is StepFun's subscription. It gives you a monthly allowance in Credits, spent across models, and you connect through a dedicated endpoint (api.stepfun.ai/step_plan/v1), not the standard API.
| Tier | Monthly | Quarterly | Yearly | Credits per month |
|---|---|---|---|---|
| Flash Mini | $6.99 | $18.99 | $69.99 | 400M |
| Flash Plus | $9.99 | $26.99 | $95.99 | 1,600M |
| Flash Pro | $29 | $79 | $289 | 8,000M |
| Flash Max | $99 | $269 | $989 | 40,000M |
(Step Plan overview, checked September 19, 2026). Booster packs add Credits for active subscribers: $6.99 for 400M, $9.99 for 1,600M.
How Credits convert. StepFun says Credits convert to payment at approximately $1 ≈ 7M Credits, meaning 7 million Credits correspond to about $1 of model usage. It also says each model's usage is "converted into Credits" and deducted. StepFun does not publish credits per token by model, whether input, output and cached tokens convert differently, or how multimodal usage converts. One Credit is not one token.
The conversion is only partly documented, so a token-level break-even cannot be calculated. The allowances have a nominal value only if "$1 of model usage" means list-price usage. These are implied credit values, not verified token values or savings:
| Tier | Price | Credits | Nominal implied value at 7M Credits/$ | Nominal multiple of price (not verified savings) |
|---|---|---|---|---|
| Flash Mini | $6.99 | 400M | ≈ $57 | 8.2× |
| Flash Plus | $9.99 | 1,600M | ≈ $229 | 22.9× |
| Flash Pro | $29 | 8,000M | ≈ $1,143 | 39.4× |
| Flash Max | $99 | 40,000M | ≈ $5,714 | 57.7× |
CALCULATION on an approximate, conditional conversion. StepFun says the monthly allowance provides more usage than pay-as-you-go at the same spend, but it does not quantify that. The multiples above are nominal credit-value ratios, not verified token allowances, token values or savings, and we do not size any workload against a tier.
Rules that limit who this is for:
- Personal use only. The Step Plan agreement limits the service to the personal use of the subscribing account and bars providing it to third parties. It works only with supported tools and integrations (OpenClaw, Claude Code, Cline, Kilo Code, Roo Code and others). You cannot build a product's cost base on it.
- Credits expire and reset monthly and do not roll over. Fees are non-refundable, and StepFun may adjust names, prices, billing and quota rules.
- It has changed recently. As of July 24, 2026, StepFun moved from a request-based "Coding Plan" (5-hour and weekly limits) to the credit-based plan its upgrade notice calls the "Token Plan." Plan terms are a moving target.
- Payment is through Stripe or account balance.
Before you subscribe, run a normal week of your own work and compare the Credits deducted against the tokens and dollars you actually used. That is the only way to learn the real conversion.
5. The open weights and the license
Step 3.5 Flash is Apache 2.0. The Hugging Face model card declares license: apache-2.0, and StepFun's GitHub repository includes the standard, unmodified Apache License 2.0 text. StepFun's Step 3.7 Flash model card also states Apache 2.0. The Hugging Face repository for 3.5 Flash holds no standalone LICENSE file; the license lives in the card metadata and on GitHub.
What that means, precisely. Commercial use, modification and redistribution are permitted under the Apache 2.0 terms. Apache 2.0 is permissive but not condition-free:
- Recipients of copies must get a copy of the license, and modified files must carry prominent change notices.
- You must keep copyright, patent, trademark and attribution notices, and carry any NOTICE file text.
- It grants a patent license, which terminates for you if you sue over the Work's patents.
- It grants no trademark rights, and the software is provided "as is" without warranty.
Do not call this "unrestricted." The Apache license covers the weights and code. Using StepFun's hosted API is governed by separate Terms of Service. Those Terms (effective August 7, 2026) permit using the services to support products you provide to your own users, but they also grant StepFun a broad license to prompts, outputs and API call logs, including for improving and promoting its services, "to the fullest extent permitted" by law. A separate Data Processing Agreement covers personal data you process for your own customers. Organizations handling customer personal data should review the DPA before choosing the hosted API. This is not legal advice, and it is one of the main non-price reasons to self-host.
6. Self-hosting economics
6.1 What the weights require
| Item | Figure | Source |
|---|---|---|
| Step 3.5 Flash, BF16 repository | About 399 GB | Hugging Face file listing |
| Step 3.5 Flash, INT4 GGUF | 111.5 GB + about 7 GB runtime overhead. StepFun's stated minimum: 120 GB of VRAM or unified memory; recommended: 128 GB | Model card |
| Step 3.7 Flash, INT4 GGUF (Q4_K_S) | 111.5 GB + 3.97 GB vision projector + about 7 GB overhead ≈ 122.5 GB; same 120 GB stated minimum, 128 GB recommended | Model card; sum is CALCULATION |
| Step 3.5 / 3.7 Flash, BF16 or FP8 in vLLM / SGLang | Official recipes use tensor parallelism across 8 GPUs | Model cards |
| Step 3.7 Flash, NVFP4 | Official recipe uses 4 GPUs with --max-model-len 8192 in the example | Model card |
Local hardware constraints:
- "INT4 fits a 128GB Mac Studio." This is documented by StepFun for single-user local inference, with modest context. The official example commands use 16K–32K contexts, not 256K. It is a local-inference statement, not a production-serving recommendation.
- "Or a dual-GPU node, e.g. 2×80GB H100." StepFun publishes no 2-GPU recipe or benchmark. 160 GB of memory exceeds the 111.5 GB of INT4 weights, but StepFun's data-center recipes use 8 GPUs. Treat 2×80GB as an illustrative setup, not a vendor recommendation.
6.2 Raw weights are not the whole bill: KV cache
From Step 3.5 Flash's config.json: 45 layers, 12 of them full attention and 33 sliding-window (window 512), with 8 KV groups and a head dimension of 128. KV cache per token on the full-attention layers is 12 × 2 × 8 × 128 values (CALCULATION):
| Context per request | KV cache, BF16 | KV cache, FP8 |
|---|---|---|
| 32K tokens | ≈ 1.6 GB | ≈ 0.8 GB |
| 128K tokens | ≈ 6.4 GB | ≈ 3.2 GB |
| 256K tokens | ≈ 12.9 GB | ≈ 6.4 GB |
The sliding-window layers add only about 70 MB per sequence at BF16. This ignores activations, MTP and framework overhead, so treat it as a floor. On a 128 GB machine, weights plus overhead leave roughly 9–10 GB, so long contexts and concurrent users get squeezed quickly. On 2×80GB with INT4 weights, about 41–49 GB would remain, enough for a few long sequences or many short ones if a serving stack supports it. That is a modeling assumption; StepFun documents no 2-GPU recipe.
6.3 Sensitivity model (all MODELING ASSUMPTION)
We do not claim a fixed minimum spec or one break-even point. Instead: how much traffic must a setup move for its fixed cost to equal your API bill?
- Workstation, 128 GB unified memory: $150–$300 a month all-in (roughly $5,000–$10,000 of hardware over 36 months, plus power). Hardware pricing varies by configuration and vendor; these figures are illustrative assumptions.
- 2× 80GB-class GPUs, rented (not a StepFun-documented configuration): 2 GPUs × $1.76–$3.70 per GPU-hour × 730 hours.
- 8× 80GB-class GPUs, rented: 8 GPUs × the same rates. StepFun's own vLLM and SGLang recipes use 8-way tensor parallelism but do not specify the GPU model or memory per GPU, so the 80GB-class pricing is our assumption.
- GPU rates are borrowed from July 2026 H200 rental datapoints ($1.76 spot-like, $3.70 on-demand-like), used as proxies for H100 rates; actual prices vary by vendor.
API-equivalent volume = Monthly fixed cost ÷ API blended $/1M. Sustained tokens/s = Volume × 1,000,000 ÷ (730 × 3,600). The blended API price for Step 3.5 Flash at an 80:20 input:output mix is 0.8 × $0.10 + 0.2 × $0.30 = $0.14 per 1M with no caching, or $0.0888 if 80% of input hits the cache.
| Setup | Fixed cost per month | Break-even volume, uncached API | Sustained blended tokens/s needed | Break-even volume, 80% cached API |
|---|---|---|---|---|
| 128 GB workstation | $150–$300 | 1.1–2.1B | 408–815 | 1.7–3.4B |
| 2×80GB rented | $2,570–$5,402 | 18–39B | 6,984–14,683 | 29–61B |
| 8×80GB rented | $10,278–$21,608 | 73–154B | 27,937–58,730 | 116–243B |
CALCULATION. No throughput figure for any of these setups is documented, so we do not claim any of them can reach the numbers in the fourth column. For context, workload C is 0.5B tokens a month, or about 190 tokens per second.
What the table says:
- The workstation loses on capacity, not just price. Even at an assumed 60 output tokens per second (a MODELING ASSUMPTION; StepFun publishes no Mac speed), it produces at most about 158M output tokens a month. The API would bill that output at about $47 on Step 3.5 Flash. The break-even needs 408–815 sustained tokens a second.
- Rented GPUs need enormous, steady volume. The 2-GPU setup's cheapest month costs about 37× workload C's $70.
- Step 3.7 Flash shifts the math, but not enough. Its blended API price is $0.39 per 1M (80:20, uncached), about 2.8× the 3.5 Flash figure. The same 2×80GB spot-like bill breaks even at about 6.6B tokens a month (about 2,500 tokens a second). That is still 13× workload C's 500M tokens.
When self-hosting is still the right call: your data cannot go to a third-party API (see the Terms discussion above); you need a fixed, private, offline or air-gapped deployment; you want to fine-tune (Apache 2.0 allows it); or your steady volume is in the tens of billions of tokens a month. Below that, StepFun's price list is hard to beat.
7. Third-party hosts and price parity
For Step 3.5 Flash, StepFun's price list and OpenRouter's SiliconFlow route show the same base rate, so this comparison does not show a material price spread.
| Model | Host (via OpenRouter) | Input | Cache read | Output | Label |
|---|---|---|---|---|---|
step-3.5-flash | SiliconFlow (the only provider listed) | $0.10 | not shown | $0.30 | THIRD-PARTY HOST PRICE |
step-3.7-flash | StepFun | $0.20 | $0.04 | $1.15 | |
step-3.7-flash | NovitaAI | $0.20 | $0.04 | $1.15 | THIRD-PARTY HOST PRICE |
step-3.7-flash | DeepInfra (20% off badge) | $0.16 | $0.032 | $0.92 | THIRD-PARTY HOST PRICE |
Three cautions:
- OpenRouter and SiliconFlow are one route for 3.5 Flash, not two independent confirmations. OpenRouter lists SiliconFlow as the sole provider, and StepFun itself is not listed for that model there. This route alone does not independently confirm StepFun's price.
- DeepInfra lists a 20% promotion on Step 3.7 Flash, not Step 3.5 Flash. Promotions can end without notice.
- Aggregator listings can conflict. One price-tracking site lists DeepInfra at about $0.09 input for 3.5 Flash (not confirmed by DeepInfra and roughly 10% under list, so it is not treated as a material spread), while another shows $0.40 output against StepFun's $0.30 list price. The latter figure appears erroneous.
International access. StepFun's own global platform (platform.stepfun.ai, API host api.stepfun.ai) lists prices in US dollars and is separate from its China platform, which requires a +86 phone number. The global Terms of Service allow commercial use to support products for your own users, and Step Plan can be paid through Stripe. StepFun's public materials do not specify signup and payment availability by country, tax and VAT handling, or whether every card is accepted. English documentation and USD prices do not prove every country can complete checkout..
8. Build & Earn: an AI refactoring assistant on the pay-as-you-go API
Scenario (all MODELING ASSUMPTION except the token prices): a codebase-refactoring assistant for independent developers, sold at $15 per customer per month. One refactor = 40K input tokens and 10K output tokens, the size of one third-party example (Kilo Code's illustration of a 10,000-line codebase). A heavy user runs 150 refactors a month.
Recalculation. Converting thousands of tokens to millions correctly: 40K input = 0.04M and 10K output = 0.01M.
- Step 3.5 Flash: 0.04 × $0.10 + 0.01 × $0.30 = $0.004 + $0.003 = $0.007 per refactor. 150 refactors × $0.007 = $1.05.
- Step 3.7 Flash: 0.04 × $0.20 + 0.01 × $1.15 = $0.008 + $0.0115 = $0.0195 per refactor, or $2.925 for 150.
| Case (per customer per month) | Model cost | Model-cost gross margin at $15 |
|---|---|---|
| 3.5 Flash, 150 refactors | $1.05 | 93.0% |
| 3.5 Flash, 150 refactors, 80% cache hits | $0.67 | 95.6% |
| 3.5 Flash, output tokens tripled (30K/refactor) | $1.95 | 87.0% |
| 3.5 Flash, 600 refactors | $4.20 | 72.0% |
| 3.5 Flash, 1,500 refactors ("unlimited" abuse) | $10.50 | 30.0% |
| 3.7 Flash, 150 refactors | $2.93 | 80.5% |
| 3.7 Flash, output tokens tripled | $6.38 | 57.5% |
| 3.7 Flash, 1,500 refactors | $29.25 | −95% |
CALCULATION. Read "model-cost gross margin" literally: it counts only model API fees. It excludes labor, infrastructure, payment fees, customer acquisition, support and other business costs. Do not call it profit margin.
What to take from it:
- Avoid "unlimited." At 1,500 refactors, the 3.5 Flash model-cost gross margin falls to 30% and the 3.7 Flash model-cost gross margin turns negative (−95%). Cap usage or meter it.
- Real agent tasks are many model calls, not one. Each tool step is billed as a separate request with a growing context, so the 40K/10K assumption could be too low.
- Your model choice moves model-cost gross margin by 12+ points. 3.7 Flash adds image and video input; 3.5 Flash is text-only and about 64% cheaper per refactor here.
- Use pay-as-you-go, not Step Plan, for the product. Step Plan is for personal use with supported tools.
10. FAQ
How much does the StepFun API cost? Step 3.5 Flash is $0.10 input, $0.02 cached input and $0.30 output per million tokens. Step 3.7 Flash is $0.20, $0.04 and $1.15.
Is Step 3.5 Flash open source? Its weights and code are released under Apache 2.0, which permits commercial use with conditions such as keeping the license and notices. StepFun's hosted API is governed by separate terms.
Is Step 3.7 Flash open weights? Yes. Its Hugging Face model card lists Apache 2.0 licensing.
Do Step Plan credits equal tokens? No. StepFun says about 7M Credits equal $1 of model usage, but it does not publish per-model credit rates, so plan totals cannot be converted into a known token allowance.
Can I run Step 3.5 Flash on a 128GB Mac Studio? StepFun lists 128 GB of unified memory as recommended for INT4 local inference (120 GB minimum), with modest context. It is not a production-serving recommendation.
Does StepFun work outside China? It runs a separate global platform with USD pricing. Signup and payment availability may vary by country; check eligibility and accepted payment methods before opening an account.
Is OpenRouter cheaper than StepFun directly? The displayed Step 3.5 Flash base rates are the same. One host currently discounts Step 3.7 Flash by 20%.
Does 11B active parameters mean it is as good as bigger models? No. Low active-parameter count explains the price, not the quality. Test it on your own tasks and count output tokens, not just price per token.