NotCheapMAKE EVERY CREDIT COUNT

Meta's consumer subscription, hosted developer API, and downloadable model weights use separate billing and licensing arrangements. Meta AI remains free for everyday use; Meta One adds paid benefits and higher usage for selected compute-intensive features. For developers, Meta offers Muse models through its hosted API, while Llama and Muse Glimmer deployment costs depend on the chosen host or hardware.

Meta AI / Llama 2026: What Does "Free" AI Actually Cost?

DIRECT ANSWER

Short answer

"Meta AI is free" is four different claims. Only one of them is a promise about money.

  1. Meta AI consumer access: Meta says Meta AI is free for everyday use, while it is testing usage limits for some compute-intensive features. Meta One is an optional subscription that adds benefits and more usage for certain features; the published USD monthly prices in Meta's announcement are $7.99 per month for Core and $19.99 per month for Premium. Availability, benefits and the price shown can vary by region, platform, account and rollout stage. (Meta AI access and limits; Meta One announcement, 15 September 2026)
  2. Downloadable models are not all "open source" in the same way. Llama 4 Scout and Maverick are downloadable under a custom Community License with conditions . Muse Glimmer 30B (10 Aug 2026) is under Apache 2.0 . Muse Spark has no released weights .
  3. Third-party APIs sell the same open models by the token: about $0.10 / $0.30 per million tokens for Llama 4 Scout and about $0.20 / $0.60 for Maverick .
  4. Self-hosting removes the per-token price and adds GPU, ops and licence obligations. On our assumptions it only beats a cheap API at very high sustained volume .

We do not invent a native Meta price for Llama. Meta's hosted Llama API was retired on 6 Jul 2026 . Meta's current paid API prices Muse Spark (text), Muse Image, Muse Voice Transcribe and SAM 3.1 , but no Llama model.

Current Meta model map (verified 26 Sep 2026)

FamilyWhat it isWeights?Where you pay
Muse Spark (1.0 Apr, 1.1 Jul, 1.2 Aug, 1.3 Sep 2026)Closed multimodal reasoning model from Meta Superintelligence LabsNo. Zuckerberg said on 10 Aug 2026 that Spark 1.2's weights would be released "soon"; no date given and no release found as of early Sep (not confirmed on a Meta primary page)Meta AI apps (free); Meta Model API (paid, public preview)
Muse Glimmer 30BDense 30B multimodal agentic model distilled from Muse Spark, released 10 Aug 2026Yes, Apache 2.0Free to download; hosted by third parties
Llama 4 Scout / MaverickMoE models, 109B and 400B total parameters, 17B active; released 5 Apr 2025Yes, Llama 4 Community LicenseFree to download; hosted by many providers
Llama 4 BehemothAnnounced, never publicly releasedNo—
Llama 3.xOlder releases; the Llama FAQ points users to Llama 3.3 70B as a replacement for 3.1 70BYes, earlier community licencesThird-party hosts

Meta's own Llama API public preview was wound down on 6 Jul 2026, per Meta's deprecation notice, which says Llama models remain available to download and to run through third-party providers . The Meta Model API (api.meta.ai) serves Muse models and no Llama models . Any guide still telling you to call Meta's Llama API is out of date.

1. Meta AI consumer access

Meta AI is Meta's consumer assistant, available through Meta's supported consumer experiences. Meta describes everyday Meta AI use as free. It also says it is testing daily or monthly limits for some compute-intensive features; when an account reaches a limit, Meta may offer a subscription for additional use before the limit resets. Meta does not publish one universal quota in this announcement, so do not assume a fixed number of prompts or generations. (Meta AI assistant FAQ)

Meta One subscription prices

Meta published these USD monthly prices for its consumer subscription plans on 15 September 2026:

PlanPublished USD monthly priceWhat the announcement says
Meta One Core$7.99/monthCore subscription benefits, including additional usage for some compute-intensive AI features
Meta One Premium$19.99/monthCore benefits plus higher usage and additional subscription benefits

These are the published USD monthly prices in Meta's announcement, not a guarantee of the amount shown to every user. Meta says availability and benefits are rolling out and may vary by region, app, account and rollout stage. Check the subscription screen for the account and platform you will use before budgeting. Meta AI everyday access is not stated to require a Meta One subscription; Meta One is an optional paid tier for additional benefits and eligible usage. (Meta's Meta One announcement; Meta AI assistant FAQ)

Do not confuse consumer access with developer costs

Meta One is a consumer subscription. It is not the rate card for API calls, and it does not make every Meta model or developer service unlimited.

Meta's separate Meta Model API is a developer service with usage-based rates and rate limits. Its documentation describes hosted Muse models and bills API use under the developer pricing schedule. Treat that bill as separate from a Meta One consumer subscription. (Meta Model API overview; Meta Model API pricing and rate limits; API billing help)

Llama is a separate model family and developer ecosystem. Meta distributes Llama model weights under its applicable licence; developers may use a third-party host or operate their own deployment, where costs depend on the host, hardware, usage and operations. Those costs are not Meta One subscription prices and should not be represented as Meta Model API rates. Check the current Llama terms and the chosen host before estimating commercial or infrastructure costs. (Meta Llama developer documentation)

2. Open-weight models: what the licences actually say

Llama 4 Community License (effective 5 Apr 2025)

Read from the licence text on Meta's Hugging Face repository:

  • Grant: non-exclusive, worldwide, royalty-free licence to use, copy, distribute, modify and create derivatives.
  • Attribution: if you distribute the materials, or a product or service containing them, you must include the licence, prominently display "Built with Llama", and keep a Notice file with Meta's copyright line.
  • Naming: if you use Llama materials or outputs to create, train or improve an AI model that you distribute, its name must begin with "Llama".
  • Scale threshold: if, on the Llama 4 release date, your (or your affiliates') products had more than 700 million monthly active users in the preceding month, you must request a licence from Meta, which it may grant "in its sole discretion". Until then you have no rights under the agreement.
  • Acceptable Use Policy is incorporated by reference. Use must comply with law, including trade compliance.
  • Termination triggers: suing Meta over IP in relation to the materials or outputs ends your licence; you agree to indemnify Meta for third-party claims arising from your use.
  • Ownership: you own derivatives and modifications you make, subject to Meta's ownership of the underlying materials.
  • Governing law: California, with exclusive California courts.
  • Meta's model card says the licence allows using outputs to improve other models, including distillation , subject to the naming rule above.

What that means. Llama 4 is commercially usable for nearly every company, but it is a custom licence with conditions, not an OSI-style open-source licence. We do not call it "open source". The 700M threshold is measured on one date, so it is a gate for very large companies, not a pay-as-you-grow fee.

EU and multimodal caveat. Model licences and regional use policies can impose separate conditions. Before deploying a multimodal model in a regulated or cross-border product, review Meta's current model-specific terms and applicable regional requirements; do not assume a model card alone covers every use case.

Muse Glimmer 30B: Apache 2.0

  • Meta's developer page and the Hugging Face model card list the licence as Apache 2.0, with an out-of-scope clause for use that violates law or Apache terms .
  • A secondary writer notes Meta layers a separate usage policy on top, which "Apache 2.0" alone would not require . Read it before shipping in a regulated product.
  • No monthly-active-user cap appears in Apache 2.0, which is the material difference from Llama .
  • Access to the base model repository on Hugging Face may require accepting access terms .

Muse Spark: no weights, no self-hosting

Meta's paid API is the only way to run Spark. Meta has said it intends to open weights for a version in future; no date has been given .

3. Third-party API prices for the exact same model

These are hosted-provider list prices per million tokens (input / output). Meta does not set them, and they move.

ModelCheapest listedRange across hostsLabel
Llama 4 Scout$0.10 / $0.30 (DeepInfra; Vercel and OpenRouter list the same)Input from $0.10 to $0.17 or more; up to 2.7x variation across providers; Amazon $0.17 input (benchmarking-site snapshot)
Llama 4 Maverick$0.1875 / $0.6525 on OpenRouter; $0.20 / $0.60 at VercelInput mostly $0.15–$0.30
Muse Glimmer 30B$0.30 / $1.10 (OpenRouter, four providers: Phala, DeepInfra, Fireworks, Together); cache read $0.04Provider prices differ

Context windows differ by host. Scout's native context is 10 million tokens, but hosts cap it at roughly 128K–1M in practice, so the headline context is mostly a self-hosting feature . Hosts also drop models: Groq's own deprecation page shows Llama 4 Scout shut down on 17 Jul 2026 and Llama 4 Maverick deprecated after a Feb 2026 notice, so Groq is no longer a valid Llama 4 price reference . Older comparison pages that list Groq prices for Llama 4 are stale.

Meta Model API: prices and limits (Meta's own documentation)

ProductPriceNotesLabel
Muse Spark Standard (1.3, 1.2, 1.1)$1.25 input / $4.25 output / $0.15 cached input per 1M tokensPrompts are not used for training; no long-context premium; web-search grounding $2.50 per 1,000 queries
Muse Spark Contributor (1.3, 1.2)$0.10 / $0.20 / $0.002 cached per 1M tokensDiscount in exchange for permission to use your prompts and completions to train future Meta models
Muse Image$0.01 per generated imageFlat price; built-in search included
Muse Voice Transcribe$0.18 per hour of audioBilled per second processed; no discounted training tier
SAM 3.1$2.50 per 1,000 images; $0.20 per 1,000 video framesPer image or frame, not per token

Rate limits (per team, not per key): Standard tier 3,000 requests per minute and 4,000,000 tokens per minute; Contributor tier 100 requests per minute and 3,000,000 tokens per minute; Muse Image 150 requests per minute; background submissions 600 per minute; Muse Voice Transcribe 128 concurrent streams and 16,000 streams per hour.

Availability. Meta's developer page describes the API as "Public preview. Now available with expanded global access" . Meta's public pages do not list which countries are supported . Press coverage of the July launch described a US-developer-only preview ; that has been superseded by Meta's current wording. Reports of $20 in free credit for new accounts come from press coverage of the launch . Preview status means no stable-pricing or SLA commitment is published . Muse Glimmer is listed in Meta's docs as an open-weight model that runs on your own hardware, not on the Meta Model API .

4. Real workload calculations

Method. Workload = tokens per month at a 3:1 input-to-output ratio. Blended price = (3 × input + output) ÷ 4. Prices are the list prices above.

ModelBlended $/M100M tokens a month1B tokens10B tokens
Llama 4 Scout (0.10 / 0.30)$0.15$15$150$1,500
Llama 4 Maverick (0.20 / 0.60)$0.30$30$300$3,000
Muse Glimmer 30B (0.30 / 1.10)$0.50$50$500$5,000
Muse Spark 1.3 Contributor (0.10 / 0.20)$0.125$12.50$125$1,250
Muse Spark 1.3 (1.25 / 4.25)$2.00$200$2,000$20,000

Using OpenRouter's Maverick price ($0.1875 / $0.6525) instead gives $0.304 blended, or $304 for 1B tokens .

Those totals ignore prompt caching, batch discounts and output-heavy workloads, and they exclude reasoning tokens for Spark and Glimmer, which are billed as output. Reasoning-heavy runs can multiply output volume, so a 3:1 ratio understates the Spark and Glimmer bill . "Free" has a price per token everywhere but your own machine.

5. Self-hosting: what Meta documents, and where we cannot go further

Hardware Meta documents.

  • Llama 4 Scout fits on a single H100 with on-the-fly int4 quantization; released as BF16.
  • Llama 4 Maverick ships as BF16 and FP8; the FP8 weights fit on a single H100 DGX host (typically eight GPUs).
  • Muse Glimmer 30B: Meta describes it as running on a single GPU or Mac . Unsloth's guide says Dynamic quants run on about 18 GB of RAM or VRAM , and a third-party writeup says a quantized build fits a single 24 GB consumer GPU .

Throughput (tokens per second per GPU) is not published in the available primary documentation, so a true self-hosting cost per token cannot be computed. Instead, compare the sustained volume at which a rented GPU costs the same as the API.

Everything in this table is a MODELING ASSUMPTION / CALCULATION, not a provider cost per token. Assumptions: rented GPUs run 730 hours a month at $2, $3 or $4 per H100-hour, or $0.50–$2 per hour for a 24 GB-class GPU. These rental rates are assumptions, not researched prices. Parity volume = monthly GPU cost ÷ blended API price.

DeploymentGPU cost per monthBlended API priceParity volumeSustained rate needed
Scout on 1 H100 at $2 / $3 / $4 per hour$1,460 / $2,190 / $2,920$0.15 per M9.7B / 14.6B / 19.5B tokens3,704 / 5,556 / 7,407 tokens a second
Maverick on 8 H100 at $2 / $3 / $4$11,680 / $17,520 / $23,360$0.30 per M38.9B / 58.4B / 77.9B14,815 / 22,222 / 29,630
Glimmer on one 24 GB-class GPU at $0.50 / $1 / $2 per hour$365 / $730 / $1,460$0.50 per M0.73B / 1.46B / 2.92B278 / 556 / 1,111

Reading the table. If your total volume is below those parity figures, the API is cheaper even before you add engineering time, monitoring, idle capacity and redundancy. Self-hosting Scout only wins above roughly 10–20 billion tokens a month and if your hardware can actually sustain thousands of tokens a second, which remains an estimate. Glimmer's small footprint makes self-hosting viable at a lower volume (around 1 billion tokens a month at a $1 GPU-hour), but you still need a 24/7 machine or a workload that can use one. Against Meta's own Spark price, Glimmer self-hosting looks attractive on paper, but Glimmer is a different, smaller model and we make no claim about equivalence.

Non-token costs of self-hosting. Egress and storage, model updates, guardrails (Meta's card recommends deploying Llama Guard, Prompt Guard and Code Shield alongside the models ), and staff time. Also, if you distribute a Llama-based product, the attribution and naming obligations above apply .

Hidden-cost analysis

  • Token prices differ between hosts by up to about 2.7x for the same Scout model. Price-shopping matters more than picking Scout versus Maverick.
  • Context caps. A 10M-token headline may be 128K on the host you can afford .
  • Reasoning tokens on Spark and Glimmer are billed as output .
  • Preview risk. Meta's API has no published SLA or stable-pricing commitment ; a model called "Muse Spark" changed version four times in five months .
  • Data terms. Meta's cheap Contributor tier trades price for training rights over your prompts . Do not send client or personal data.
  • Free-tier data. Any consumer "free" Meta AI use is subject to Meta's privacy terms .

International access

  • Meta Model API: public preview with "expanded global access" per Meta ; the supported-country list was not found . Launch-time coverage said US-only .
  • Open weights: downloadable globally, subject to trade-compliance clauses in the licences . Llama's EU multimodal restriction question is unresolved above .
  • Third-party hosts: availability and data-residency vary by provider; check each.
  • Consumer Meta AI: country availability was not researched .

Licensing and commercial-use summary

QuestionLlama 4Muse GlimmerMuse Spark
Open source?No (custom licence)Apache 2.0 (plus a separate usage policy )No weights
Commercial useYes, with conditionsYesVia paid API, subject to Meta's terms
Attribution / naming"Built with Llama" and Notice file; derivative model names start with "Llama"Apache attribution normsn/a
Scale limit700M MAU on the release dateNone in Apache 2.0n/a
Can outputs train other models?Yes, with naming ruleNot confirmedNot confirmed
Litigation clauseSuing Meta ends your licenceApache patent termsn/a

None of this is legal advice.

BUILD & EARN scenario

You sell a document-summarization service to small businesses. Assumptions: 1B tokens processed a month, sold at $1.00 per million ($1,000 revenue), all via API.

EngineGeneration costShare of revenueModel-cost gross margin
Llama 4 Scout (hosted)$15015.0%85.0%
Llama 4 Maverick (hosted)$30030.0%70.0%
Muse Glimmer 30B (hosted)$50050.0%50.0%
Muse Spark 1.3 Contributor$12512.5%87.5%, but Meta may train on customer data
Muse Spark 1.3 (standard)$2,000200%−100%

The last row is the lesson: at that price point the closed frontier API loses money on every token, while an open model on a cheap host leaves room for a margin. This is model-cost gross margin only: it excludes engineering, support, compliance, tax and payment fees, and it ignores quality differences between the models. It also assumes Scout or Maverick actually meets your accuracy bar, which should be tested against your own accuracy requirements.

FAQ

Is Llama open source? Downloadable and commercially usable, but under a custom licence with attribution, naming and a 700M-MAU condition. We do not call it open source .

Is Muse Glimmer open source? It is released under Apache 2.0 ; check Meta's separate usage policy.

Does Meta sell an API? Yes: the Meta Model API prices Muse Spark, Muse Image, Muse Voice Transcribe and SAM 3.1 (public preview) . It does not sell hosted Llama inference; that API was retired on 6 Jul 2026 .

What does Llama 4 cost through an API? About $0.10 / $0.30 per million tokens for Scout and $0.20 / $0.60 for Maverick at cheap hosts .

Is self-hosting cheaper? Only at very high sustained volume, on our rental-price assumptions .

Can I use Llama in the EU? Unresolved for Llama 4 multimodal use; get a legal read .