NotCheapMAKE EVERY CREDIT COUNT

Z.ai / GLM Coding Plans 2026: Real Cost vs Claude Code, Codex and Cursor

Direct answer

Z.ai's flagship model, GLM-5.3, costs $1.40 per million input tokens and $4.40 per million output tokens, with $0.26/M cached input, according to Z.ai's API pricing. That's roughly a fifth of Claude Opus 5's output rate. Separately, Z.ai sells a flat-fee GLM Coding Plan (Lite/Pro/Max) restricted to officially supported coding tools like Claude Code and Cline. For anyone coding daily at real volume, the comparison below shows when the subscription can beat metered API billing — but Z.ai does not publish a numerical usage allowance for the plan, an important limit when estimating value.

Consumer / subscription pricing

Z.ai's API pricing page lists the current rate card below. Z.ai also runs a free consumer chat interface at chat.z.ai, separate from both the API and the Coding Plan.

Independent price trackers list three GLM Coding Plan tiers — Lite $18/mo, Pro $72/mo, Max $160/mo — and billing-cycle discounts of roughly 10% monthly, 20% quarterly and 30% annually, which would bring the effective monthly prices to $12.60/$50.40/$112 on annual billing. The listed model lineup includes GLM-5.3 and GLM-5.3-Flash, with tier differences mainly in usage quota. Check Z.ai's checkout for current prices, discounts and included features before subscribing.

Z.ai's public pricing materials do not state the Lite tier's usage allowance as a number. Pro and Max are described as multiples of that undisclosed base (5x and 20x). Without a published base allowance, there is no reliable way to convert the plan into a token estimate.

GLM Coding Plan pricing has changed materially over the past year. Older prices may refer to a previous plan or promotion, so compare the date, tier and billing period before relying on a quote.

API pricing — full rate card

Z.ai lists the following API prices per 1M tokens:

ModelInputCached inputOutput
GLM-5.3$1.40$0.26$4.40
GLM-5.2$1.40$0.26$4.40
GLM-5.3-Flash$0.15$0.03$0.50
GLM-5.3-FlashX$0.37$0.075$1.25
GLM-5.1$1.40$0.26$4.40
GLM-5$1.00$0.20$3.20
GLM-4.7 / GLM-4.6 / GLM-4.5$0.60$0.11$2.20
GLM-4.5-Air$0.20$0.03$1.10
GLM-4.7-Flash / GLM-4.5-FlashFreeFreeFree

The same pricing page lists Web Search at $0.01/use. It also publishes rates for vision (GLM-4.6V), image generation (GLM-Image $0.015/image, CogView-4 $0.01/image), video (CogVideoX-3, $0.2/video), and audio (GLM-ASR-2512, $0.03/MTok).

Artificial Analysis reports that GLM-5.3, despite having the same per-token price as GLM-5.2, costs more per completed task in its benchmark ($0.44 → $0.68) because it generates longer responses. A flat per-token rate does not guarantee a flat per-task cost.

Real workload calculations

Using GLM-5.3's listed rate with no caching applied:

WorkloadFormulaCost
A — Everyday (1M in, 250K out)(1×$1.40) + (0.25×$4.40)$2.50
B — Coding/reasoning (10M in, 3M out)(10×$1.40) + (3×$4.40)$27.20
C — High volume (400M in, 100M out/mo)(400×$1.40) + (100×$4.40)$1,000.00

D — Provider-specific: Coding Plan vs. metered API. For this modeled scenario, assume a solo developer running heavy agentic coding sessions generates roughly 5M input and 1M output tokens per day (full-repo context resent on many turns); this is an illustrative assumption, not a benchmark. Over 22 working days: 110M input + 22M output tokens.

  • Metered API cost: (110×$1.40) + (22×$4.40) = $154.00 + $96.80 = $250.80/month
  • GLM Coding Plan Max: $160/month list, or $144/month with the standard monthly discount — cheaper than the metered API under this assumption, which matches the general pattern independent reviewers report: steady, heavy daily coding favors the subscription; light or bursty use favors the metered API.

Comparison with Claude Code, Codex and Cursor

Claude Code and Codex can be paid through subscriptions or usage-based billing, while Cursor combines a lower entry price with usage-based overages. These plans are not directly comparable: included usage and billing rules differ, and the price comparison does not imply feature parity or equivalent coding quality. Verify each provider's current checkout price and usage terms before choosing a plan.

Regional pricing

Z.ai's public USD rate card does not list a separate China-only GLM API rate. Regional availability or account-specific pricing may differ, so users operating a China-region account should confirm the applicable rate in their account before budgeting.

BUILD & EARN scenario

MODELING ASSUMPTION, clearly labeled, not a market benchmark: a small coding agency offers "AI-assisted PR review and small-feature delivery" to clients, built on the GLM Coding Plan Max ($160/mo fixed cost) for a team of 3 developers sharing the pool. Assumed client price: $500/month retainer per client (illustrative only). At 4 clients: $2,000/month revenue against $160/month in model cost — a wide margin, but this model deliberately ignores developer labor cost, which dominates a real agency's P&L; the point is only that the AI-tooling cost itself is a rounding error next to a $2,000 revenue line, not a claim that this business is profitable as described.

FAQ

Is GLM-5.3 available on the pay-as-you-go API, or only in the Coding Plan? Both are available on the pay-as-you-go API. Z.ai lists GLM-5.3 at $1.40/$4.40 per million tokens, the same rate as GLM-5.2.

Why doesn't Z.ai say how many tokens the Coding Plan actually includes? Z.ai does not publish a numerical Lite allowance in its public Coding Plan information. Pro and Max are marketed only as multiples ("5x," "20x") of that undisclosed baseline. Budget with real usage testing, not the marketing multiplier.

Is the GLM Coding Plan cheaper than Claude Code or Cursor? At the subscription level, GLM Coding Plan Max's listed price is lower than some premium tiers from Claude and ChatGPT. Prices and included usage vary by plan and can change; this is a price comparison, not a claim of feature parity or equivalent coding quality.