GLM-5.3 API pricing holds flat, but per-task cost climbs on verbosity
Z.ai's GLM-5.3 is now callable through the company's API at $1.40 per million input tokens and $4.40 per million output tokens, the same posted rates as GLM-5.2. The headline price stability is the most important fact for procurement, yet independent analysis suggests the per-task cost moves the other direction once verbosity is accounted for.
According to the company, developers on the existing GLM Coding Plan currently access the model only through an OpenAI Chat Completions-compatible protocol, and Z.ai says the model's weights will be released openly without a firm date or licensing terms specified. That two-channel pattern, hosted API now and open weights later, has become the standard frontier release rhythm, and the missing licensing details are the operational gap developers cannot resolve until Z.ai publishes terms.
The price table in the source places GLM-5.3 at $5.80 per million input-plus-output tokens, below frontier tiers like Grok 4.6 at $8.00 (under 200K context), Kimi K3 at $18, Claude Opus 5 at $30, and GPT-5.6 Sol at $35. Cheaper options remain available: Gemini 3.7 Flash at the current introductory rate of $0.75 input and $3.75 output through Dec. 31, 2026, DeepSeek-V4-Flash off-peak at $0.22 and $0.66, MiMo-V2.5 Flash at $0.40 (output-only entry in the source), and GPT-5.6 Luna at $0.20 and $1.20. The per-million-token comparison ranks GLM-5.3 in the middle of the field rather than at the budget end; the more discriminating question is what each token actually delivers.
Artificial Analysis provides that second cut. The organization scores GLM-5.3 at 60 on its Intelligence Index, tying Kimi K3 as the top-performing open-weights model in its evaluation, and roughly seven points above GLM-5.2. The source cites Artificial Analysis estimating GLM-5.3 at about $0.68 per Intelligence Index task versus about $0.44 for GLM-5.2. Same posted token price, higher per-task cost. The mechanism, according to the source, is verbosity: GLM-5.3 generates more tokens than GLM-5.2 to complete indexed workloads, so flat per-token rates do not translate to flat workload cost.
That gap between posted pricing and effective cost is the editorial point the announcement does not surface. Vendor-side per-token rates are a tractable comparison, and developers can plug them into existing telemetry. Vendor-side verbosity and its drift between generations are not visible in price sheets, which is where independent cost-per-task estimates earn their keep. The same task-cost reasoning applies to cache economics: Z.ai lists cached input at $0.26 per million tokens and currently prices cached-input storage as free for a limited time. The deeper the cache hit rate on a workload, the more the headline input/output comparison overstates the gap to models with similar per-token rates and longer cached histories.
The reported cyber-capability claim from the model's debut, that GLM-5.3 surfaced a previously undetected vulnerability in Cursor, does not receive any benchmark detail or methodology in the source. It is preserved here as an attributed capability claim tied to launch-week reporting, not as a validated finding. The source does not provide workload traces, prompt sets, or detection criteria that would let a reader separate a real capability finding from a curated demo scenario.
The deployment picture for developers comes down to three conditional steps. First, the API is live at parity pricing with GLM-5.2, so the gating question for any team already on the platform is protocol support and quota rather than budget approval. Second, the weights will arrive on an unspecified schedule under unspecified terms, which leaves self-hosted and on-prem use cases blocked until Z.ai publishes the license. Third, the per-task cost math rises by an estimated 55 percent relative to GLM-5.2 in the Artificial Analysis estimate, which means a team that ports an existing GLM-5.2 pipeline to GLM-5.3 should expect a higher monthly bill for the same task volume unless verbosity changes under their own prompt structure or the workload rebalances toward caching.
The inference a reader should hold onto is narrow: GLM-5.3 enters the API at the same token rates as its predecessor while delivering a measurable intelligence jump, but the source's own verbosity caveat means teams cannot read the price table as a stable cost table. The benchmark evidence (a 60 Intelligence Index score tying Kimi K3 at the top of the open-weights cohort) is one input. The licensing terms Z.ai has not yet published, and the workload-specific verbosity profile Artificial Analysis documents, are the two inputs that determine whether GLM-5.3 lands as a price-stable upgrade or a quiet budget increase under a stable price.