Chinese LLM API Pricing Comparison 2026: The Definitive Buyer's Guide

Chinese LLM API pricing comparison 2026: a practical guide from HeFu.

HeFu · Published 2026-09-02

11 min read

As of Aug 21, 2026 (vendor pricing pages remain the final authority), Chinese LLM APIs are among the cheapest in the world: flagship input prices range from ¥4.00 to ¥12.00 per million tokens (ERNIE 5.1 ¥4.00, GLM-5.1 ¥6.00, Kimi K2.6 ¥6.50, DeepSeek V4 Pro ¥9.00, Qwen3.7 Max ¥12.00), the cheapest budget-tier input is ¥0.20 (Qwen3.5 Flash), and value models like DeepSeek V4 are 80–98% cheaper than GPT-5.5-class peers; however, final selection should weigh cache hit rates, endpoint access, and tool-calling fit rather than nominal list prices. The figures above were verified against official pricing pages by llmabacus on 2026-08-21; Chinese vendors have turned quarterly price cuts into a structural competitive weapon—DeepSeek V4 Flash, for example, offers cache input at ¥0.10 per million tokens, just 1/30th of its standard input price.

2026 Chinese LLM API Pricing Landscape: An Overview

The 2026 Chinese LLM market is shaped by three forces: hardware cost deflation, escalating price wars among domestic vendors, and the rise of aggregator endpoints that arbitrage price gaps. As of Aug 2026, tracking firm pricepertoken lists 610+ models globally, 43 of them free, with paid input prices ranging from roughly $0 to $150 per million tokens; Chinese vendors sit in the lowest price band, and many update prices quarterly—as Morph noted in its 2026-06-28 analysis: "LLM prices change every quarter." Final prices are subject to each vendor's official pricing page: DeepSeek, Alibaba Cloud Bailian/Qwen, Moonshot/Kimi, Zhipu GLM, Baidu ERNIE, Tencent Hunyuan.

The main camps remain unchanged: DeepSeek and Alibaba's Qwen family dominate the extreme value tier; Kimi (Moonshot) differentiates on ultra-long context; GLM (Zhipu), Doubao, and Tencent Hunyuan serve the domestic enterprise market; OpenAI, Claude, and Gemini hold the high-end capability tier. Through the HeFu unified gateway, all of the above families are accessible via one API—including GPT-5.6 (Terra/Sol/Luna), Claude Opus 5 and Sonnet 4.6, DeepSeek-V4-Pro/V4-Flash, Kimi K2.5/K2.6/K3, Gemini 3.6 Flash, Qwen3.7-Max, GLM-5.x, and MiniMax M2.5–M3 (the exact model list and versions are subject to HeFu and vendor official pages).

Headline Findings: What the 2026 Market Data Reveals

Three findings matter most. First, flagship Chinese pricing has collapsed. According to llmabacus, which verified official pricing pages on 2026-08-21, per-million-token list prices are: DeepSeek V4 Pro ¥9.00 input / ¥27.00 output; Qwen3.7 Max ¥12.00 / ¥36.00; Kimi K2.6 ¥6.50 / ¥27.00; GLM-5.1 ¥6.00 / ¥24.00; Baidu ERNIE 5.1 ¥4.00 / ¥18.00. Second, the budget tier is now priced in "cents": the lowest verified input price is Qwen3.5 Flash at ¥0.20 ($0.030) per million tokens, the lowest output price is iFlytek Spark Ultra at ¥0.80, and the lowest cache-input price is DeepSeek V4 Flash at ¥0.10 ($0.015) per million tokens—1/30th of its standard input price. Third, the gap with Western flagships is roughly 40×: IntuitionLabs (as of Feb 2026) calculated that processing 1M input + 1M output tokens costs about $0.70 with DeepSeek (cache miss), versus $5 + $25 = $30 with Claude Opus 4.6. List-price arithmetic confirms that, as of Aug 2026, the DeepSeek V4 series is 80–98% cheaper than GPT-5.5-class models.

Methodology: How This Pricing Comparison Was Conducted

This comparison uses official list prices verified on 2026-08-21 (via llmabacus, cross-checked against vendor pricing pages and public API docs) and is cross-referenced with aggregator endpoint records from Morph (2026-06-28) and the pricepertoken model database. Evaluation dimensions include: per-million-token input/output/cache-input prices (CNY and USD); context window (DeepSeek V4 Flash supports 1M tokens, per DeepSeek official docs); benchmark-adjusted value using SWE-bench Verified as a coding proxy; and channel differences between first-party endpoints and aggregators like OpenRouter, Requesty, and Eden AI. Latency is assessed separately because it varies by endpoint, region, and load; first-party endpoints usually deliver more predictable latency, and the HeFu Hong Kong node provides direct low-latency access to OpenAI models without requiring an overseas credit card (subject to the official HeFu page). USD conversions retain the original rounding from each source, so minor discrepancies may exist.

Side-by-Side Pricing Table (As of Aug 2026)

Model (official list price, per 1M tokens)InputOutputCache inputContextVerified
DeepSeek V4 Pro (official)¥9.00¥27.00—*Official docs2026-08-21
DeepSeek V4 Flash (official)¥3.00 ($0.45)¥9.00 ($1.34)¥0.10 ($0.015)1M2026-08-21
Qwen3.7 Max (official)¥12.00¥36.00—*Official docs2026-08-21
Qwen3.5 Flash (official)¥0.20 ($0.030)—*—*Official docs2026-08-21
Kimi K2.6 (official)¥6.50¥27.00—*Official docs2026-08-21
GLM-5.1 (official)¥6.00¥24.00—*Official docs2026-08-21
ERNIE 5.1 (market benchmark, official)¥4.00¥18.00—*Official docs2026-08-21
MiniMax M3 (US list price, via Z.AI)$0.60$2.40—*Official docs2026-06-28
GLM-5.2 (US list price, via Z.AI)$1.40$4.40—*Official docs2026-06-28

*Not disclosed by vendor; check official pricing pages. Official CNY prices from llmabacus, verified 2026-08-21; USD conversions from the same source. MiniMax M3 and GLM-5.2 prices from Morph, recorded 2026-06-28. Free tiers vary by vendor; most offer limited trial quotas, and aggregator endpoints often run promotional credits. Bold models are available through the HeFu unified gateway; ERNIE 5.1, MiniMax M3, and GLM-5.2 are market benchmarks only.

Performance-to-Price Ratio: Which Models Deliver Best Value

Token prices without a capability anchor are of limited use. As of mid-2026, the most important coding benchmark signal is SWE-bench Verified: Morph (2026-06-28) found that MiniMax M3, at $0.60/$2.40, is the cheapest model with a SWE-bench Verified score above 80%; GLM-5.2 then appeared on Z.AI at $1.40/$4.40, showing that Chinese vendors keep repricing intra-quarter. Within the HeFu catalog, the value picks are (HeFu USD list prices, verified 2026-09-02; subject to the official pricing page): DeepSeek-V4-Pro for general reasoning at $1.87/$3.74 per million tokens; DeepSeek-V4-Flash for high-throughput, high-cache-hit workloads at $0.16/$0.31; Qwen3.5-Plus for input-bursty applications at $0.12/$0.75, where cost becomes nearly negligible; and Kimi K2.6/K3 ($1.01/$4.21 and $3.12/$15.59) for uncompromising ultra-long-context scenarios. Coding teams that need reliable tool calling should review our systematic evaluation Chinese LLM Tool Calling Compatibility—a model with good prices and benchmark scores but poor function-call reliability will cost more in engineering time than it saves in tokens.

Hidden Cost Factors: Concurrency, Storage, and Fine-Tuning Fees

The most important caution in the 2026 data comes from BenchLM (as of Aug 22, 2026): the lowest list price is not the lowest production cost. Output volume, cache hit rate, retry count, and task quality all significantly change the final bill. Five factors deserve attention. Cache input: DeepSeek V4 Flash charges ¥0.10 per million cached tokens versus ¥3.00 per million standard input—a 30× spread that strongly favors long, stable system prompts and RAG contexts. Concurrency: RPM/TPM rate limits differ by tier, and exceeding them triggers retries that amplify costs; DeepSeek's official batch API costs about half the standard price (official docs). Fine-tuning: most Chinese vendors bill per training token plus a hosting surcharge, often quoted separately from inference. Data egress: some vendors charge egress/migration fees for moving fine-tuned weights or logs. Peak-hour surcharges exist on several platforms, but none was prominently disclosed in the 2026 headline price lists—confirm with your account manager and rely on the contract.

Endpoint choice is another hidden variable. pricepertoken notes that the same model can be cheaper on alternative endpoints; Morph recorded DeepSeek V4 Flash at $0.14/$0.28 via aggregator endpoints on 2026-06-28, versus roughly ¥3.00/¥9.00 official—a spread of more than 3×. But aggregator discounts can disappear at any time, and some aggregators actually add a markup. Model both official and aggregator pricing with your own usage profile (cache hit rate, output share, retry rate) before making any commitment.

Regional and International Pricing Differences

Chinese domestic list prices are quoted in CNY and typically assume mainland infrastructure; overseas access is usually billed in USD with a regional premium. IntuitionLabs (as of Feb 2026) measured the cross-border gap: processing 1M input + 1M output tokens costs about $0.70 with DeepSeek versus $5 + $25 = $30 with Claude Opus 4.6—a roughly 40× gap. For international teams, Chinese models remain competitive even after currency conversion and compliance costs. MiniMax M3 ($0.60/$2.40) and GLM-5.2 ($1.40/$4.40) on Z.AI are examples of USD-denominated overseas list prices. On the HeFu platform, models connect directly through the Hong Kong node, reducing cross-border latency—and since September 2026 the entire HeFu catalog, Chinese models included, is billed in USD from a single price list (e.g., DeepSeek-V4-Flash at $0.16/$0.31 per million tokens, verified 2026-09-02), so overseas teams no longer need a CNY balance or a mainland payment rail to reach Chinese-vendor pricing. Most first-party Chinese vendors still publish separate domestic and international price lists—always confirm your account region before comparing numbers.

How to Switch or Migrate: Vendor Lock-in Risks

Migration costs in 2026 are lower than many buyers fear, because the OpenAI-compatible chat-completions format has become a de facto standard among Chinese vendors. Most models—DeepSeek V4, Qwen3.x, GLM-5.x, and others—accept OpenAI-style requests (DeepSeek API, Alibaba Cloud Bailian OpenAI-compatible interface, Moonshot Chat API) with only minor configuration tweaks. The real lock-in risks lie elsewhere: fine-tuned weights are tied to a specific vendor's training pipeline; prompts are often optimized around a model's particular tool-calling habits; and cached-prompt infrastructure only pays off if you stay with the original vendor. Our Chinese LLM Tool Calling Compatibility report (as of Aug 2026) shows that function-call reliability varies significantly across vendors, which directly affects migration costs for applications that depend on structured tool use. A multi-vendor orchestration strategy—DeepSeek V4 Flash for high-throughput, high-cache traffic; Claude Opus 5 or GPT-5.6 for high-end reasoning; Kimi for long-context tasks—reduces dependence on any single vendor's pricing decisions. The HeFu unified gateway is designed exactly for this, letting you switch models without rewriting application code. For background on the high end, see Claude Opus 5 API Access.

Forecast: Price Trends for Late 2026 and 2027

The price direction is clear. Morph documents that Chinese LLM pricing changes quarterly; GLM-5.2's USD list price on Z.AI appeared shortly after GLM-5.1's domestic CNY price, showing that vendors adjust regional pricing within a quarter. Note that cross-model, cross-currency comparisons cannot automatically support a "price cut" conclusion—track the same model's regional history to see the real trend, and defer to official pages. As of Aug 2026, there is no evidence of a floor: hardware costs are still falling, and competition among DeepSeek, Qwen, Kimi, GLM, and Doubao shows no sign of easing. Looking toward late 2026 and 2027, the battleground will shift from input list prices to cache-input prices, batch discounts, and fine-tuning economics—the cost levers that actually determine production bills. Buyers should put quarterly price-review clauses in contracts rather than locking annual prices, and should be wary of vendors that refuse to negotiate on cache pricing.

FAQ

What is the cheapest Chinese LLM API in 2026?

By scenario, as of Aug 2026: on the input side, Qwen3.5 Flash at ¥0.20 per million tokens; on the output side, iFlytek Spark Ultra at ¥0.80 per million tokens; on the cache-input side, DeepSeek V4 Flash at ¥0.10 per million tokens (all verified by llmabacus on 2026-08-21). On a performance-adjusted basis, MiniMax M3 at $0.60/$2.40 is the cheapest model above 80% on SWE-bench Verified (Morph, 2026-06-28). Within the HeFu catalog, start with DeepSeek-V4-Flash for cache-heavy workloads and Qwen3.5 for input-bursty tasks; final prices are subject to official pricing pages.

Why do OpenRouter, Requesty, and Eden AI show different prices for the same model?

Aggregators use different upstreams, apply different markups or subsidies, and refresh prices on different schedules, so the same model can be cheaper or more expensive than the official price (pricepertoken). Morph recorded DeepSeek V4 Flash at $0.14/$0.28 via aggregator endpoints on 2026-06-28, versus roughly ¥3.00/¥9.00 official—a spread of more than 3×. The key point is that the lowest list price is not the lowest production cost: cache hit rate, output volume, and retries all change the final bill (BenchLM, as of Aug 22, 2026). Model both channels with your own usage curve before deciding.

How much cheaper are Chinese models than OpenAI or Claude, and is migration worth it?

At the raw token level, as of Aug 2026, the DeepSeek V4 series is 80–98% cheaper than GPT-5.5-class models (derived from official list prices), and for a 1M input + 1M output workload it is roughly 40× cheaper than Claude Opus 4.6 (IntuitionLabs, as of Feb 2026).

FAQ

What is the cheapest Chinese LLM API in 2026?

By scenario, as of Aug 2026: on the input side, Qwen3.5 Flash at ¥0.20 per million tokens; on the output side, iFlytek Spark Ultra at ¥0.80 per million tokens; on the cache-input side, DeepSeek V4 Flash at ¥0.10 per million tokens (all verified by [llmabacus](https://llmabacus.com) on 2026-08-21). On a performance-adjusted basis, MiniMax M3 at $0.60/$2.40 is the cheapest model above 80% on SWE-bench Verified ([Morph](https://morph.so), 2026-06-28). Within the HeFu catalog, start with DeepSeek-V4-Flash for cache-heavy workloads and Qwen3.5 for input-bursty tasks; final prices are subject to official pricing pages.

Why do OpenRouter, Requesty, and Eden AI show different prices for the same model?

Aggregators use different upstreams, apply different markups or subsidies, and refresh prices on different schedules, so the same model can be cheaper or more expensive than the official price ([pricepertoken](https://pricepertoken.com)). [Morph](https://morph.so) recorded DeepSeek V4 Flash at $0.14/$0.28 via aggregator endpoints on 2026-06-28, versus roughly ¥3.00/¥9.00 official—a spread of more than 3×. The key point is that the lowest list price is not the lowest production cost: cache hit rate, output volume, and retries all change the final bill ([BenchLM](https://benchlm.com), as of Aug 22, 2026). Model both channels with your own usage curve before deciding.

How much cheaper are Chinese models than OpenAI or Claude, and is migration worth it?

At the raw token level, as of Aug 2026, the DeepSeek V4 series is 80–98% cheaper than GPT-5.5-class models (derived from official list prices), and for a 1M input + 1M output workload it is roughly 40× cheaper than Claude Opus 4.6 ([IntuitionLabs](https://www.intuitionlabs.ai), as of Feb 2026).

Related reading

Best Pay-as-You-Go LLM API for Indie Developers in 2026

Best pay-as-you-go LLM API for indie developers: a practical guide from HeFu.

How Can a Startup Reduce LLM API Costs: A 2026 Playbook

How can a startup reduce LLM API costs: a practical guide from HeFu.

Cheap DeepSeek API Pay As You Go: Complete Cost Guide (As of Aug 2026)

cheap DeepSeek API pay as you go: a practical guide from HeFu.

Want to try these models yourself?

HeFu aggregates every major LLM behind one OpenAI-compatible API — pay as you go.

Prices quoted are official list prices for reference — see the main site pricing page for actual rates.

Start Free TrialBook an Enterprise Demo