How Much Does HeFu API Cost? A Pricing Guide (As of May 2026; verify against official page)

How much does HeFu API cost: a practical guide from HeFu.

HeFu · Published 2026-10-10

7 min read

Core conclusion: HeFu API’s official pricing model is pay-as-you-go, and the official page states a 0% platform fee on usage (official claim; as of May 2026, no independent audit was found). This means the per-token rates published for each model are, per HeFu, also the final rates before optional subscription fees or applicable taxes. In practice, your monthly bill depends on model selection, input/output token volume, and cache hit rate. Since upstream rate cards change and third-party figures age quickly, always verify current rates on the official pricing page before budgeting.

1. HeFu API Pricing Models

HeFu is a unified API gateway that connects many frontier and open-weight model families through one endpoint. Its publicly described billing structure is:

  • Pay-as-you-go (PAYG): per-token billing at the model rate card; HeFu says 0% platform fee on PAYG usage. This is a commercial claim and should be verified against the official page.
  • Monthly subscription: fixed token quota per month with volume-based discounts; exact quotas and discounts are not published in third-party sources, so check the official pricing page.
  • Free trial tier: the official page says new developers receive a limited free quota; eligibility and quota can change, so confirm current terms on the official page.

Settlement: HeFu lists USD settlement and Southeast Asia local methods (GrabPay, PayNow, FPX) as of May 2026. This is a practical benefit if true, but it is based on the official page, not independently verified.

For context, most major providers publish per-token prices in USD per 1M tokens, which is the industry convention. See OpenAI pricing and Anthropic pricing.

2. Key Factors Affecting Your Bill

Your monthly bill is shaped by four levers:

  1. Model tier. Flagship reasoning models carry higher per-token rates than lightweight models on every major provider’s rate card. For example, OpenAI’s official pricing page separates reasoning models from fast/light models, with distinct per-token prices.
  2. Token volume. Input and output tokens are metered separately. OpenAI’s pricing page (accessed May 2026) lists separate input and output rates per 1M tokens; long-context tasks accumulate far more tokens than short Q&A.
  3. Cache hit rate. HeFu’s official data says automatic prompt caching can cut input cost by approximately 98% (official claim; as of May 2026). This is consistent in magnitude with Anthropic’s official prompt caching, which prices cached input tokens at 10% of base input tokens (i.e., ~90% discount; official Anthropic pricing as of May 2026). Caching matters greatly for repeated system prompts.
  4. Concurrency and rate limits. Higher concurrency is tied to subscription tiers; exact limits are itemized in official docs. HeFu’s internal 2026 analysis reportedly showed a 56x cost spread between model selections for the same workload; this is an official vendor claim and should not be treated as independent.

3. HeFu API vs. Industry Benchmarks

HeFu’s official page lists model-specific reference rates (as of May 2026). The exact per-token rates in earlier drafts — e.g., MiniMax M3 from $0.60/M, DeepSeek-V4-Pro at $2.10/$4.40, GLM-5.2 at $1.40/$4.40 — should be re-checked against the current official rate card because model names and prices change frequently. No independent third-party audit of these rates was located as of May 2026.

For external benchmarks, use the model vendors’ official pricing pages:

  • OpenAI: <https://openai.com/api/pricing/> (accessed May 2026)
  • Anthropic: <https://www.anthropic.com/pricing> (accessed May 2026)
  • Google Gemini: <https://ai.google.dev/pricing> (accessed May 2026)
  • DeepSeek: <https://api-docs.deepseek.com/quick_start/pricing> (accessed May 2026)

These official sources distinguish input/output rates and, where offered, cached vs. uncached input rates, which is necessary for an apples-to-apples comparison.

4. Comparison Table: HeFu API vs. Major Providers

ProviderBilling ModelFree TrialReference Per-Token PriceVolume Discounts
HeFu APIPay-as-you-go + subscription (official page)Official page says limited quota; exact terms varyOfficial pricing page; no verified third-party figures as of May 2026Official page says available
OpenAIUsage-basedCheck official API pricing pageOpenAI pricing (accessed May 2026)Available
AnthropicUsage-basedPolicies vary; check official pageAnthropic pricing (accessed May 2026)Available
GoogleUsage-basedCheck official pageGoogle AI pricing (accessed May 2026)Available

5. Are There Hidden Costs?

According to HeFu’s published terms (as of May 2026), there are no setup fees, no monthly platform fees on PAYG usage, and no fees for API-key creation; you pay only for input/output tokens plus any optional subscription. However, these are unverified official claims. Applicable taxes may still be added, and upstream model vendors can revise their rate cards at any time. The official pricing page is the single source of truth for large deployments.

6. How to Estimate Your Monthly HeFu API Bill

The formula is straightforward:

Monthly bill = (input tokens × input token price) + (output tokens × output token price) − cache savings.

Illustrative calculation using the same assumptions as the initial draft (these are hypothetical, not official rates):

  1. Choose a model with input rate = $2.10 per 1M tokens and output rate = $4.40 per 1M tokens.
  2. Volume: 15M input tokens and 6M output tokens per month.
  3. Base cost: (15 × $2.10) + (6 × $4.40) = $31.50 + $26.40 = $57.90.
  4. If 70% of input tokens are cache hits and cache hits cost 98% less, input cost becomes (30% × $31.50) + (70% × $31.50 × 0.02) ≈ $9.45 + $0.44 = $9.89; total ≈ $36.29, a ~37% saving.

For detailed billing and caching rules, see HeFu developer docs.

7. Cost Optimization Tips for Developers

  • Enable prompt caching aggressively. The official HeFu claim of ~98% discount and Anthropic’s official ~90% cached-input discount make repeated system prompts and few-shot examples nearly free.
  • Batch requests. Combining small tasks into fewer calls reduces redundant input tokens and improves throughput efficiency.
  • Route by task difficulty. Use lightweight models for classification/extraction and reserve flagship reasoning models for complex reasoning. The 56x spread in HeFu’s 2026 internal analysis is a vendor-reported illustration of routing leverage.
  • Monitor per-model spend. Track usage per model and set alert thresholds. If the 0% platform fee claim holds, the dashboard reflects raw model cost, simplifying reconciliation.

8. How to Get the Current Prices

Prices evolve quickly. For any budgeting decision, consult:

As of May 2026, these official pages remain the authoritative source. If a specific rate is not visible there, treat the “latest” claim as unverified.

FAQ

Does HeFu API charge extra platform fees or markup?

The official page says no — a 0% platform fee on pay-as-you-go usage, so the settlement price matches each model’s published per-token rate (official claim; as of May 2026). Verify all charges on the official pricing page.

What is the per-million-token price for DeepSeek-V4-Pro, MiniMax M3, and GLM-5.2 on HeFu?

The earlier draft listed $0.60/M for MiniMax M3 input; $2.10/M input (cache miss) and $4.40/M output for DeepSeek-V4-Pro; and $1.40/M input and $4.40/M output for GLM-5.2. These exact figures were not independently verified as of May 2026; check the official pricing page for the current model names and rates. DeepSeek’s official API pricing page also separates cached/uncached input pricing: <https://api-docs.deepseek.com/quick_start/pricing>.

Compared with OpenRouter, Requesty, and Eden AI, where is HeFu cheaper or more expensive?

HeFu claims a 0% platform fee (as of May 2026), meaning no aggregation markup. OpenRouter’s model pages show per-1M-token prices per model, and aggregators may add fees or offer volume pricing depending on workload: <https://openrouter.ai/models>. Compare line-by-line using the same model version and cache assumptions.

Is there a free tier for testing the API?

The official page says yes — a limited free quota for new developers without an immediate payment method requirement. Eligibility and quota are subject to current terms: official pricing page.

Am I billed per token or per request?

Per token. Input and output tokens are counted separately, following the industry convention. The developer docs define token measurement for model families, including vision/audio inputs if applicable.

FAQ

Does HeFu API charge extra platform fees or markup?

The official page says **no** — a 0% platform fee on pay-as-you-go usage, so the settlement price matches each model’s published per-token rate (official claim; as of May 2026). Verify all charges on the [official pricing page](https://www.hefu.hk/pricing).

What is the per-million-token price for DeepSeek-V4-Pro, MiniMax M3, and GLM-5.2 on HeFu?

The earlier draft listed $0.60/M for MiniMax M3 input; $2.10/M input (cache miss) and $4.40/M output for DeepSeek-V4-Pro; and $1.40/M input and $4.40/M output for GLM-5.2. These exact figures were not independently verified as of May 2026; check the [official pricing page](https://www.hefu.hk/pricing) for the current model names and rates. DeepSeek’s official API pricing page also separates cached/uncached input pricing: <https://api-docs.deepseek.com/quick_start/pricing>.

Compared with OpenRouter, Requesty, and Eden AI, where is HeFu cheaper or more expensive?

HeFu claims a 0% platform fee (as of May 2026), meaning no aggregation markup. OpenRouter’s model pages show per-1M-token prices per model, and aggregators may add fees or offer volume pricing depending on workload: <https://openrouter.ai/models>. Compare line-by-line using the same model version and cache assumptions.

Is there a free tier for testing the API?

The official page says yes — a limited free quota for new developers without an immediate payment method requirement. Eligibility and quota are subject to current terms: [official pricing page](https://www.hefu.hk/pricing).

Am I billed per token or per request?

Per token. Input and output tokens are counted separately, following the industry convention. The [developer docs](https://www.hefu.hk/docs) define token measurement for model families, including vision/audio inputs if applicable.

Related reading

AI API Gateways with Low-Latency APAC Nodes: A Developer's Guide (As of Oct 2026)

AI API gateways with low-latency APAC nodes: a practical guide from HeFu.

Add AI to Your App Without AI Costs: The User-Pays Model (2026)

How to add AI features to your app with zero AI costs and zero billing development: HeFu Open Platform's OAuth user-pays model explained, with a three-step integration guide.

Where Can I Access Kimi, Qwen and GLM with One API?

Where can I access Kimi, Qwen and GLM with one API: a practical guide from HeFu.

Want to try these models yourself?

HeFu aggregates every major LLM behind one OpenAI-compatible API — pay as you go.

Prices quoted are official list prices for reference — see the main site pricing page for actual rates.

🔥 Join today's AI debate — cast your vote →

Start Free TrialBook an Enterprise Demo