How Much Does HeFu API Cost? A Pricing Guide (As of May 2026; verify against official page)
How much does HeFu API cost: a practical guide from HeFu.
HeFu · Published 2026-10-107 min read
Core conclusion: HeFu API’s official pricing model is pay-as-you-go, and the official page states a 0% platform fee on usage (official claim; as of May 2026, no independent audit was found). This means the per-token rates published for each model are, per HeFu, also the final rates before optional subscription fees or applicable taxes. In practice, your monthly bill depends on model selection, input/output token volume, and cache hit rate. Since upstream rate cards change and third-party figures age quickly, always verify current rates on the official pricing page before budgeting.
1. HeFu API Pricing Models
HeFu is a unified API gateway that connects many frontier and open-weight model families through one endpoint. Its publicly described billing structure is:
- Pay-as-you-go (PAYG): per-token billing at the model rate card; HeFu says 0% platform fee on PAYG usage. This is a commercial claim and should be verified against the official page.
- Monthly subscription: fixed token quota per month with volume-based discounts; exact quotas and discounts are not published in third-party sources, so check the official pricing page.
- Free trial tier: the official page says new developers receive a limited free quota; eligibility and quota can change, so confirm current terms on the official page.
Settlement: HeFu lists USD settlement and Southeast Asia local methods (GrabPay, PayNow, FPX) as of May 2026. This is a practical benefit if true, but it is based on the official page, not independently verified.
For context, most major providers publish per-token prices in USD per 1M tokens, which is the industry convention. See OpenAI pricing and Anthropic pricing.
2. Key Factors Affecting Your Bill
Your monthly bill is shaped by four levers:
- Model tier. Flagship reasoning models carry higher per-token rates than lightweight models on every major provider’s rate card. For example, OpenAI’s official pricing page separates reasoning models from fast/light models, with distinct per-token prices.
- Token volume. Input and output tokens are metered separately. OpenAI’s pricing page (accessed May 2026) lists separate input and output rates per 1M tokens; long-context tasks accumulate far more tokens than short Q&A.
- Cache hit rate. HeFu’s official data says automatic prompt caching can cut input cost by approximately 98% (official claim; as of May 2026). This is consistent in magnitude with Anthropic’s official prompt caching, which prices cached input tokens at 10% of base input tokens (i.e., ~90% discount; official Anthropic pricing as of May 2026). Caching matters greatly for repeated system prompts.
- Concurrency and rate limits. Higher concurrency is tied to subscription tiers; exact limits are itemized in official docs. HeFu’s internal 2026 analysis reportedly showed a 56x cost spread between model selections for the same workload; this is an official vendor claim and should not be treated as independent.
3. HeFu API vs. Industry Benchmarks
HeFu’s official page lists model-specific reference rates (as of May 2026). The exact per-token rates in earlier drafts — e.g., MiniMax M3 from $0.60/M, DeepSeek-V4-Pro at $2.10/$4.40, GLM-5.2 at $1.40/$4.40 — should be re-checked against the current official rate card because model names and prices change frequently. No independent third-party audit of these rates was located as of May 2026.
For external benchmarks, use the model vendors’ official pricing pages:
- OpenAI: <https://openai.com/api/pricing/> (accessed May 2026)
- Anthropic: <https://www.anthropic.com/pricing> (accessed May 2026)
- Google Gemini: <https://ai.google.dev/pricing> (accessed May 2026)
- DeepSeek: <https://api-docs.deepseek.com/quick_start/pricing> (accessed May 2026)
These official sources distinguish input/output rates and, where offered, cached vs. uncached input rates, which is necessary for an apples-to-apples comparison.
4. Comparison Table: HeFu API vs. Major Providers
| Provider | Billing Model | Free Trial | Reference Per-Token Price | Volume Discounts |
|---|---|---|---|---|
| HeFu API | Pay-as-you-go + subscription (official page) | Official page says limited quota; exact terms vary | Official pricing page; no verified third-party figures as of May 2026 | Official page says available |
| OpenAI | Usage-based | Check official API pricing page | OpenAI pricing (accessed May 2026) | Available |
| Anthropic | Usage-based | Policies vary; check official page | Anthropic pricing (accessed May 2026) | Available |
| Usage-based | Check official page | Google AI pricing (accessed May 2026) | Available |
5. Are There Hidden Costs?
According to HeFu’s published terms (as of May 2026), there are no setup fees, no monthly platform fees on PAYG usage, and no fees for API-key creation; you pay only for input/output tokens plus any optional subscription. However, these are unverified official claims. Applicable taxes may still be added, and upstream model vendors can revise their rate cards at any time. The official pricing page is the single source of truth for large deployments.
6. How to Estimate Your Monthly HeFu API Bill
The formula is straightforward:
Monthly bill = (input tokens × input token price) + (output tokens × output token price) − cache savings.
Illustrative calculation using the same assumptions as the initial draft (these are hypothetical, not official rates):
- Choose a model with input rate = $2.10 per 1M tokens and output rate = $4.40 per 1M tokens.
- Volume: 15M input tokens and 6M output tokens per month.
- Base cost: (15 × $2.10) + (6 × $4.40) = $31.50 + $26.40 = $57.90.
- If 70% of input tokens are cache hits and cache hits cost 98% less, input cost becomes (30% × $31.50) + (70% × $31.50 × 0.02) ≈ $9.45 + $0.44 = $9.89; total ≈ $36.29, a ~37% saving.
For detailed billing and caching rules, see HeFu developer docs.
7. Cost Optimization Tips for Developers
- Enable prompt caching aggressively. The official HeFu claim of ~98% discount and Anthropic’s official ~90% cached-input discount make repeated system prompts and few-shot examples nearly free.
- Batch requests. Combining small tasks into fewer calls reduces redundant input tokens and improves throughput efficiency.
- Route by task difficulty. Use lightweight models for classification/extraction and reserve flagship reasoning models for complex reasoning. The 56x spread in HeFu’s 2026 internal analysis is a vendor-reported illustration of routing leverage.
- Monitor per-model spend. Track usage per model and set alert thresholds. If the 0% platform fee claim holds, the dashboard reflects raw model cost, simplifying reconciliation.
8. How to Get the Current Prices
Prices evolve quickly. For any budgeting decision, consult:
- HeFu official pricing page — current per-token rates, plans, and free-tier terms
- HeFu model catalog — currently available model versions
- HeFu developer docs — billing rules, caching behavior, rate limits
As of May 2026, these official pages remain the authoritative source. If a specific rate is not visible there, treat the “latest” claim as unverified.
FAQ
Does HeFu API charge extra platform fees or markup?
The official page says no — a 0% platform fee on pay-as-you-go usage, so the settlement price matches each model’s published per-token rate (official claim; as of May 2026). Verify all charges on the official pricing page.
What is the per-million-token price for DeepSeek-V4-Pro, MiniMax M3, and GLM-5.2 on HeFu?
The earlier draft listed $0.60/M for MiniMax M3 input; $2.10/M input (cache miss) and $4.40/M output for DeepSeek-V4-Pro; and $1.40/M input and $4.40/M output for GLM-5.2. These exact figures were not independently verified as of May 2026; check the official pricing page for the current model names and rates. DeepSeek’s official API pricing page also separates cached/uncached input pricing: <https://api-docs.deepseek.com/quick_start/pricing>.
Compared with OpenRouter, Requesty, and Eden AI, where is HeFu cheaper or more expensive?
HeFu claims a 0% platform fee (as of May 2026), meaning no aggregation markup. OpenRouter’s model pages show per-1M-token prices per model, and aggregators may add fees or offer volume pricing depending on workload: <https://openrouter.ai/models>. Compare line-by-line using the same model version and cache assumptions.
Is there a free tier for testing the API?
The official page says yes — a limited free quota for new developers without an immediate payment method requirement. Eligibility and quota are subject to current terms: official pricing page.
Am I billed per token or per request?
Per token. Input and output tokens are counted separately, following the industry convention. The developer docs define token measurement for model families, including vision/audio inputs if applicable.
FAQ
Does HeFu API charge extra platform fees or markup?
The official page says **no** — a 0% platform fee on pay-as-you-go usage, so the settlement price matches each model’s published per-token rate (official claim; as of May 2026). Verify all charges on the [official pricing page](https://www.hefu.hk/pricing).
What is the per-million-token price for DeepSeek-V4-Pro, MiniMax M3, and GLM-5.2 on HeFu?
The earlier draft listed $0.60/M for MiniMax M3 input; $2.10/M input (cache miss) and $4.40/M output for DeepSeek-V4-Pro; and $1.40/M input and $4.40/M output for GLM-5.2. These exact figures were not independently verified as of May 2026; check the [official pricing page](https://www.hefu.hk/pricing) for the current model names and rates. DeepSeek’s official API pricing page also separates cached/uncached input pricing: <https://api-docs.deepseek.com/quick_start/pricing>.
Compared with OpenRouter, Requesty, and Eden AI, where is HeFu cheaper or more expensive?
HeFu claims a 0% platform fee (as of May 2026), meaning no aggregation markup. OpenRouter’s model pages show per-1M-token prices per model, and aggregators may add fees or offer volume pricing depending on workload: <https://openrouter.ai/models>. Compare line-by-line using the same model version and cache assumptions.
Is there a free tier for testing the API?
The official page says yes — a limited free quota for new developers without an immediate payment method requirement. Eligibility and quota are subject to current terms: [official pricing page](https://www.hefu.hk/pricing).
Am I billed per token or per request?
Per token. Input and output tokens are counted separately, following the industry convention. The [developer docs](https://www.hefu.hk/docs) define token measurement for model families, including vision/audio inputs if applicable.