Where Can I Access Kimi, Qwen and GLM with One API?
Where can I access Kimi, Qwen and GLM with one API: a practical guide from HeFu.
HeFu · Published 2026-10-078 min read
As of May 2026, the simplest answer is a unified model gateway: instead of maintaining separate accounts, keys, and invoices for each Chinese LLM vendor, point an OpenAI-compatible client at a single base_url. HeFu is one such gateway: it exposes Kimi, Qwen, and GLM (along with GPT, Claude, DeepSeek, Gemini, and more) through the same API key at https://api.hefu.hk/v1; switching models only requires changing the model parameter, not rebuilding the integration. Specific model identifiers and available versions are subject to the HeFu official model catalog.
The Growing Demand for Multi-Model API Access
Teams evaluating models like Kimi K2.6, Qwen 3.8 Max, and GLM 5 typically end up with three dashboards, three keys, and three invoices (model names are illustrative; actual identifiers depend on each gateway's official catalog). This friction has driven the rise of aggregation platforms. As of September 30, 2026, the AIsa model gateway brings six Chinese LLM families — Qwen, DeepSeek, Kimi, GLM, MiniMax, and Hunyuan — under one API key, with unified wallet billing and OpenAI-compatible SDKs (AIsa documentation). In 2026, Tencent Cloud TokenHub launched a similar unified OpenAI-style endpoint: you only change base_url and the model name to move between DeepSeek, GLM, Kimi, MiniMax, and Qwen (Tencent Cloud official docs). Yotta AI Gateway also claimed in 2026 that one OpenAI-compatible key can access 50+ models, including GLM 5.2 and DeepSeek V3.2 (Yotta blog).
This convergence has been discussed before: our DeepSeek, Qwen, Kimi single API key guide covers the architecture, while how to use DeepSeek, Kimi, and Qwen in one app demonstrates the client-side pattern.
HeFu API: A Single Endpoint for All Your Models
HeFu applies the aggregation idea directly to the three model families discussed in this article. One account (https://www.hefu.hk) and one key are enough to call:
- Kimi (Moonshot AI): Kimi K2.5 / K2.6 / K3 and other versions — long context is its main selling point, suitable for Chinese long-document processing. Actual version identifiers are subject to the official catalog.
- Qwen (Alibaba): Qwen 3.8 Max, Qwen 3.5–3.7 series, and the Qwen 3 open-source line (8B–235B Thinking / Instruct / VL / Coder variants). Specific identifiers are subject to the official catalog.
- GLM (Zhipu AI): GLM 5 and GLM-5.x — competitive in coding and agentic tool-calling scenarios. Specific identifiers are subject to the official catalog.
All requests share https://api.hefu.hk/v1 and an OpenAI-compatible request/response structure. Under the same account you can also route Claude, GPT, DeepSeek, Gemini, Wan, Seed, Grok, and MiniMax, so adding non-Chinese models later requires no extra integration. The Hong Kong node is designed for direct connections outside mainland China, and registration does not require an overseas credit card (official site).
Quick Start: From Registration to First Call
Going from zero to a working first request takes about 5 minutes:
- Create an account at https://www.hefu.hk and generate an API key from the dashboard.
- Set your client's
base_urltohttps://api.hefu.hk/v1. - Set the
modelparameter to an identifier from the catalog, for example a Kimi K2.6, Qwen 3.8 Max, or GLM 5 entry (check https://www.hefu.hk/models for the exact string). - Send your first chat completion.
Using the official OpenAI Python SDK:
from openai import OpenAI
client = OpenAI(
api_key="your-hefu-key",
base_url="https://api.hefu.hk/v1"
)
resp = client.chat.completions.create(
model="kimi-k2.6", # illustrative; verify the exact identifier at /models
messages=[{"role": "user", "content": "Compare Kimi, Qwen and GLM for long-document analysis."}]
)
print(resp.choices[0].message.content)
That's the entire migration surface: base URL and key. Your prompt templates, streaming loops, and tool-calling code do not need to change.
Which Models Are Available and How to Switch
The authoritative model catalog is at https://www.hefu.hk/models. Because version identifiers rotate, always check that page before choosing a model name. Switching between model families only requires changing the model parameter:
- Kimi identifiers → long-context reading, Chinese summarization, deep document Q&A.
- Qwen identifiers → general reasoning, multilingual tasks, tool use across the 3.8/3.7/3.5 generations.
- GLM identifiers → coding, structured output, agentic workflows.
No code rewrite and no second key are needed. This matches the industry pattern we documented in GLM, Zhipu & MiniMax access outside China: 2026 guide.
The table below compares publicly available unified gateway information as of May 2026; platforms may adjust later, so refer to official pages.
| Gateway | Model families on one key | Billing model | Account / regional barrier |
|---|---|---|---|
| HeFu | Kimi, Qwen, GLM, plus routing to GPT, Claude, DeepSeek, Gemini, Wan, Seed, Grok, MiniMax | Pay-as-you-go; pricing subject to the official pricing page | Hong Kong node; registration does not require an overseas credit card (official site) |
| AIsa (source dated 2026-09-30) | Qwen, DeepSeek, Kimi, GLM, MiniMax, Hunyuan | Unified wallet, pay-as-you-go | AIsa documentation |
| Tencent Cloud TokenHub (2026) | DeepSeek, GLM, Kimi, MiniMax, Qwen | Single billing path | Requires a Tencent Cloud account (official docs) |
| Yotta (2026) | 50+ models, including GLM 5.2, DeepSeek V3.2 | OpenAI-compatible key | Yotta blog |
Performance, Pricing, and Choosing the Right Model
Each of the three model families has its own strength. Kimi focuses on long context, good for processing hundreds of pages of Chinese material; Qwen has the widest coverage of variants, including 8B–235B open-source weights for offline deployment and fine-tuning; GLM 5 is a strong candidate for coding and agentic workloads. On cost, the ecosystem is experimenting with subscriptions: as of February 25, 2026, the GLM Lite international coding plan was about $10/month, while the Chinese domestic version was about $3/month; Alibaba's $10/month coding plan covers roughly 18,000 requests and can be used across Qwen3.5-plus, Kimi K2.5, GLM-5, and MiniMax-M2.5, billed per request rather than per token (CodingPlan.org). On March 17, 2026, the AINFT platform added MiniMax-M2.5, Kimi-K2.5, and GLM-5 to a unified multi-model entry with usage-based billing (per ChainCatcher / CryptoRank.io reports), showing that "one key for multiple models" has become the default consumption pattern.
HeFu's own rates are not printed here as fixed numbers because prices change over time; visit https://www.hefu.hk/pricing for current pay-as-you-go figures. A pragmatic approach: in the catalog, benchmark Kimi, Qwen, and GLM against your main workload, pick the best one, and keep the other two one parameter away for A/B testing and fallback.
Enterprise-Grade Stability and Developer Support
Aggregation only matters when it works in production. HeFu routes traffic through its Hong Kong node with direct connections to upstream Chinese model vendors, and the platform is built for concurrent multi-model workloads, not just prototyping. API keys can be isolated and rotated as needed; https://www.hefu.hk/docs covers OpenAI-compatible clients, streaming, tool calling, error handling, and version notes, so teams can maintain one codebase while receiving model updates transparently.
FAQ
Do I need a Chinese phone number or real-name verification to use these aggregate APIs?
It depends on the gateway. TokenPAPA officially states "no Chinese phone number required," while A2Agent emphasizes "registration within 1 minute" (both subject to their current official wording). HeFu opens registration to developers worldwide through https://www.hefu.hk, does not require an overseas credit card, and its Hong Kong node is designed for users outside mainland China. Tencent Cloud TokenHub (2026), by contrast, is tied to Tencent Cloud accounts and may have regional verification requirements (official docs).
Can I switch existing OpenAI SDK code to HeFu with minimal changes?
Yes. Because HeFu exposes an OpenAI-compatible API at https://api.hefu.hk/v1, migration only involves two configuration changes: set base_url to that address and replace the key with a HeFu key. Switching between Kimi, Qwen, and GLM only changes the model string; prompts, streaming, and function calling code stay untouched.
Which is cheaper: a $10/month subscription or per-token pay-as-you-go?
It depends on call volume and how often you switch model families. As of February 25, 2026, Alibaba's $10/month plan covers about 18,000 requests, suitable for high-frequency multi-model experiments: you can switch between Qwen3.5-plus, Kimi K2.5, GLM-5, and MiniMax-M2.5 without watching token consumption (CodingPlan.org). For low-volume production traffic that mainly calls a single model, per-token billing is usually cheaper. HeFu's current pricing is on the official pricing page; estimate based on your actual request volume.
How do I know which model version identifier to use?
Version rotation is fast — MiniMax-M2.5, Kimi-K2.5, and GLM-5 entered some gateways on March 17, 2026 (per ChainCatcher / CryptoRank.io reports), and identifiers like Kimi K2.6 and Qwen 3.8 Max appeared later. HeFu's authoritative list is https://www.hefu.hk/models; always use the strings from that page, and do not rely only on second-hand blogs.
Is my data isolated when I route through an aggregator?
Reputable gateways isolate keys, logs, and billing by account. With HeFu, requests are billed to your own key, and https://www.hefu.hk/docs describes data handling and retention policies. For regulated workloads, read both the platform documentation and the model vendor's data-processing terms before going live.
FAQ
Do I need a Chinese phone number or real-name verification to use these aggregate APIs?
It depends on the gateway. TokenPAPA officially states "no Chinese phone number required," while A2Agent emphasizes "registration within 1 minute" (both subject to their current official wording). HeFu opens registration to developers worldwide through [https://www.hefu.hk](https://www.hefu.hk), does not require an overseas credit card, and its Hong Kong node is designed for users outside mainland China. Tencent Cloud TokenHub (2026), by contrast, is tied to Tencent Cloud accounts and may have regional verification requirements ([official docs](https://cloud.tencent.com/document)).
Can I switch existing OpenAI SDK code to HeFu with minimal changes?
Yes. Because HeFu exposes an OpenAI-compatible API at `https://api.hefu.hk/v1`, migration only involves two configuration changes: set `base_url` to that address and replace the key with a HeFu key. Switching between Kimi, Qwen, and GLM only changes the `model` string; prompts, streaming, and function calling code stay untouched.
Which is cheaper: a $10/month subscription or per-token pay-as-you-go?
It depends on call volume and how often you switch model families. As of February 25, 2026, Alibaba's $10/month plan covers about 18,000 requests, suitable for high-frequency multi-model experiments: you can switch between Qwen3.5-plus, Kimi K2.5, GLM-5, and MiniMax-M2.5 without watching token consumption ([CodingPlan.org](https://codingplan.org)). For low-volume production traffic that mainly calls a single model, per-token billing is usually cheaper. HeFu's current pricing is on the [official pricing page](https://www.hefu.hk/pricing); estimate based on your actual request volume.
How do I know which model version identifier to use?
Version rotation is fast — MiniMax-M2.5, Kimi-K2.5, and GLM-5 entered some gateways on March 17, 2026 (per [ChainCatcher](https://www.chaincatcher.com) / [CryptoRank.io](https://www.cryptorank.io) reports), and identifiers like Kimi K2.6 and Qwen 3.8 Max appeared later. HeFu's authoritative list is [https://www.hefu.hk/models](https://www.hefu.hk/models); always use the strings from that page, and do not rely only on second-hand blogs.
Is my data isolated when I route through an aggregator?
Reputable gateways isolate keys, logs, and billing by account. With HeFu, requests are billed to your own key, and [https://www.hefu.hk/docs](https://www.hefu.hk/docs) describes data handling and retention policies. For regulated workloads, read both the platform documentation and the model vendor's data-processing terms before going live.