How to Use DeepSeek, Kimi and Qwen in One App

How to use DeepSeek, Kimi and Qwen in one app: a practical guide from HeFu.

HeFu · Published 2026-10-03

10 min read

Building an app that uses DeepSeek, Kimi and Qwen does not require three SDKs, three API keys or three separate codebases. HeFu exposes all three model families through one OpenAI-compatible endpoint (https://api.hefu.hk/v1), as documented in the official HeFu API docs(截至 2026 年 9 月): you integrate once, and you can route a request to DeepSeek-V4-Pro, Kimi K3 or Qwen 3.8 Max by simply changing the model name in each call, with one authentication flow, one billing relationship and one support channel.

Why Consolidate Three Models into One App

Each of the three model families was built for a different job. DeepSeek is the reasoning and code specialist (DeepSeek official site), Kimi is the long-context champion for Chinese-language scenarios (Moonshot platform docs), and Qwen is the balanced generalist with the broadest multimodal and open-source ecosystem (Qwen project page). Routing tasks across all three instead of committing to a single vendor improves output quality per task and protects you from vendor lock-in, while a unified API layer keeps maintenance costs close to zero.

The technical differences are substantial. As of Sep 2, 2026, Wavect's model comparison (wavect.ai) reports that Kimi K3 is a 2.8-trillion-parameter native multimodal mixture-of-experts model with a 1M-token context window, released under a proprietary license. DeepSeek V4, by contrast, ships Pro and Flash variants under the permissive MIT license with open weights and the same 1M-token context (DeepSeek GitHub), making it the natural choice for teams that care about transparency and cost. Qwen 3.8 offers an Apache-2.0 27B multimodal model with a 262K native window plus a 2.4T sparse model under the Max license (Qwen GitHub). As always, confirm these model-card figures on the vendors' official pages.

A free-tier comparison published by Bryan Software on Aug 14, 2026 adds a practical distinction that often decides the architecture: DeepSeek V4's two variants are pure text models and accept no image, audio or video input, while Kimi K3 and Qwen 3.8 both accept image and video input, with Qwen 3.8 handling videos up to two hours long (as reflected in the HeFu model catalog as of 2026 年 8 月). If your app needs multimodal understanding, that single capability gap determines which model must be in the pipeline. For a full walkthrough of the unified-key approach, see our earlier guide: DeepSeek, Qwen, Kimi on One API Key: The Complete Guide.

Getting Started with the HeFu Unified API

The setup takes three steps:

  1. Create a HeFu account and obtain one API key at https://www.hefu.hk.
  2. Point your OpenAI-compatible client to https://api.hefu.hk/v1.
  3. Call any model by using its identifier from the official model catalog.

The Python example below shows the entire pattern:

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_HEFU_API_KEY",
    base_url="https://api.hefu.hk/v1",
)

client.chat.completions.create(
    model="deepseek-v4-pro",
    messages=[{"role": "user", "content": "Solve this math problem step by step..."}],
)

client.chat.completions.create(
    model="kimi-k3",
    messages=[{"role": "user", "content": "Summarize this 500-page PDF and extract action items..."}],
)

client.chat.completions.create(
    model="qwen-3.8-max",
    messages=[{"role": "user", "content": "Localize this campaign copy for three markets..."}],
)

Only the model field changes between requests. This is the same pattern the wider industry already uses: as of 2026, Together AI's official documentation shows how setting OPENAI_API_KEY, OPENAI_BASE_URL and OPENAI_MODEL lets you switch between models such as DeepSeek V4 Pro, Kimi K2.7 Code and Qwen Code without touching request logic(实际支持名单以官方文档为准). HeFu applies the same principle across DeepSeek, Kimi and Qwen, so any SDK or framework that speaks the OpenAI protocol works unchanged. Request-level details, rate limits and authentication flows are documented at https://www.hefu.hk/docs.

Which Model Should You Call for Each Task

Choosing the right model per request is where most of the quality gain lives.

DeepSeek is the default for complex logical reasoning, mathematics, code generation and high-throughput text pipelines. Use DeepSeek-V4-Pro for the hardest reasoning tasks, and V4-Flash when latency or cost per token matters more than peak accuracy. Because DeepSeek is text-only as of Aug 2026 (DeepSeek API docs), keep it in the text-processing layer of your app rather than in multimodal features.

Kimi is the default for long documents and long conversations. With a 1M-token context window, Kimi K3 absorbs large PDFs, codebases and multi-turn chat histories, which makes it the strongest choice for Chinese-language business scenarios. According to our systematic comparison, Kimi leads in long-horizon agent and tool-calling scenarios as of Aug 2026, while DeepSeek concentrates its capabilities on reasoning and context processing, and Qwen's ecosystem emphasizes agent environments.

Qwen is the balanced generalist. For multilingual content, summarization, classification and any task that needs image or video input, Qwen 3.8 Max is the reliable default, and the open Qwen 3 family (8B–235B, including Thinking, Instruct, VL and Coder variants) gives you a self-host path when data isolation matters. When in doubt, start with Qwen for general tasks and escalate only the tasks that genuinely need DeepSeek's reasoning depth or Kimi's context length.

Model Comparison Table

Model family (HeFu catalog)Core strengthBest forInput modalitiesContext windowAPI integration
DeepSeek-V4-Pro / V4-FlashReasoning, code, cost-efficient textMath, codegen, high-throughput textText only (as of Aug 2026)1M tokensLow — same endpoint
Kimi K3Long context, agent capabilityLarge PDFs, multi-turn chat, Chinese scenariosImage + video input1M tokensLow — same endpoint
Qwen 3.8 Max / Qwen 3.8-27BBalanced, multilingual, multimodalGeneral tasks, video analysis up to 2 hoursImage + video (up to 2 h)262K native → 1MLow — same endpoint

*Model parameters and context windows: as of Sep 2026, per vendor model cards and the HeFu model catalog; subject to official pages.* API integration difficulty is identical across the three families because HeFu serves them behind the same endpoint and the same authentication. As of Sep 2026, DeepSeek V4 and Kimi K3 both offer 1M-token contexts, while the Qwen 3.6 line starts at 262K native and can be extended to 1M. License differences matter mainly for self-hosting: DeepSeek V4 is MIT-licensed, Qwen 3.8-27B is Apache-2.0, while Kimi K3 and the Qwen Max tier are proprietary. The authoritative current list of models and identifiers is the model catalog page(以官方页面为准).

Switching Models Dynamically in Production

Because the interface is uniform, routing decisions can be made at runtime with a few lines of code. A common pattern is a router that classifies each request and selects a model accordingly:

def route(task_type: str) -> str:
    if task_type in ("long_doc", "chat_history"):
        return "kimi-k3"
    if task_type in ("math", "code", "reasoning"):
        return "deepseek-v4-pro"
    return "qwen-3.8-max"  # general and multimodal

Three production use cases make this pattern valuable. First, cost control: send routine text volume to DeepSeek-V4-Flash and reserve the flagship reasoning model for the requests that actually need it. Second, fallback resilience: if one model family is degraded or rate-limited, you re-route traffic to another without redeploying. Third, quality A/B testing: generate the same response from two models and compare side by side before committing to a default — the same idea popularized by consumer multi-model apps, which let users compare answers from several vendors in one conversation thread.

Leveraging HeFu's Built-in Tools

A unified API is most useful when the surrounding workflow is already built. HeFu ships ready-made tools that connect to the same models, so you do not have to assemble the glue code yourself:

  • Marketing Studio — generate and localize campaign content; use Qwen for multilingual copy and DeepSeek for structured analysis.
  • Seller Assistant — build Amazon listings; Kimi's long context handles product briefs, manuals and review digests.
  • HR Assistant — screen resumes and prepare interview summaries with DeepSeek's reasoning power.
  • Dev Toolkit — code review, test generation and documentation powered by DeepSeek-V4-Pro.
  • Creative Studio — multimodal concept work with Qwen's image and video understanding.
  • Support Widget — customer-service conversations with long memory, where Kimi's context window shines.

Each tool is pre-wired to the unified API, so the model choice is a configuration change rather than a code change.

Pricing and Configuration Notes

HeFu's API pricing and subscription plans are subject to change, so treat the official pages as the source of truth: https://www.hefu.hk/pricing for pricing, and the model catalog for the current model list. For a broader cost benchmark across vendors, read our Chinese LLM API Pricing Comparison 2026(截至 2026 年,以官方页面为准).

Two additional data points are worth keeping in mind. First, the multi-model bundle is already validated in the consumer market: as of 2026, Fello AI charges US$9.99 per month for access to Claude, ChatGPT, Gemini, Grok and DeepSeek in a single Mac/iOS app with unified conversation history and a global shortcut(价格以 Fello AI 官方页面为准). Second, self-hosting heavy models is not realistic for most small teams: according to a hardware analysis by Fello AI (as of Sep 2, 2026), DeepSeek-V4-Pro weights occupy roughly 865 GB across 64 shards (about 4 bits per parameter), and Kimi K2.6 around 595 GB across 64 shards — neither can be compressed through further quantization to fit a single consumer GPU. Exact shard/weight figures vary by release and should be checked in the DeepSeek GitHub repository and the Moonshot platform docs(以官方仓库信息为准). An API-based aggregation layer is therefore the most economical route for production apps that need more than one model.

FAQ

Do I need separate API keys for DeepSeek, Kimi and Qwen?

No. One HeFu API key works across all three model families. You authenticate once against https://api.hefu.hk/v1, and the model identifier in each request determines which model handles that call.

Can I switch models without rewriting my code?

Yes. DeepSeek, Kimi and Qwen are all served through the same OpenAI-compatible interface on HeFu, so the only thing that changes between calls is the model string. Your existing OpenAI SDK code, prompt templates and inference pipelines continue to work unchanged.

Which model is best for long-document analysis?

Kimi. The Kimi K3 model provides a 1M-token context window (Moonshot platform docs) and is specifically optimized for long-context processing, making it the default choice for large PDFs, codebase reviews and long multi-turn conversations, especially in Chinese-language scenarios.

How do DeepSeek, Kimi and Qwen compare on tool calling and agent tasks?

Based on our systematic comparison as of Aug 2026, Kimi leads in long-horizon agent and tool-calling scenarios; DeepSeek focuses on reasoning and long-context processing; and Qwen's ecosystem emphasizes agent environments. Full results are in our tool-calling compatibility guide.

Where can I find the latest model list and pricing?

Check the official model catalog at https://www.hefu.hk/models and the pricing page at https://www.hefu.hk/pricing. For cost benchmarks across vendors, read our Chinese LLM API pricing comparison.

FAQ

Do I need separate API keys for DeepSeek, Kimi and Qwen?

No. One HeFu API key works across all three model families. You authenticate once against `https://api.hefu.hk/v1`, and the model identifier in each request determines which model handles that call.

Can I switch models without rewriting my code?

Yes. DeepSeek, Kimi and Qwen are all served through the same OpenAI-compatible interface on HeFu, so the only thing that changes between calls is the `model` string. Your existing OpenAI SDK code, prompt templates and inference pipelines continue to work unchanged.

Which model is best for long-document analysis?

Kimi. The Kimi K3 model provides a 1M-token context window ([Moonshot platform docs](https://platform.moonshot.cn/docs)) and is specifically optimized for long-context processing, making it the default choice for large PDFs, codebase reviews and long multi-turn conversations, especially in Chinese-language scenarios.

How do DeepSeek, Kimi and Qwen compare on tool calling and agent tasks?

Based on our systematic comparison as of Aug 2026, Kimi leads in long-horizon agent and tool-calling scenarios; DeepSeek focuses on reasoning and long-context processing; and Qwen's ecosystem emphasizes agent environments. Full results are in our [tool-calling compatibility guide](/en/blog/chinese-llm-tool-calling-compatibility-comparison).

Where can I find the latest model list and pricing?

Check the official model catalog at [https://www.hefu.hk/models](https://www.hefu.hk/models) and the pricing page at [https://www.hefu.hk/pricing](https://www.hefu.hk/pricing). For cost benchmarks across vendors, read our [Chinese LLM API pricing comparison](/en/blog/chinese-llm-api-pricing-comparison-2026).

Related reading

What Is the Cheapest Way to Access DeepSeek API in Oct 2026?

What is the cheapest way to access DeepSeek API: a practical guide from HeFu.

Developing with GPT-6: API Access Guide for Developers (As of Oct 2026)

GPT-6 API access: a practical guide from HeFu.

GPT-6 Sol and Luna Models — A Practical Guide for Developers and Businesses

GPT-6 Sol and Luna models: a practical guide from HeFu.

Want to try these models yourself?

HeFu aggregates every major LLM behind one OpenAI-compatible API — pay as you go.

Prices quoted are official list prices for reference — see the main site pricing page for actual rates.

🔥 Join today's AI debate — cast your vote →

Start Free TrialBook an Enterprise Demo