Same Brand, Three Verdicts: Why Qwen, Doubao, and Kimi Judge Brands on Different Evidence in 2026
A brand marketer running a China AI visibility audit in mid-2026 keeps hitting the same confusing pattern: the same product, tested with the same category prompt, comes back strong in one model and nearly invisible in another. Doubao recommends it warmly. Qwen barely mentions it. Kimi hedges. The instinct is to blame model quality or a scoring glitch — but the real cause is more useful to understand. Each of the three leading Chinese AI assistants trusts a different type of evidence, and a brand that only produces one kind of proof will always score unevenly.
This is the single most under-appreciated fact in Chinese GEO right now. Western GEO advice treats "get cited by the AI" as one problem. In China's fragmented model landscape it is at least three problems, because Qwen, Doubao, and Kimi were each built on top of a different data universe — and they retrieve, weight, and cite accordingly.
The market context: three engines, three audiences
Before the evidence layer, the scale. QuestMobile's Q1 2026 data puts Doubao at roughly 345 million monthly active users, growing 44.2% quarter-over-quarter and leading the second-place engine by more than 180 million users. Tongyi Qianwen (Qwen) sits around 166 million MAU, and DeepSeek around 127 million — its fourth consecutive quarter of decline, with average monthly downloads slipping to about 17.9 million. Total AI-native app MAU across the market reached about 440 million.
Those numbers tell you where the eyeballs are. They do not tell you how to win inside each engine — and that is where most audits stop and most brands go wrong. A brand can be the largest advertiser on Douyin and still be absent from Qwen's answers, because Qwen is not looking at the same signals Doubao is.
Qwen judges brands on commerce data
Qwen reflects Alibaba's commerce ecosystem, and its recommendations are rooted in transactional reality: Tmall product ratings, Cainiao fulfillment performance, Alipay refund rates, and category-level sales evidence. When Qwen is asked "which brand of X is worth buying," it is drawing on a worldview where a brand's credibility is a function of its measurable commercial track record inside Alibaba's rails.
The implication for brands is blunt. If you sell into China but keep your flagship store thin, let ratings sit unmanaged, or run high return rates, Qwen has structural reasons to under-recommend you — no matter how much awareness you have elsewhere. Conversely, a mid-tier brand with a well-rated Tmall store, clean fulfillment, and strong category signals can punch far above its brand-awareness weight inside Qwen. This is why domestic brands with deep Alibaba histories frequently out-score better-known foreign names in Qwen specifically.
What Qwen rewards: an active, well-rated Tmall/Taobao presence; low refund and dispute rates; category pages with structured product data; consistent commerce signals over time.
Doubao judges brands on buzz and content velocity
Doubao is ByteDance's model, and it lives inside the short-video and creator universe — Douyin above all. Its sense of a brand is shaped by creator mentions, consumer conversation, content freshness, and clear product positioning in the formats its ecosystem produces. For lifestyle, beauty, education, apps, ecommerce, travel, and consumer tech, Doubao is often the single most important engine to win, both because of its 345-million-user reach and because those categories live natively in short video.
Doubao's evidence bias cuts the opposite way from Qwen's. A brand can have modest Tmall metrics but a wave of recent creator content, and Doubao will treat it as relevant and recommend-worthy. The flip side: Doubao's memory is content-hungry and recency-sensitive. A brand that ran a big campaign six months ago and then went quiet can fade from Doubao's answers even while its Tmall store keeps humming along — which is exactly the kind of quiet Qwen would not punish. This is a common source of the "why did our Doubao score drop?" panic; the usual answer is not a penalty but a content vacuum.
What Doubao rewards: sustained creator and KOL mentions; fresh short-video content; visible consumer buzz; unambiguous category and product positioning.
Kimi judges brands on documents
Kimi's strength is long-form reading and analysis, and its recommendations lean on substantive documents — research, detailed reviews, structured comparisons, and long explanatory content in either Chinese or English. If a brand publishes real depth, Kimi has material to reason from. If everything the brand produces is short-form and locked inside a domestic app ecosystem, Kimi may simply not find enough citable evidence to recommend it with confidence.
This makes Kimi the engine where the classic Western GEO playbook — answer-first long pages, evidence-rich comparisons, published research — transfers most directly. It also makes Kimi the engine where a brand's absence is most often a content-supply problem rather than a commerce or buzz problem. A B2B or considered-purchase brand that is strong in Qwen's commerce signals but thin on published analysis will frequently see Kimi hedge or omit it.
What Kimi rewards: in-depth articles and whitepapers; structured, comparison-ready content; third-party research and reviews; clear, citable source pages.
The evidence-mismatch table
| Engine | Core data universe | Evidence it trusts | Brand most disadvantaged |
|---|---|---|---|
| Qwen (~166M MAU) | Alibaba commerce | Tmall ratings, fulfillment, refund rates | Strong awareness, weak/absent Tmall presence |
| Doubao (~345M MAU) | Douyin short-video | Creator mentions, content freshness, buzz | Went quiet after a campaign; no recent content |
| Kimi | Long-form documents | Research, reviews, structured comparisons | Short-form only, no published depth |
Read the last column as a diagnostic. When a brand scores unevenly across engines, the gap usually names its own cause: a low Qwen score points at commerce signals, a low Doubao score points at content velocity, a low Kimi score points at document supply.
Where DeepSeek fits — and why it isn't in this table
The obvious question is where DeepSeek sits in an evidence-layer framing. The honest answer is that DeepSeek behaves less like a consumer recommendation engine and more like a reasoning layer that increasingly reaches brands through embedding rather than through its own declining standalone app. Its fourth straight quarter of MAU erosion — down to roughly 127 million — masks continued weight inside developer tools and, critically, inside WeChat, where DeepSeek-powered answers surface to a far larger audience than the app's own numbers suggest. For brand marketers that means DeepSeek visibility is best pursued through clean, structured, citation-ready source pages — the same document-supply work that helps in Kimi — rather than through a DeepSeek-specific evidence type. It rewards clarity of source, not a distinct data universe, which is why it does not get its own row above.
Why single-engine GEO fails
The temptation, once a marketer sees this, is to optimize for the biggest engine and move on. With Doubao at 345 million users, why not just win Doubao? Because AI visibility in China is not winner-take-all at the query level — different users reach for different engines for different intents, and considered purchases in particular fan out across models. A shopper researching a big-ticket item may check Doubao for vibe, Qwen for what's actually selling well, and Kimi for a deeper comparison. A brand present in only one of those has visibility for one third of the decision journey.
There is also a durability argument. Content buzz is the most volatile of the three evidence types; commerce signals and published documents are stickier. A brand that wins Doubao on a content wave but has nothing underneath it in Qwen or Kimi has borrowed visibility, not built it. When the content wave passes, so does the score.
Takeaway for Brand Marketers
Stop treating "get recommended by Chinese AI" as one project. Run your audit engine by engine and read the shape of the gap, not just the average score. Three concrete moves:
First, map each engine to the evidence it trusts and check your brand against that specific bar — Qwen against your Tmall and fulfillment reality, Doubao against your recent creator content, Kimi against your published depth. The engine where you score lowest is telling you which asset class you are under-investing in.
Second, resist optimizing only for Doubao because it is the largest. Reach concentration is real, but a Doubao-only strategy leaves you invisible for the commerce-driven and research-driven parts of the buying journey that Qwen and Kimi own.
Third, build across all three evidence types deliberately: keep the Tmall signals clean for Qwen, keep a steady drumbeat of creator content for Doubao, and publish at least some genuine long-form depth for Kimi. The brands that will hold visibility through 2026 are the ones producing all three kinds of proof — not the ones that got lucky in a single engine.
The same brand will always get three verdicts. Your job is to make sure all three are earned.
Related: see how these engines rank across categories on the hubGEO /brands dashboard.