When Chinese AI apps were free, every user hit the same model with the same retrieval budget. That assumption quietly died this year. Doubao now sells three paid tiers at RMB 68, 200, and 500 per month. Kimi sells four, from RMB 49 to RMB 699. And the features gated behind those paywalls — deep research, multi-step agents, document and deck generation — are exactly the features that change how many sources a model reads before it names a brand.
Almost every brand visibility score in circulation today, including the ones on this site, was measured on the free tier. That is no longer the same product your highest-intent Chinese buyers are using.
The tier split arrived faster than the measurement caught up
Doubao disclosed its paid plan in an App Store service statement on May 4, 2026: a standard tier at RMB 68/month, an enhanced tier at RMB 200/month, and a professional tier at RMB 500/month. The free base tier stays. What moves behind the paywall is high-compute work — deck generation, data analysis, video production, and the long-running agentic tasks that consume real inference budget.
Kimi got there first, with a four-tier membership at RMB 49, 99, 199, and 699 per month. Its structure is more revealing than Doubao's because it is metered rather than flat: all paid features draw from a single shared quota pool, and deep research, Office file handling, site deployment, and Kimi Code all deduct from it. Basic chat with the current K-series model stays free and does not touch the quota.
The economics behind this are not Chinese-specific. GPU rental rates and storage prices rose sharply through 2026, and a platform serving Doubao's volume — roughly 345 million MAU as of March 2026, with daily token calls past 120 trillion and doubling every three months — cannot absorb unlimited deep-research runs on a free plan. Paid tiers were inevitable. What was not inevitable is that the industry kept scoring brands as if one model equals one behavior.
Free tier and paid tier are different retrieval regimes
The gap is not model quality. On most Chinese platforms, free and paid users often reach the same base model. The gap is retrieval breadth, and it is large.
A standard chat answer typically resolves against parametric memory plus a handful of retrieved pages. A deep research run behaves differently in kind: it decomposes the question, issues many parallel queries, and reads dozens to hundreds of sources before writing a structured, cited report. Published comparisons of Western deep research implementations put the range at roughly 50 to 200 sources per run, against a small single-digit count for a normal answer. Chinese platforms have not published equivalent figures, and I would not assume the numbers transfer exactly — but the architectural difference is the same one, and the directional consequence is unavoidable.
| Free chat tier | Paid deep research tier | |
|---|---|---|
| Sources consulted | Handful, or none | Dozens to hundreds |
| Dominant signal | Parametric memory, brand priors | Retrievable evidence, freshness |
| Query cost to user | Zero | Metered quota or subscription |
| Typical use | Quick question, casual browsing | Comparison, shortlist, purchase justification |
| Output | Conversational answer | Structured report, often shared internally |
Two answers to the same brand question, produced by the same model on the same day, are drawing on evidence pools that differ by an order of magnitude or more.
This inverts which brands win
Here is the part that should worry anyone reporting a single brand score.
On the free tier, a model leaning on parametric memory rewards name recognition. Global brands with decades of Chinese-language presence, dense Baidu Baike and Wikipedia entries, and heavy mention volume in training data surface effortlessly. This is why our own category benchmarks keep producing perfect or near-perfect scores for BMW, Chanel, KFC, and Estée Lauder while credible competitors land at zero. The free tier is, in effect, a brand-equity mirror.
Deep research mode does not work that way. It goes looking. And when it goes looking, the ranking reshuffles along a different axis: whichever brand has the most retrievable, current, Chinese-language evidence wins the citation — regardless of how famous it is.
That produces two failure modes brands are not currently testing for:
The incumbent's blind spot. A household name that scores 100 on free-tier prompts can be dismantled inside a deep research comparison, because when the model actually reads sources, it finds thin, outdated, or English-only Chinese-market material and defers to a competitor with better documentation. High free-tier score, poor paid-tier outcome. This is the more dangerous of the two, because the brand's dashboard says everything is fine.
The challenger's invisible win. A brand scoring zero on free-tier prompts may already be surfacing in deep research reports, because it publishes detailed specification pages, comparison content, and third-party review coverage that a retrieval-heavy run will find. The visibility exists; the measurement misses it entirely.
Both failure modes come from the same mistake: treating "does Doubao mention my brand" as a single question when the platform now answers it two different ways depending on who is asking.
Metering makes paid queries the ones that matter
The obvious pushback is volume. Free users vastly outnumber paid ones, so surely free-tier visibility is what counts.
Kimi's quota design argues the opposite. When deep research draws down a finite monthly allowance, users do not spend it casually. They spend it on questions worth the quota — vendor comparisons, purchase justifications, category research they intend to act on. Metering acts as an intent filter. The small minority of queries that consume paid quota are disproportionately the queries attached to a decision.
There is a second amplifier. Deep research outputs are not ephemeral chat. They are structured, cited documents, and they get forwarded — pasted into procurement decks, dropped into WeChat work groups, attached to internal recommendations. One deep research report that omits your brand can foreclose consideration across a buying committee that never queried an AI model at all. A free-tier chat answer evaporates in the scroll; a paid-tier report becomes an artifact with a shelf life.
For B2B and considered-purchase categories in China, the paid tier is not a niche. It is the tier where shortlists are formed.
Your surface count just went up again
We flagged in early August that Qwen's open-weight releases multiply the number of distinct "Qwen surfaces" a brand needs to test. Tier stratification stacks on top of that, and the arithmetic gets uncomfortable fast.
Six major Chinese models, tested only at their default free-tier chat surface, gives you six measurement points. Add a paid deep research surface for each platform that offers one and you are at eleven or twelve. Add the open-weight and third-party-hosted variants, plus the in-app assistants embedded in Taobao, Douyin, and WeChat, and the honest surface count for a brand serious about China is somewhere north of twenty.
Nobody should test all twenty at equal frequency. But testing exactly one and calling it a brand score is now indefensible. The practical minimum is two surfaces per platform — free chat and paid deep mode — on the two or three platforms that actually carry your category.
Takeaway for brand marketers
Re-baseline on both tiers before you set 2027 budget. Take your ten highest-value category queries and run each twice on Doubao and once more on whichever second platform matters for you: once in default chat, once in deep research mode. Compare not just whether you appear, but which sources the deep run cites. Two scores per platform, not one.
Read the gap, not the score. If your free-tier score is high and your deep-tier result is weak, you have a brand-equity buffer that is actively hiding an evidence problem — and it will erode as paid adoption grows. If the gap runs the other way, your content is working and your measurement is broken. The direction of the gap tells you which team to fund.
Fund retrievable evidence, not mentions. Free-tier visibility responds to mention volume and brand fame, which are slow and expensive to move. Paid-tier visibility responds to whether detailed, current, Chinese-language material about your products exists somewhere a retrieval run can reach it — specification pages, comparison content, third-party reviews, structured data. That is a much cheaper lever, and it is the one that compounds as deep research adoption rises.
Budget for a moving target. Paid tiers launched in May. Adoption curves for consumer AI subscriptions in China are still early, and Doubao's advertising system is expected to land in Q4. The measurement regime that was adequate in 2025 has now been invalidated twice in one year. Assume it will be invalidated again.
The uncomfortable summary: the number in your brand visibility dashboard is probably accurate, and probably describes a tier your buyers are leaving.
Related: See how individual brands score across all six Chinese AI models on our brand rankings, and read our earlier analysis on surface proliferation from open-weight releases and Doubao's Q4 ad rollout.
Sources: Doubao App Store service statement (May 4, 2026), reported by 36Kr and Sina Tech; Kimi official membership pricing documentation; QuestMobile March 2026 China AI-native app MAU data; published deep research mode implementation comparisons.