GEO Blog

When the Answer Is Spoken: Doubao's SeedRealtime Upgrade and the Collapse of Brand Shortlists in China

2026/8/18 上午1:03:31

Doubao shipped free full-duplex audio-video AI to ~345M users on Aug 6, 2026. Spoken answers name one brand, not five — and text-based visibility scores no longer measure what brands actually get.

When the Answer Is Spoken: Doubao's SeedRealtime Upgrade and the Collapse of Brand Shortlists in China

On August 6, 2026, ByteDance shipped what looks on the surface like a consumer feature and is, for brand marketers, a structural change to how visibility works in China. Doubao's video call function was upgraded to run on SeedRealtime, a natively multimodal, full-duplex audio-video model — and it went out free to Doubao's entire user base, which QuestMobile put at roughly 345 million monthly active users in its March 2026 ranking, the largest of any AI-native app in China.

The technical claim is that audio, video, and text are handled end-to-end in one model, so Doubao can decide who to listen to, when to answer, and when to stay silent while the scene in front of the camera keeps changing. ByteDance called it the first deployment of audio-video full-duplex interaction at this scale.

Here is why that matters more than another benchmark win: a spoken answer has no room for a shortlist. Text answers return three to five brands. Voice answers return one, occasionally two. Every brand that has been quietly surviving on position 3 through 5 in Chinese AI recommendations is about to find out what its visibility is actually worth.

The shortlist was your safety net

Most of the brand visibility work done in China over the past eighteen months has implicitly optimized for inclusion, not primacy. That was rational. When a user types "推荐几个高端护肤品牌" into Doubao or Qwen, the model returns a paragraph naming several brands. Being named fourth still gets you into the consideration set. Our tracked panel across six Chinese models — Doubao, Kimi, DeepSeek, Qwen, Wenxin, and Hunyuan — shows the same pattern repeatedly: in text-mode category queries, a typical answer surfaces four to six named brands, and the gap in outcome between rank 1 and rank 4 is much smaller than the gap in score.

That is the safety net. A brand scoring in the 50s on our 0–100 index rarely leads an answer, but it is frequently in the answer. It gets seen.

Real-time voice destroys that arithmetic. Nobody wants a spoken assistant to read out six brand names while they are holding a phone up to a shelf. The conversational norm — set by the interaction design, not by the model's knowledge — is one recommendation, with a reason, and then a follow-up question. The model still knows about the other five. It just will not say them unless asked.

This is the single most important thing to internalize about the SeedRealtime shift: the retrieval did not change, the delivery budget did. Your brand's underlying score may be identical. Your realized exposure is not.

Camera-first queries change the question itself

The second change is subtler and arguably more consequential. Full-duplex video means the user no longer has to describe the product. They point at it.

Compare the two query shapes:

Text queryCamera + voice query
What the user suppliesCategory, budget, intent, sometimes a brand nameA physical object, a shelf, a room, a face
What the model must resolveBrand entities from text tokensBrand entities from visual features, then recommend
Where citation happensRetrieved web/UGC sourcesVisual recognition first, retrieval second
Typical output length150–400 characters, multi-brandOne to two sentences, single brand
Brand's leverage pointText corpus, structured data, UGC volumeVisual identity recognizability + text corpus

The last row is the one most brands are unprepared for. In a camera-first query, the model must first decide what it is looking at. If your packaging, logo lockup, or product silhouette is not reliably recognized, you are not competing on recommendation quality — you never entered the retrieval step at all. This is a failure mode that does not exist in text GEO and that no current visibility dashboard, including ours, fully measures yet.

International brands entering China are unusually exposed here. A product sold in China under a localized Chinese name, with different packaging than its global SKU, generates visual training signal that is thinner and more fragmented than the global version. The model may recognize the global bottle and not the China-market bottle — or recognize it and attach the global brand entity, which then retrieves global-language sources that the Chinese-language answer layer weights poorly.

What the numbers say about urgency

The scale argument is straightforward. Doubao's reported base of roughly 345 million MAU (QuestMobile, March 2026) sits far ahead of Qwen at approximately 166 million and DeepSeek at approximately 127 million in the same ranking. Other Q2 2026 reads put DeepSeek above 180 million and Baidu's Ernie Bot near 210 million — the counts differ by methodology, but the ordering at the top does not. Doubao is the volume leader, and it is the one shipping full-duplex video for free to everyone.

The free rollout matters. In our earlier analysis of Doubao's paid tier, we flagged that paywalled surfaces create measurement blind spots because they cannot be observed with standard API probes. SeedRealtime goes the other way: free, default-on after an app update, and therefore likely to see fast adoption across exactly the mainstream consumer segment brands care about. Adoption speed is the variable that turns a feature launch into a visibility event.

There is also a compounding effect with ByteDance's content graph. Doubao is trained on and connected to the Douyin, Toutiao, and Xigua ecosystem. A camera-based product query resolves against a corpus where short-form video reviews of that product already exist in volume. Brands with heavy Douyin presence get a double advantage in this surface: better visual recognition and better retrieval. Brands that treated Douyin as a paid-media channel rather than an owned-content asset get neither.

The measurement problem nobody has solved

Being candid about the limits of the current tooling — including ours — is more useful than pretending otherwise.

Every China GEO measurement approach in production today, ours included, probes models through text. You send a query, you parse the answer, you score whether the brand appeared and in what position. That methodology cannot observe:

  1. Spoken-answer truncation. Whether the voice surface returns one brand where the text surface returned five.
  2. Visual entity resolution. Whether the model recognizes your product from an image at all.
  3. Turn-level follow-up. Whether a user who hears one brand asks "什么其他选择?" and how the model orders the second answer.

Until someone builds probes that actually run against the multimodal surface, brand scores measured on text are a proxy that is quietly drifting from the thing they claim to measure. Treat any China AI visibility score you see this quarter — including on our own /brands pages — as a text-surface score, and assume the voice-surface distribution is more concentrated at the top.

Takeaway for Brand Marketers

Four things worth doing in the next quarter, ordered by effort-to-impact:

1. Re-read your scores as a primacy problem, not an inclusion problem. Pull your brand's position distribution in category queries across the six major models. If you are usually named but rarely named first, your text score is overstating your real-world exposure on voice and camera surfaces. Set the internal target to rank 1 share, not mention rate.

2. Audit visual recognizability of your China SKU. Take ten photographs of your actual China-market packaging — on shelf, in hand, poorly lit, partially occluded — and ask Doubao in video mode what it is. This costs an afternoon. If it misidentifies the product or attaches the wrong brand entity, that is now a visibility defect, and it will not show up on any text dashboard.

3. Treat Douyin content as retrieval infrastructure. Not as a campaign channel. The question is not "did the video perform" but "does a durable corpus of Chinese-language content about this product exist inside ByteDance's graph." Volume of evergreen, specific, product-named content beats a small number of high-performing hero videos.

4. Push your evidence into the sources voice answers can compress. A spoken recommendation needs one defensible reason — a certification, an award, a specific spec, a ranking. Brands whose Chinese-language corpus contains a single crisp, repeatable claim get quoted; brands whose corpus is diffuse brand-adjective marketing do not survive compression into one sentence.

The broader point: China's AI surfaces are moving from lists to answers. The list era rewarded being present. The answer era rewards being first. Most brand GEO programs in China are still optimizing for the previous era, and the gap will not be visible in the dashboards until it has already cost something.


Related: See how brands score across all six Chinese AI models on /brands, and our earlier analysis of Doubao's Seedance and Seedream video models — generation-side, where this piece covers the understanding side.

Sources: QuestMobile China AI app MAU rankings (March 2026); ByteDance SeedRealtime announcement, August 6, 2026, via Sina Tech and Tencent News; QuestMobile / company Q2 2026 disclosures.