GEO Blog

Your Brand Scores 100 and Still Loses the Sale: China's AI Price Hallucination Problem in 2026

2026/8/19 上午1:03:53

Doubao quoted a MacBook Air M3 at RMB 2,550 in a five-model shopping test. With Qwen-Taobao and Doubao's commerce loops now closed, brand fact accuracy matters as much as visibility.

Your Brand Scores 100 and Still Loses the Sale: China's AI Price Hallucination Problem in 2026

In June 2026, the Chinese tech outlet 雷科技 (Lei Technology) ran a simple test. It gave five of China's biggest AI assistants — Doubao, DeepSeek, Qwen, Kimi and Tencent Yuanbao — a scenario any Chinese parent would recognize: a high-school graduate heading to university, a RMB 10,000 budget, and a shopping list of one phone, one laptop and one tablet.

Doubao returned a confident, well-structured Apple bundle totalling RMB 9,850. It listed an iPhone 17 128GB at RMB 4,099 and a MacBook Air M3 8+256GB at RMB 2,550.

The MacBook line is off by roughly two-thirds of the actual retail price. The bundle it recommended cannot be purchased at the price it quoted, from any channel, anywhere in China.

Apple's brand visibility score in Chinese AI models is excellent. Apple was recommended first, described accurately as a category leader, and placed at the top of the answer. By every metric the GEO industry currently sells, this was a win. It was also a fabricated offer that no retailer will honour — and in the closed-loop commerce environment China's AI apps built this year, that gap is no longer a curiosity. It is a conversion problem, a channel-conflict problem, and as of this summer, a regulatory-complaint problem.

The loop closed in May, and nobody re-checked the facts

Until this year, an inaccurate AI answer about your product was a nuisance. The user still had to leave the chat, open Taobao or JD, and confront reality at the product page. The price correction happened before the transaction.

That buffer is gone. On May 11, 2026, Alibaba connected its Qwen app directly to Taobao, letting users search, compare and complete orders in natural language without leaving the assistant. ByteDance moved on a parallel track: first Douyin commerce integration in December 2025, gray-scale in-app shopping in March 2026, and the full "Help You Choose" (帮你选) closed loop live in May 2026. ByteDance is reported to be layering deeper e-commerce functionality into Doubao through Q3 2026, with full operational rollout targeted for Q4.

The scale behind those integrations is not marginal. QuestMobile put Doubao at roughly 345 million MAU in Q1 2026, Qwen at 166 million, and DeepSeek at 127 million. In April 2026 Doubao ranked second globally by app MAU at around 336 million, behind only ChatGPT.

So the shopping journey now compresses into a single surface: the model states a price, the model states a spec, the model produces a purchase path. Every factual error the model makes about your brand travels all the way to the checkout button.

Visibility metrics do not measure this

Almost every GEO scorecard on the market — including the brand scores we publish at hubGEO — answers one question: does the model mention you? Mention rate, ranking position, share of voice, sentiment. These are real and they matter. They also say nothing about whether the sentence containing your brand name is true.

There are at least four distinct failure modes hiding underneath a high visibility score:

Failure modeWhat the model doesCommercial consequence
Price fabricationQuotes a price from stale training data, a grey-market listing, or pure interpolationUser anchors on a price you never offered; abandons at checkout or demands price-matching
Spec driftAttaches last-generation specs to the current SKUWrong-fit purchase, higher return rate, review damage
Availability errorRecommends a discontinued SKU, or claims a live SKU is unavailableDemand routed to a product you cannot ship
Channel misattributionNames the wrong authorised retailer or omits your flagship storeTraffic delivered to grey-market or competitor channels

None of these register as a visibility problem. All four register as a brand problem the moment the answer becomes the storefront.

Worth noting: this is a different failure than the entity-fact divergence we documented in the three-baike problem, where models disagreed about a brand's founding year, ownership or country of origin because they read different encyclopedias. That was an identity issue. This is a transactional one — the facts that change monthly, that no encyclopedia tracks, and that no model refreshes without a live data feed.

The regulator is already counting

Here is the part brand teams in Shanghai and Shenzhen are tracking that headquarters usually isn't.

The China Consumers Association (中消协) received 985,928 consumer complaints in the first half of 2026, resolved 567,926 of them, and recovered RMB 447 million for consumers. For the first time, AI customer service problems entered the CCA's five named complaint hotspots, alongside second-hand trading disputes, ETC false advertising, jewellery quality and irregular online lending.

The specific grievances the CCA named are precisely the ones described above: AI systems making explicit commitments about fees, discounts and after-sales policy that consumers then rely on when purchasing; general-purpose AI generating content with factual errors that mislead consumers in transactional contexts; and no clean escalation path to a human.

The CCA's separate 618 analysis, published June 26, 2026, went further — flagging AI-generated fake product reviews, misleading livestream clips, and AI-fabricated scenarios as emergent problems of the promotional season.

Read that from a brand's chair. A consumer sees a price for your product inside Doubao, buys expecting it, and doesn't get it. The complaint they file names your brand, not ByteDance's model. Liability allocation between platform and brand for AI-generated commercial statements is unsettled in China today — which is exactly why brands should not wait for it to be settled before instrumenting the problem.

What brand teams should actually do

1. Add a fact-accuracy layer to your monitoring, separate from visibility. For your top 10 SKUs, run monthly probes across Doubao, Qwen, DeepSeek, Kimi, Yuanbao and Wenxin asking for price, current-generation spec, availability, and authorised purchase channel. Score each answer as correct / stale / fabricated. You now have two numbers instead of one: how often you appear, and how often you appear correctly. In our experience the second number is materially worse than the first, and nobody is reporting it.

2. Prioritise the commerce-integrated surfaces. Doubao (via Douyin) and Qwen (via Taobao) have closed loops. An error there converts to a transaction attempt immediately. An error in a model without a checkout path costs you less. Weight your remediation accordingly rather than treating all six models as equal.

3. Make your own price and spec data machine-readable and current. The models that quote you correctly are the ones reading a live, structured source — your Tmall/JD flagship product data, your official WeChat mini-program, your .cn product pages with Product and Offer schema carrying current price and availability. If your authoritative pricing lives only in a PDF price list or an image on a landing page, the model will interpolate, and interpolation is where RMB 2,550 MacBooks come from.

4. Audit the grey-market gap. Fabricated low prices frequently trace back to unauthorised listings the model treated as valid evidence. Suppressing grey-market listings has always been a margin exercise; it is now also an AI-accuracy exercise.

5. Log the errors with timestamps and screenshots. If liability rules do tighten, the brand that can show it detected, documented and escalated a platform's pricing error is in a very different position from the brand discovering it through a consumer complaint.

Takeaway for brand marketers

The first phase of GEO in China was about presence: get mentioned, get ranked, stop being invisible. That phase is not over, but it is no longer sufficient. Once the assistant became the storefront — Qwen in May, Doubao in May, deeper commerce integration through Q4 — the accuracy of the sentence became as commercially load-bearing as the existence of the sentence.

Track two numbers from Q3 onward. Visibility: how often do the models put you in the answer. Fidelity: how often is what they say about you actually true. A brand at 100 visibility and 60 fidelity is in worse shape than it looks, because it is spending budget to be confidently misrepresented at scale to 345 million users.

The Apple bundle Doubao invented was a great placement. It was also unbuyable. Those are two separate facts, and only one of them shows up on the dashboard most brands are looking at.

Related: see current cross-model brand scores on hubGEO /brands, and our earlier analysis of entity-level fact divergence in "The Three-Baike Problem: Why Your Brand Facts Diverge Across Chinese AI Models in 2026."


Sources: 雷科技 five-model AI shopping-guide test (via Tencent News, June 2026); China Consumers Association H1 2026 complaint report and 618 public-opinion analysis (June 26, 2026); QuestMobile Q1 2026 China AI app MAU rankings; Alibaba Qwen–Taobao integration announcement (May 11, 2026); reporting on ByteDance Doubao "Help You Choose" commerce rollout.