GEO Blog

DeepSeek and Qwen Agree on Luxury Brands, Disagree on Sources: What the 2026 Divergence Data Means for GEO

2026/8/6 上午1:02:15

New data: DeepSeek and Qwen pick the same luxury brands 70% of the time, but cite matching sources only 27% of the time.

DeepSeek and Qwen Agree on Luxury Brands, Disagree on Sources: What the 2026 Divergence Data Means for GEO

New research testing 1,620 brand-recommendation responses across DeepSeek and Qwen found something that should reshape how brand teams think about China AI visibility measurement: the two engines recommend the same luxury brands about 70% of the time, but they cite matching third-party sources behind those recommendations only 27% of the time. A brand audit run on DeepSeek tells you almost nothing about which content is actually driving your Qwen visibility — because the two engines are, in effect, reading different libraries to reach similar conclusions.

That gap — roughly 3x wider on the source side than the brand side — is the core finding from a Q2 2026 study by Eastbound Research, which ran 45 natural Chinese consumer prompts per category (watches, luggage, handbags) six times each on DeepSeek and Qwen, then normalized which websites each engine self-attributed as the basis for its answer. The result is one of the more granular public datasets on how Chinese AI engines actually source their brand recommendations, and it has direct implications for anyone running a GEO program across more than one model.

The gap between "who gets recommended" and "why"

Brand-level and source-level visibility are not the same measurement, and treating them as interchangeable is the mistake the data flags most clearly.

CategoryBrand overlap (DeepSeek vs Qwen)Source overlapDivergence ratio
Watches77%36%2.1x
Luggage67%25%2.7x
Handbags67%20%3.4x

The pattern holds across all three categories and gets more pronounced as the category becomes more lifestyle-driven rather than technical. In plain terms: DeepSeek and Qwen tell a Chinese consumer largely the same shortlist of brands to consider, but they justify those recommendations by citing almost entirely different corners of the web. For a brand's GEO team, this means source-side interventions — which platforms to publish on, which communities to seed with content — cannot be planned from a single engine's citation pattern and assumed to transfer.

DeepSeek concentrates, Qwen distributes — two different citation shapes

The study also found the two engines pull from structurally different-sized citation pools. DeepSeek draws from 200–400 unique sources per category and its top 15 sources capture 60–80% of all citations — a concentrated pattern where ranking into a small set of dominant sources goes a long way. Qwen draws from a much wider pool, 480–940 unique sources per category, with its top 15 capturing only 40–65% of citations. Getting into Qwen's top ranks requires broader coverage across a longer tail, not just dominance in a handful of channels.

The platform mix differs sharply too. Bilibili appears in DeepSeek's top 5 sources in every category tested (35–54% citation share) but is nearly absent from Qwen (3–12%). Reddit citation share runs high on DeepSeek (61–66%) and mid-to-variable on Qwen (4–56%). Brand-owned websites are more likely to surface in Qwen's citations — Rimowa's official site appeared in 14% of Qwen luggage citations, Tumi's in 6% — while DeepSeek tends to suppress brand-owned sites whenever a vertical specialist publication exists for the category.

Where a single "top source" strategy breaks: the ultra-luxury cliff

Perhaps the sharpest finding concerns SMZDM (什么值得买), the Chinese deal-and-review aggregator that dominates aspirational-luxury citations — appearing as a top-3 source in 12 of 12 test cells at mention rates of 80–99% for watches, luggage, and handbags on both engines. A brand team that saw only the category-level average would reasonably conclude SMZDM is the one channel worth prioritizing.

But when the handbag panel was re-cut by price tier, that pattern reversed sharply at the ultra-luxury threshold (30,000+ RMB, the Birkin/Kelly/exotic-Chanel range):

Engine × tierSMZDM mention rateHigh-weight replacement sources
DeepSeek, aspirational (≤10,000 RMB)100%
DeepSeek, ultra (30,000+ RMB)33%The Purse Forum (38%), Vogue Business/WWD (62%), Sotheby's/Christie's (8%)
Qwen, aspirational (≤10,000 RMB)100%
Qwen, ultra (30,000+ RMB)71%Auction-house archives (17%), Vogue Business/WWD (21%)

At the DeepSeek ultra-luxury tier, one representative answer to a prompt about 100,000+ RMB bags explicitly downgraded SMZDM to "low weight, used only for second-hand market discussion" while elevating The Purse Forum to "high weight, used for rare-skin authentication discussion among global collectors." The engines are not just citing different sources by chance — they appear to be applying a self-consistent hierarchy of source credibility that shifts by price tier, and a GEO plan built around one tier's source map will misfire at another.

Category shape changes the lever that matters

The study's category comparisons add a second dimension worth noting for brand teams working across product lines. In the technical, heritage-driven watch category, a single Chinese vertical specialist — Watch Home (腕表之家) — was cited in nearly 100% of responses on both engines, and brand-owned sites essentially disappeared from the top-15 source list as a result. In luggage and handbags, where no comparable specialist dominates, brand-owned sites and resale/authentication platforms (The RealReal, Vestiaire Collective, Plum/红布林) entered the citation mix meaningfully. Xiaohongshu's weight also rises with how lifestyle-driven the category is — from 17–55% on watches to 75–96% on handbags — while Wikipedia's weight moves the opposite direction, mattering more for craft and heritage categories than for fashion-forward ones.

One more structural point worth flagging for brands running English-only global content: the study found 46–63% of all cited sources were Chinese-language platforms, even though the brands recommended were overwhelmingly Western names (Rolex, Hermès, Louis Vuitton, Rimowa). The output substrate is Chinese; the subject matter is global. A brand with no Chinese-language presence is effectively invisible to the layer of sources doing the actual recommendation work, regardless of how strong its English-language site or PR footprint is.

Why this matters more in a six-model market

Eastbound's study covers two engines, DeepSeek and Qwen, and explicitly deferred Doubao to a follow-up round. That's a meaningful caveat for brands operating in mainland China specifically, where Doubao is the largest consumer-facing AI surface by usage and pulls the citation mix further toward Bilibili, Dianping, and other Mainland-native platforms than either DeepSeek or Qwen do on their own. If two engines already diverge this much on sources while broadly agreeing on brands, a six-model tracking approach — DeepSeek, Kimi, Doubao, Qwen, Wenxin, and Hunyuan — should be expected to fragment the source picture even further, not converge on a single "correct" channel list. Brand teams that build a GEO content plan around whichever engine they happened to audit first are, on this evidence, optimizing for roughly a quarter of the actual source landscape.

The reliability of the underlying data matters here too. Eastbound re-ran each of its six test cells as two independent panels specifically to check whether the citation patterns were stable rather than noise — the source-side findings replicated with a mean Spearman correlation of 0.97 and a top-15 Jaccard overlap of 0.74 across runs, meeting the conventional threshold for a strong result in 11 of 12 sub-cells. That's a meaningfully higher bar than a single-pass audit, and it's part of why the ultra-luxury tier collapse in particular reads as a structural pattern in how these engines weight source credibility, rather than a one-off artifact of a specific prompt set.

Category-specific playbook, not a universal source list

Put together, the data argues against a single "top 5 sources for China AI" checklist that a brand applies uniformly. The right source mix depends on three variables stacking on top of each other: which engine, which price tier, and how lifestyle-driven the category is. A watch brand chasing DeepSeek visibility should prioritize Watch Home (腕表之家) coverage over its own site, because vertical specialists crowd out brand-owned pages entirely once one exists for a category. A handbag brand selling into the aspirational tier should prioritize SMZDM editorial relationships; the same brand's ultra-premium exotic-skin line should instead prioritize collector-forum authentication threads and trade press like Vogue Business and WWD, where SMZDM's citation weight has already collapsed by two-thirds or more. Running both plays under one generic "China AI SEO" strategy misallocates budget toward the wrong tier's source map.

Takeaway for Brand Marketers

Three practical implications follow directly from this data. First, measure source-side visibility separately per engine — a DeepSeek citation audit does not predict Qwen citation coverage, and the gap is roughly three times wider than the brand-recommendation gap, so budget for engine-specific source tracking rather than a single combined report. Second, segment any GEO content plan by price tier if the category spans aspirational and ultra-premium positioning — a brand publishing only SMZDM-style comparison content will win the mass-market layer and remain invisible at the tier where scarcity and provenance matter more, where collector forums, trade press, and auction-house documentation carry the citation weight instead. Third, treat Chinese-language publication as non-negotiable even for globally positioned brands, since the citation layer behind Chinese AI answers is predominantly Chinese-language regardless of the brand's origin market.

For brands tracking their own citation and score data across China's AI engines — including Doubao, which this particular study set aside for a follow-up round — hubGEO's brand visibility rankings break down score movement by model so teams can see where these source-level gaps are actually showing up for their category.

Related: hubGEO Brand Rankings