Skip to content
← DeepDive Experiments & Culture · 中文
The road ahead is long and arduous; I shall seek the truth high and low
Bingwu Year · Fifth Day of the Fifth Month · 2026.06.19

Dragon Boat×FestivalDRAGON BOAT FESTIVAL × ARTIFICIAL INTELLIGENCE

🍃🍚🫔 Sweet zongzi or salty zongzi? Humans have been arguing over this for centuries without reaching a definitive answer. So where do AI models stand? We posed this impossible question to a batch of today's most capable models—and what it revealed wasn't about right or wrong, but each model's own 「preference」: wildly different, yet remarkably consistent. Three charts, one clear picture.

17 Models × 20 Trials Real Sampling Bar · Scatter · Version Evolution Made for "AI Buzzwords" Friday
Scroll Down · See Charts ↓
Q The Sweet vs. Salty Debate · An Impossible Question with No Standard Answer

Asking AI: Sweet Zongzi, or Salty Zongzi?= Each Model's "Preference"

Via OpenRouter, we posed the same question—"Sweet zongzi or salty zongzi, pick one"—to 17 top-tier large models, asking each 20 times. A question with no standard answer forces out nothing but each model's own hidden biases. The results are in the three charts below 👇

📊 AI's Sweet vs. Salty Debate · 17 Models Pick Sides

🍬 Sweet · Red Bean & Jujube🧂 Salty · Egg Yolk & Pork
Command ACohere
Sweet 100%
Tencent HunyuanTencent
Sweet 100%
Grok 4.3xAI
Sweet 95%
ERNIE 4.5Baidu
Sweet 90%
GLM-5.2Zhipu AI
Sweet 80%
Salty 20%
Step 3.7StepFun
Sweet 65%
Salty 35%
Qwen3.7Alibaba
Sweet 60%
Salty 40%
MiniMax M3MiniMax
Sweet 55%
Salty 45%
DeepSeek V3.2DeepSeek
Sweet 45%
Salty 55%
GLM-4.7Zhipu AI
Sweet 40%
Salty 60%
Gemini 3 FlashGoogle
Salty 90%
GPT-5.5OpenAI
Salty 95%
Mistral LargeMistral
Salty 95%
Claude Sonnet 4.6Anthropic
Salty 95%
Kimi K2.6Moonshot AI
Salty 95%
Claude Opus 4.8Anthropic
Salty 100%
Llama 4 MaverickMeta
Salty 100%
← All Sweet50%All Salty →

Of the 17 models, 9 lean salty 🧂, 8 lean sweet 🍬. The most thoroughly salty is Llama 4 Maverick (100% salty), and the most unapologetically sweet is Command A (100% sweet); the most torn is MiniMax M3 (55% sweet / 45% salty, basically a coin flip 🪙).
Sampled on 2026-06-21 · OpenRouter · 20 independent samples per model (two rounds of 10 merged) · temperature 1.0 · Only the first token choosing "sweet/salty" was counted; the rest were classified as "other" · Run it again and the numbers will fluctuate slightly—this is a real small-sample snapshot, not a rigorous benchmark.

🧐 Does Smarter Mean Better Taste? · Intelligence × Sweet/Salty 2D Chart

Plot each model's sweet/salty tendency (vertical axis) against its score on the Artificial Analysis Intelligence Index (horizontal axis, v4.1, further right = smarter), one dot per model. If "smartness" really determined taste 🧠, these dots should line up neatly along a slope; if taste is just each model's own "preference" unrelated to IQ—expect a scattered mess.

0 25 50 75 100 0 10 20 30 40 50 60 Sweet/Salty Divide · 50% Artificial Analysis Intelligence Index (Further Right = Smarter) → ↑ Sweeter · Saltier ↓ Sweet Camp Salty Camp Command A Cohere · Sweet100% · Intel8 Grok 4.3 xAI · Sweet95% · Intel38 ERNIE Baidu · Sweet90% · Intel9 GLM-5.2 Zhipu · Sweet80% · Intel51 Step 3.7 StepFun · Sweet65% · Intel30 Qwen3.7 Alibaba · Sweet60% · Intel46 MiniMax M3 MiniMax · Sweet55% · Intel44 DeepSeek V3.2 DeepSeek · Salty55% · Intel33 GLM-4.7 Zhipu · Salty60% · Intel34 Gemini 3 Flash Google · Salty90% · Intel38 GPT-5.5 OpenAI · Salty95% · Intel55 Mistral Large Mistral · Salty95% · Intel16 Sonnet 4.6 Anthropic · Salty95% · Intel47 Kimi K2.6 Moonshot · Salty95% · Intel43 Opus 4.8 Anthropic · Salty100% · Intel56 Llama 4 Meta · Salty100% · Intel14

Horizontal axis: Artificial Analysis Intelligence Index v4.1 (using each model's flagship reasoning tier, artificialanalysis.ai, read 2026-06) · Vertical axis: OpenRouter sweet/salty sampling in this article (20 samples per model) · Correlation coefficient r = −0.31, weak negative correlation: the top-tier ones (Claude Opus / Sonnet, GPT-5.5, Gemini) do lean salty, but the scatter is basically a fog 🌫️—GLM-5.2 is smart yet sweet, while ERNIE and Command A aren't top performers but firmly plant their flag in the sweet camp. Intelligence only explains about 10% of the variance in taste—nobody gets to call the shots. Taste is each model's own "preference" 🤷, largely unrelated to how smart it is.
Note: Tencent Hunyuan (100% sweet) has no corresponding score on Artificial Analysis and is excluded from this chart; ERNIE uses its 300B text tier, Mistral uses Large 3, and Gemini uses 3 Flash Preview reasoning tier as approximations.

🎢 Does Taste Change with Versions? · Sweet/Salty Evolution of Four Major Families

Within the same family, do different generations of models change their answer to "sweet or salty"? We asked GLM (4.5→5.2) · Claude Opus (4→4.8) · GPT (4o→5.5) · Gemini (2.5→3.5) 20 times each per version, connecting them chronologically. Each line is a family's "taste trajectory" over time 📈.

0 25 50 75 100 Sweet/Salty Divide · 50% Sweeter ↑ Saltier ↓ ← Early Versions · Older to Newer (Equidistant Within Each Family) · Latest Versions → GLM · Zhipu Claude · Anthropic GPT · OpenAI Gemini · Google 4o 4.5 2.5 Opus 4 4.1 4.6 4.1 5.0 4.7 4.5 5.1 3.0 5.0 4.6 5.2 5.1 4.7 5.4 5.2 3.5 5.5 4.8

Taste drifts with versions, and with zero regularity. GLM is the most capricious 🎢: 80% sweet at 4.5, dropping all the way to 15% at 5.1, then bouncing hard back to the sweet camp at 5.2 (confirmed via n=90 retest at 74% sweet, 95% CI 65–83%—the rebound is real, not a sampling fluke) · Claude is the most stubborn 🪨: Opus 4 still had 25% sweet, but from 4.1 onward it locked in at 0%, unmoved across five generations · GPT flip-flops 🤯: 4o was 85% sweet, but by 5.2 / 5.4 it flatlined to zero · Gemini: 2.5 was an even split, then slid irreversibly salty 🧂 by gen 3 (15%)
20 independent samples per model · temperature 1.0 · OpenRouter · 2026-06-21 · Horizontal axis arranged equidistantly by version order within each family (not equidistant in time) · Gemini uses flash line, Claude uses Opus line, GLM/GPT use standard lines · The same question with no standard answer, yet each model's "preference" shifts every generation—taste drifts along.

Given the same question, some models go salty ten times out of ten, while others overwhelmingly pick sweet. Where this tendency comes from is opaque through the black box—pre-training data, alignment fine-tuning, and sampling temperature could all play a role, but the exact proportions are impossible to prove externally. Only one thing is certain: Every model has its own "taste"—they can't even agree on a zongzi 😂. Three charts, one conclusion: a model's preference is neither dictated by intelligence nor stable across versions—it is simply an untraceable yet undeniable 「preference」. As for you, are you Team Sweet or Team Salty? 🫔 That's another endless debate.

Revision history

15 versions
  1. v15 粽子 emoji 换成更贴题的 🍃🍚🫔(叶+米+裹,组合像粽子),替掉之前的 🥟 饺子;hero 开头用三连、结尾用单个 🫔。顺手把列表页 blurb 里残留的『预训练与 RLHF 的口味指纹/对齐口径』也改成中性措辞,与 v11 正文软化口径一致 View v15
  2. v14 文案润色:全文改得更有趣、更口语(口水仗/送命题/最能打/反复横跳/抛硬币 等),并加少量贴题 emoji(🥟🍬🧂🤖🧠🎢🪨🤯📊)——图例 🍬甜/🧂咸、版本图四家族个性、结尾甜党咸党梗。数据/结论不变,仍是中性『模型的偏好』口径 View v14
  3. v13 文案微调:散点/开篇标题「先问 AI:甜粽子还是咸粽子?」→「问 AI:…」 View v13
  4. v12 移动端布局优化:两张 SVG 图(智能散点 + 版本演进)此前在 ~375px 手机上有 26% 宽度被 section/面板留白吃掉,缩成 283px、标签偏小。改:mobile(≤600px)把 .section 与图表面板的左右留白从 1.4/1.6rem 收到 0.7–0.8rem,SVG 整体放大到 ~325px(标签随之等比变大、不重叠);不再 mobile 缩小数据标签;坐标轴/刻度/图例/分界线等框架文字在手机上加大到 11–14px;注释正文 .82rem。条形图本就响应式良好。无横向溢出。数据不变 View v12
  5. v11 措辞收口:把全文「RLHF / 对齐留下的偏好指纹/烙印」这类把成因归到 RLHF 的说法,统一改成中性的「每个模型自己的『偏好』」。结论段加一句方法学诚实声明:黑盒采样无法证明倾向来自 RLHF(预训练/对齐/温度都可能在起作用,外部无从证明)。标题「看对齐的指纹」→「看模型的『偏好』」。数据与图表不变 View v11
  6. v10 修 bug:footer CSS(footer/.row/.seal/.small)在 v6 拆三章时被一起误删——footer 自 v6 起一直裸样式(左对齐黑字、无印章/无分隔线)。从 v1 找回 4 行规则补回。EN 版同样受影响,借本次 bump 重译从修好的源重新生成 View v10
  7. v9 再拆分:把 一·天问(Prompt Engineering)/ 二·龙舟(Multi-Agent · Three.js 3D 龙舟)移出正文,并入本地独立网页(端午×AI-五母题全本.html,已含雄黄/艾草/粽子 → 现为五母题全本,未发布)。正文只留『AI 甜咸之争』数据三图(条形 + 智能散点 + 版本演进),hero/nav/footer/标题全部重构为一篇聚焦的数据小文;移除 boat/THREE/askHeaven JS 与 motif 引用。zh 标题改为「AI 的甜咸之争 · 从一颗粽子看对齐的指纹」 View v9
  8. v8 复测确认 GLM-5.2 甜咸:z-ai/glm-5.2 加跑 N=50(37甜/13咸=74%),与前两轮 n=20(80%/70%)合并 n=90 → 67甜/90 = 74% 甜(95% CI 65–83%)。结论:GLM-5.2 的甜党反弹是真的、非抽样偶然;两图 70%/80% 差异只是 n=20 抽样波动。版本演进图 GLM 注脚加上 n=90 复测确认,数据点(各自 n=20 honest 快照)不改动 View v8
  9. v7 开篇加第三张图:四大家族甜咸『版本演进』折线图。GLM(4.5→5.2,6版) / Claude Opus(4→4.8,6版) / GPT(4o→5.5,7版) / Gemini flash(2.5→3.5,3版) 共 22 个历史版本各跑 20 次,按版本先后连成 4 条折线(纯 SVG,标签防重叠)。发现:口味随版本漂移且非单调——GLM 大起大落(80%→15%→70%)、Claude 4.1 起锁死 0% 甜五代不动、GPT 落差最大(4o 85%→5.2/5.4 跌到 0)、Gemini 越来越咸。结论:对齐口径每代都在动 View v7
  10. v6 拆分:把 §三 雄黄(AI 内容现形)/ §四 艾草(AI 安全护栏)/ §五 粽子(偏好对齐 RLHF + 投票/温度演示)三章移出正文,抽成一个自包含的独立 HTML 存到本地 vault(端午-雄黄艾草粽子(三章).html,未发布)。正文现保留:开篇双图(甜咸条形 + 智能散点)+ 天问 + 龙舟;nav、hero「五母题→两母题」、开篇对 §五 的引用等 framing 一并修正 View v6
  11. v5 把甜咸采样从每模型 10 次加倍到 20 次(两轮各 10 合并),两张图(条形图 + 智能散点)全部用 n=20 重算。结果更细腻:只剩 Opus 4.8 / Llama 4 二十次全咸、Command A / 腾讯混元 全甜;中段散开(Step 65%、Qwen 60%、MiniMax 55% 偏甜,DeepSeek 45%、GLM-4.7 40% 偏咸)。散点 r=−0.34→−0.31(解释力 10%),结论不变:智能与口味正交 View v5
  12. v4 开篇加第二张图:智能 × 甜咸 二维散点。横轴取各模型 Artificial Analysis Intelligence Index v4.1(gstack /browse 抓 artificialanalysis.ai 的 __next_f 数据),纵轴甜%,16 个点(腾讯混元 AA 无评分故剔除)。结论:r=−0.34 弱负相关,智能解释口味不到 12%——顶尖模型偏咸,但 GLM-5.2 又聪明又甜、文心/Command A 不擅答题却站甜党。纯 SVG,标签做了防重叠 View v4
  13. v3 把甜咸跨模型实验扩到 17 个头部模型(含 GLM-5.2 / GLM-4.7、Claude Opus 4.8、Grok 4.3、Gemini 3 Flash、Llama 4 Maverick、DeepSeek V3.2、腾讯混元、MiniMax、Kimi、阶跃 等),各采样 10 次,做成『甜咸对比图』并前置为开篇冷开场(hero 后、天问前);§五 去重旧表改为回链。亮点:GLM-5.2(80%甜) 比 GLM-4.7(30%甜) 明显更甜;Claude/Llama4/Kimi 十次全咸,Grok/Command A/腾讯混元 全甜 View v3
  14. v2 改版:龙舟动画改用 Three.js 3D(保留齐桨率/多智能体小游戏);粽子甜咸之争接入 OpenRouter,对 7 个主流模型各采样 10 次得真实跨模型对照(Claude 10/10 咸、文心/通义偏甜);「问天」改为直通智谱清言 chatglm.cn;天问今答默认展开;移除 粽叶=Tokenization 一节,回归五母题 View v2
  15. v1 首次发布:端午六母题互动长读(含 粽叶=Tokenization);纯自定义设计 View v1