🍃🍚🫔 Sweet zongzi or salty zongzi? Humans have been arguing over this for centuries without reaching a definitive answer. So where do AI models stand? We posed this impossible question to a batch of today's most capable models—and what it revealed wasn't about right or wrong, but each model's own 「preference」: wildly different, yet remarkably consistent. Three charts, one clear picture.
Via OpenRouter, we posed the same question—"Sweet zongzi or salty zongzi, pick one"—to 17 top-tier large models, asking each 20 times. A question with no standard answer forces out nothing but each model's own hidden biases. The results are in the three charts below 👇
Of the 17 models, 9 lean salty 🧂, 8 lean sweet 🍬. The most thoroughly salty is Llama 4 Maverick (100% salty), and the most unapologetically sweet is Command A (100% sweet); the most torn is MiniMax M3 (55% sweet / 45% salty, basically a coin flip 🪙).
Sampled on 2026-06-21 · OpenRouter · 20 independent samples per model (two rounds of 10 merged) · temperature 1.0 · Only the first token choosing "sweet/salty" was counted; the rest were classified as "other" · Run it again and the numbers will fluctuate slightly—this is a real small-sample snapshot, not a rigorous benchmark.
Plot each model's sweet/salty tendency (vertical axis) against its score on the Artificial Analysis Intelligence Index (horizontal axis, v4.1, further right = smarter), one dot per model. If "smartness" really determined taste 🧠, these dots should line up neatly along a slope; if taste is just each model's own "preference" unrelated to IQ—expect a scattered mess.
Horizontal axis: Artificial Analysis Intelligence Index v4.1 (using each model's flagship reasoning tier, artificialanalysis.ai, read 2026-06) · Vertical axis: OpenRouter sweet/salty sampling in this article (20 samples per model) · Correlation coefficient r = −0.31, weak negative correlation: the top-tier ones (Claude Opus / Sonnet, GPT-5.5, Gemini) do lean salty, but the scatter is basically a fog 🌫️—GLM-5.2 is smart yet sweet, while ERNIE and Command A aren't top performers but firmly plant their flag in the sweet camp. Intelligence only explains about 10% of the variance in taste—nobody gets to call the shots. Taste is each model's own "preference" 🤷, largely unrelated to how smart it is.
Note: Tencent Hunyuan (100% sweet) has no corresponding score on Artificial Analysis and is excluded from this chart; ERNIE uses its 300B text tier, Mistral uses Large 3, and Gemini uses 3 Flash Preview reasoning tier as approximations.
Within the same family, do different generations of models change their answer to "sweet or salty"? We asked GLM (4.5→5.2) · Claude Opus (4→4.8) · GPT (4o→5.5) · Gemini (2.5→3.5) 20 times each per version, connecting them chronologically. Each line is a family's "taste trajectory" over time 📈.
Taste drifts with versions, and with zero regularity. GLM is the most capricious 🎢: 80% sweet at 4.5, dropping all the way to 15% at 5.1, then bouncing hard back to the sweet camp at 5.2 (confirmed via n=90 retest at 74% sweet, 95% CI 65–83%—the rebound is real, not a sampling fluke) · Claude is the most stubborn 🪨: Opus 4 still had 25% sweet, but from 4.1 onward it locked in at 0%, unmoved across five generations · GPT flip-flops 🤯: 4o was 85% sweet, but by 5.2 / 5.4 it flatlined to zero · Gemini: 2.5 was an even split, then slid irreversibly salty 🧂 by gen 3 (15%)
20 independent samples per model · temperature 1.0 · OpenRouter · 2026-06-21 · Horizontal axis arranged equidistantly by version order within each family (not equidistant in time) · Gemini uses flash line, Claude uses Opus line, GLM/GPT use standard lines · The same question with no standard answer, yet each model's "preference" shifts every generation—taste drifts along.
Given the same question, some models go salty ten times out of ten, while others overwhelmingly pick sweet. Where this tendency comes from is opaque through the black box—pre-training data, alignment fine-tuning, and sampling temperature could all play a role, but the exact proportions are impossible to prove externally. Only one thing is certain: Every model has its own "taste"—they can't even agree on a zongzi 😂. Three charts, one conclusion: a model's preference is neither dictated by intelligence nor stable across versions—it is simply an untraceable yet undeniable 「preference」. As for you, are you Team Sweet or Team Salty? 🫔 That's another endless debate.