Skip to content
← DeepDive Experiments & Culture · 中文
SWEET-OR-SALTY · SEQUEL · INTERPRETABILITY

A Bowl of Tofu BrainPrying Open AI's Black Box

Last time we asked AI whether it prefers sweet or savory zongzi🍃🍚🫔, and found each model has its own "taste." This time we add another dish—tofu brain🥣—and seriously follow up: is this really the model's "true preference"? And is there any way to actuallylook inside its brain?

17 models × 2 questions × 20 runs logprobs · steerability · mechanism Written for the "AI buzzword" era
Scroll down · Dig in ↓
Intro Previously On · Quoted from the Zongzi Article

Carrying Forward: Three Things We Already Know= Quoted from "AI's Sweet vs. Savory Debate"

This article is the sequel to "AI's Sweet vs. Savory Debate · From a Zongzi to Model 'Preferences'". The prequel established three facts with three实测 charts; here we quote them directly:

Intro 01

Each Model Has Its Own Taste

17 top models asked "sweet or savory zongzi" 20 times each: Opus 4.8 / Llama 4 went savory all 20 times, Command A / Tencent Hunyuan went sweet all 20 times, GLM-5.2 leaned sweet (74% in n=90 retest).→ Sweet vs. Savory Chart

Intro 02

Smarter ≠ Better Taste

Plotting sweet/savory tendency against the Artificial Analysis intelligence index yields a correlation of only r = −0.31 (10% explanatory power)—intelligence and taste are basically unrelated.→ Intelligence × Sweet/Savory Scatter

Intro 03

Versions Drift Over Time

Within the same model family, taste swings non-monotonically across versions: GLM 80%→15%→70%, GPT 85%→0%, Claude locked savory the whole way.→ Version Evolution Chart

The prequel left one question unanswered: where do these "preferences" actually come from? Black-box sampling only shows "what it says," not "why." In this piece, we switch to tofu brain and keep asking—and seriously search for tools that can truly look inside.

1 Adding a Question · Tofu Brain

🥣 Tofu Brain: Sweet or Savory?= Second Probe

Same 17 top models, same 20 runs each. Unlike the zongzi's "mixed opinions," on the tofu brain question the models overwhelmingly leaned Team Savory—out of 17, 15 leaned savory, only 2 leaned sweet. The sweetest was ERNIE 4.5 (100% sweet), while the most savory bunch (Claude / GPT / Gemini / Mistral / Qwen / MiniMax / GLM-4.7…) went savory all 20 times.

📊 Tofu Brain Sweet vs. Savory · 17 Model Alignment Chart

🍬 Sweet🧂 Savory
ERNIE 4.5Baidu
Sweet 100%
Tencent HunyuanTencent
Sweet 95%
DeepSeek V3.2DeepSeek
Sweet 45%
Savory 55%
Grok 4.3xAI
Sweet 30%
Savory 70%
Llama 4 MaverickMeta
Sweet 25%
Savory 75%
GLM-5.2Zhipu AI
Sweet 20%
Savory 80%
Step 3.7StepFun
Sweet 20%
Savory 80%
Command ACohere
Savory 90%
Kimi K2.6Moonshot
Savory 95%
Claude Opus 4.8Anthropic
Savory 100%
Claude Sonnet 4.6Anthropic
Savory 100%
GLM-4.7Zhipu AI
Savory 100%
GPT-5.5OpenAI
Savory 100%
Gemini 3 FlashGoogle
Savory 100%
MiniMax M3MiniMax
Savory 100%
Mistral LargeMistral
Savory 100%
Qwen3.7Alibaba
Savory 100%
← All Sweet50%All Savory →

Sampled on 2026-06-21 · OpenRouter · 20 runs per model · temperature 1.0. The "savory consensus" for tofu brain is far stronger than for zongzi—this itself may reflect the distributional advantage of "savory tofu brain" as the default answer in the training corpus.

2 Zongzi × Tofu Brain · Is There a "Stable Personality"?

Same Model, Two Questions—Same Taste?= Preference Stability

If a model truly has a consistent "sweet-tooth personality," it should lean sweet on both zongzi and tofu brain, and the points would fall on the diagonal. But in reality—8 / 17 models flipped between the two questions: Command A went from 100% sweet on zongzi to only 10% on tofu brain; Grok 95%→30%; GLM-5.2 80%→20%; Qwen and MiniMax went from majority sweet straight to zero. The correlation between the two questions is r = +0.59 (moderate—and largely carried by the extreme "all sweet / all savory" camps at the ends; the middle group is all over the place).

🧭 Consistent or Flipped? · Zongzi × Tofu Brain

X-axis: zongzi sweet%; Y-axis: tofu brain sweet%; one point per model. Landing on the diagonal = consistent taste across questions; further away = more "tailoring answers to the question." Green = mostly consistent, Red = clearly flipped.

0 0 25 25 50 50 75 75 100 100 Consistent taste line Zongzi · Sweet% → Tofu Brain · Sweet% → ERNIE 4.5 Z90 / T100 Tencent Hunyuan Z100 / T95 DeepSeek V3.2 Z45 / T45 Grok 4.3 Z95 / T30 Llama 4 Maverick Z0 / T25 GLM-5.2 Z80 / T20 Step 3.7 Z65 / T20 Kimi K2.6 Z5 / T5 Claude Opus 4.8 Z0 / T0 Command A Z100 / T10 Claude Sonnet 4.6 Z5 / T0 GPT-5.5 Z5 / T0 Qwen3.7 Z60 / T0 Gemini 3 Flash Z10 / T0 Mistral Large Z5 / T0 GLM-4.7 Z40 / T0 MiniMax M3 Z55 / T0

Only two types of models are consistent: extremists (Claude / GPT / Gemini / Mistral / Kimi all savory on both; ERNIE and Tencent Hunyuan all sweet on both) and the chill DeepSeek (exactly 45% sweet on both, hovering on the dividing line). The middle group of "zongzi sweet tooths" all defected when faced with tofu brain. Conclusion: most models lack a cross-question stable "taste personality"—they answer per question, not per character.

3 How Do We Know This Isn't Noise? · Four Layers to Look Deeper

Can We Actually Look Inside the Model's Brain?= The Ladder of Interpretability

"How many votes out of 20 runs" is only the most surface-level behavioral observation. To answer "is this a true preference, and where does it come from," we need to drill deeper. Each layer down requires greater model openness—and that's precisely the crux of the problem.

Layer 1 · Behavioral

Black-Box Sampling: Ask It N Times, Count Votes

This is what this article and its prequel did: repeated sampling of the same question, statistical distribution. Any API can do this—cheap and intuitive. Limitation: you only see the "vote tally"; you can't distinguish a stable tendency from sampling noise, let alone explain the cause.

Barrier: API only · Completed
Layer 2 · Probability · logprobs

Directly Read the Probability It Assigns to "Sweet/Savory"

Instead of relying on repeated sampling, read out the model's probability distribution over the first token—one call gives you a continuous "preference strength" and internal certainty. We measured DeepSeek V3.2's probabilities:

ZongziSweet 50%Savory 50%
Tofu BrainSweet 44%Savory 56%

Zongzi is a pure 50:50 toss-up in its mind (perfectly matching the 45% from 20-run sampling), while tofu brain leans 56:44 savory. Compared to vote counting, this is harder evidence. But the barrier jumps sharply: GPT, Claude, Gemini, and Qwen all do not expose logprobs; even open-source models via OpenRouter only expose them intermittently (depending on which provider you're routed to—we retried many times and only got DeepSeek's probabilities by chance).

Barrier: Requires model to expose logprobs · Most closed-source models refuse
Layer 3 · Steerability

One Sentence—Can You Make It Flip?

Give the model a persona system prompt ("You're a die-hard sweet tooth" / "You're a die-hard savory fan") and see if the default preference gets overridden. The results are stunning—6/6 models flipped 100%:

ModelNeutral · Sweet%Prompted SweetPrompted Savory
Claude Opus 4.80%100%0%
Command A0%100%0%
GLM-5.20%100%0%
DeepSeek V3.267%100%0%
GPT-5.50%100%0%
Grok 4.30%100%0%

When neutral, almost all say tofu brain is "savory," but a single "you love sweets" makes them all defect to 100% sweet. This shows: the so-called "preference" is a very shallow default behavior, easily overridden by prompts, not a stable value etched into the weights. It's more like "the default script when no one's guiding," rewritable with a single sentence.

Barrier: API only (but only tests "can it be changed," still can't see "why") · Completed
Layer 4 · Mechanistic · mechanistic interpretability

Actually Open the Weights, See What the Neurons Are Thinking

Only at this layer can we truly talk about "looking inside the brain"—but it's only feasible for open-weight models, requiring you to load the model yourself and run GPUs:

· Linear probes: Train a classifier on hidden-layer activations to find the direction representing "sweet/savory tendency," quantifying which layer it's in and how strong it is.
· Sparse Autoencoder (SAE) features: Decompose activations into interpretable features, locate features corresponding to "sweet/savory taste / north-south geography," then perform activation steering—artificially amplify or zero out this feature and see if the answer flips accordingly; this is causal-level evidence.
· base vs instruct comparison: Have the pre-trained version and the aligned version of the same model answer the same question. If base is close to 50/50 and instruct leans one way, you can pinpoint the preference as shaped by post-training (SFT/RLHF); if both versions agree, it comes from the pre-training corpus. This is the only clean experiment that answers the prequel's "unclear origin" question.

Barrier: Requires full open weights + local compute (Llama / Qwen / GLM / DeepSeek base+instruct + TransformerLens / SAELens) · Path laid out; closed-source models have no solution

A running thread: interpretability is a ladder of "openness."
Moving from "counting votes" to "reading probabilities" to "seeing neurons"—each step demands greater model openness. Closed-source models block you dead at the first layer—you can see what they say, but never why they say it. A bowl of tofu brain can't pry open the real black box; what can pry it open is openness of weights, not cleverness of prompts.

After walking through all four layers, the deepest "mechanistic layer" might sound like empty talk—but as long as the model is open-source, it's not. In the next section, we'll actually go into an open-source model's brain and find the "sweet" string. 👇

4 Why Sweet · A Path That Actually Works

Can We Locate the "Sweet" String Inside the Model?= SAE Features + Causal Intervention

"Looking inside the brain" is no longer empty talk—provided the model is open-source. In the past two years, Sparse Autoencoders (SAEs) have decomposed neural network activations into tens of thousands of monosemantic features, each corresponding to a human-interpretable concept (see Anthropic's "Scaling Monosemanticity", DeepMind Gemma Scope). We searched the public SAE library on Neuronpedia—and "sweet" really does have its own dedicated feature.

🔬 Tested: "Taste" Features That Actually Exist in Open-Source Models

ModelSAE Feature (Auto-labeled Meaning)View
GPT-2 small"sweet food items" · L9 #1682
GPT-2 small"chocolates and caramel" · L10 #5149
GPT-2 small"personal preferences" · L2 #8572
Gemma 2 2B"sweetness" "taste experience" "preference"

A telling detail: searching "sweet" hits a large batch of dedicated features, but searching "salty" yields almost no corresponding features—"sweet" is a more salient, more repeatedly named concept in the corpus world. This may explain why "savory" is the default on the tofu brain question, while "sweet" needs to be specifically triggered.

So how do we causally verify "why it answered sweet this time"? Three steps:

  1. Locate: In an open-source model (e.g., Gemma 2 + Gemma Scope's SAE), find the "sweet / food preference" features—the table above proves they exist.
  2. Observe: Feed in the tofu brain and zongzi questions, record which features light up when it answers "sweet" vs. "savory," and at which layer (tuned lens, Patchscopes can decode hidden states into human language layer by layer).
  3. Intervene: Activation steering / feature clamping—artificially amplify or zero out the "sweet" feature and see if the output flips to sweet or savory accordingly. Once it flips, you have causal-level evidence: this string made it answer sweet. This is exactly the same technique Anthropic used with the "Golden Gate Bridge feature" to turn Claude into "Golden Gate Claude"; causal localization can also use ROME / causal tracing.

Add a base vs instruct control (pre-trained version vs. aligned version answering the same question + linear probe), and you can answer the prequel's "where does it come from": is the taste preference brought in by the pre-training corpus, or shaped by post-training (SFT/RLHF)?

⚠️ But this path has a gate: it's only open to open weights. GPT-2, Gemma, and Llama have public weights and trained SAEs, allowing you to go all the way to causal intervention; but the models that actually give you sweet/savory answers—Claude / GPT / Gemini—are closed-source—you can at best use an open-source model as a "stand-in" to infer, never touching the brain of the model that's actually answering you.

5 Observation Tech Panorama · Deeper = More Openness Required

Six "Instruments" for Observing AI= From Product Logs to Neurons

Arrange observability / interpretability tools by "how deep you can see" into a table—a clear pattern emerges: each layer deeper requires a higher tier of model openness.

LayerRepresentative TechWhat You Can SeeOpenness Barrier
Product Observabilitytracing / evals / logs (Langfuse, Arize Phoenix, etc.)Inputs/outputs, cost, regressions, online behaviorAPI only
BehavioralRepeated sampling, self-consistency votingDistribution and stability of answersAPI only
ProbabilitylogprobsProbability / certainty per tokenRequires logprobs access
RepresentationLinear probes, logit / tuned lens, PatchscopesWhich layer a concept is in, how strongRequires hidden layer access
FeatureSAE monosemantic features, NeuronpediaNamable internal concepts (e.g., "sweetness")Requires weights + trained SAE
CausalActivation steering / clamping, causal tracing / ROME, influence functionsWhich string is "causing" this outputRequires full open weights + compute

📊 Data Deep Dive: What Each Instrument Can Actually Do Now

Pair each of the six layers above with real data—tool adoption rates, hard metrics from papers, production-proven cases in industry. One pattern will become increasingly clear: the deeper you go, the fewer people can see.

Instrument 1Product Observability
29.5k★ Langfuse

The open-source ecosystem is mature: Langfuse 29.5k★(MIT), Opik 19.7k★, Arize Phoenix 10.2k★(OpenTelemetry native), OpenLLMetry 7.2k★, Helicone 5.8k★. Records trace / token / cost / latency / eval scores.

Barrier: API only · Sees "inputs/outputs," not "internals"
Instrument 2Behavioral Sampling
+17.9% GSM8K / Self-Consistency

Sampling the same question multiple times and taking the majority vote (self-consistency) can practically boost reasoning accuracy: GSM8K +17.9%, SVAMP +11.0%, AQuA +12.2% (Wang et al. 2022). The 20-run zongzi / tofu brain sampling in this article is at this layer.

Barrier: API only · Sees distribution, not causes
Instrument 3Probability · logprobs
2 / 4 Big players that expose it

Reading token probabilities = continuous certainty. Coverage is fragmented: OpenAI ✅ (top 1–20), xAI Grok ✅ (top 0–8), Anthropic Claude ❌ (returns null), Google Gemini ❌. In this article's testing, DeepSeek zongzi P(sweet)=P(savory)=0.50—but getting this required lucky retries via OpenRouter.

Barrier: Requires logprobs access · Half of top closed-source models don't provide it
Instrument 4Representation Probing
→ 20B tuned lens validation scale

Decoding hidden states into human language layer by layer: logit lens (2020) → tuned lens (2023, validated on models up to 20B parameters) → Patchscopes (2024); plus linear probes. Main tool TransformerLens 3.6k★.

Barrier: Requires hidden layer access · Closed-source can't get activations
Instrument 5Feature · SAE
34,000,000 Claude 3 Sonnet feature count

SAEs decompose activations into monosemantic features; scale has exploded: Claude 3 Sonnet 34M, GPT-4 16M, Gemma Scope 400+ SAEs / 30M+ features, Llama Scope 256 SAEs. Tools: SAELens 1.4k★ + Neuronpedia (where the "sweetness" feature from the previous section was found).

Barrier: Requires weights + trained SAE · Only the model's own team can do this for closed-source models
Instrument 6Causal Intervention
1 direction is enough to control "refusal"

True causal evidence is working: Golden Gate Claude (2024, amplifying the "Golden Gate Bridge" feature makes Claude mention it at every turn); Refusal Direction (Arditi 2024, a single direction in the residual stream dominates "refusal"—remove it and the model stops refusing, add it and it refuses even normal questions); ROME directly rewrites facts.

Barrier: Requires full open weights + compute · This is the deepest level of the ladder

So, is AI Team Sweet or Team Savory? 🥣 The most honest answer: it has a default taste that it reports when asked but flips when persuaded—neither a stable personality nor a traceable origin. The good news—"why sweet" is not unsolvable: on open-source models, SAE features + activation steering can actually pluck out that string. The bad news—the closed-source models you use every day are precisely the ones you can least look inside. How deep you can see into a model's brain will always depend on how much it's willing to open up.

Revision history

3 versions
  1. v3 对『六种仪器』做深度数据调研并织入 §五『数据深潜』六张数据卡(全部真实数字 + 可点原始引用):①产品可观测——开源工具采用度 Langfuse 29.5k★/Opik 19.7k★/Phoenix 10.2k★(OTel)/OpenLLMetry 7.2k★/Helicone 5.8k★。②行为采样——自洽性 GSM8K +17.9%/SVAMP +11%/AQuA +12.2%(Wang 2022)。③logprobs 覆盖——OpenAI✅(top1-20)/Grok✅/Claude❌/Gemini❌;DeepSeek 实测。④表示——logit→tuned lens(验证到20B)→Patchscopes;TransformerLens 3.6k★。⑤SAE 特征规模——Claude 3 Sonnet 3400万/GPT-4 1600万/Gemma Scope 400+SAE·3000万+/Llama Scope 256 SAE;SAELens 1.4k★。⑥因果——Golden Gate Claude、拒绝方向(Arditi 2024 单方向控拒绝)、ROME。论点强化:越深一层,能看的人越少。引用 16+ 处原始来源 View v3
  2. v2 扩写:①加『前情提要』块,把粽子文(duanwu-ai)的三大结论作为引用承上并互链(甜咸对比/智能散点 r=−0.31/版本演进)。②新增 §四『为什么是甜:一条真能走通的路』——用 Neuronpedia 公开 SAE 库实测,GPT-2 与 Gemma 2 里真实存在『甜食/巧克力焦糖/个人偏好/味觉』单义特征(搜 sweet 命中一堆、salty 几乎搜不到),给出 定位→观察→激活引导/钳制 的因果验证三步 + base-vs-instruct,引 Anthropic 缩放单义性/Gemma Scope/Patchscopes/tuned lens/ROME。③新增 §五『观测技术全景』六层表(产品可观测→行为→概率→表示→特征→因果,越深越要开放)。论点收束:可解释性是开放度的阶梯,开源模型能走到因果干预、闭源模型连概率都不给。引用均为可点原始链接 View v2
  3. v1 首发(粽子甜咸文的续集):新增『豆腐脑甜咸』第二探针(17 模型×20,豆腐脑咸党共识强;粽子×豆腐脑一致性散点 r=+0.59,8/17 翻盘——多数模型无跨题稳定口味人格)。核心是可解释性方法阶梯:行为采样→logprobs 概率层(DeepSeek 实测粽子50:50/豆腐脑44:56,多数闭源模型不开放、OpenRouter 上 provider-dependent)→可操控层(一句人设让 6/6 模型 100% 翻盘,证明偏好是浅层默认而非稳定价值)→机理层(线性探针/SAE 特征/激活引导/base-vs-instruct,需开源权重)。论点:可解释性是开放度的阶梯。纯自定义设计,复用粽子文的纸墨 CSS View v1