Proof of Thinking: How to Prove an Article Contains Human Thought

Core Position: We are solving the wrong problem with the wrong methods. All current technical routes for "verifying human writing"—AI detectors, keystroke monitoring, content provenance—verify human presence, not human thinking. Between the two lies a gorge far wider than we imagine. And when this gorge is finally seen clearly, a deeper question emerges: What is it we truly want to protect?

Date: 2026-05-19 · Category: [TRENDS] Cognitive Economy
Length: ~36,000 words · Est. reading 96 minutes

Executive Summary

Dimension Core Finding
Problem Misdefinition Current verification systems ask "who pressed the keys" rather than "who generated the thinking"—this is a hierarchical error
Academic Naming Wojtowicz & DeDeo (AAAI 2025) named this "Mental Proof": using observable actions to authenticate unobservable mental facts
Blockchain Analogy PoW resists Sybil attacks by "making forgery more expensive than honesty"; AI makes cognitive "computation" nearly free, causing Mental Proof mechanisms to collapse
Ceiling of Technical Means AI detector accuracy has hit a systemic bottleneck; keystroke monitoring only proves human presence, not thinking; C2PA provenance tells you the "tool chain" not the "cognitive chain"
Conclusion-First and Proof of Experience Human thinking is "conclusion-first" (Haidt/Kahneman)—intuition is a compressed hash of experience; LLM CoT is also "conclusion-first," but sourced from statistical weights rather than embodied experience
The Double Edge of Interpretability CoT does not equal interpretability (Oxford AIGI 2025); reasoning horizon decays after 70-85% of chain length; but attribution graphs are making AI's internal reasoning readable—ultimately AI's thinking may be more verifiable than human thinking
Reflexivity Crisis of Cognitive Atrophy Human Effective Context Span (ECS) dropped from ~16,000 tokens in 2004 to ~1,800 tokens in 2026; the delegation loop is accelerating the erosion of the very object being protected
The Unforgeable Core of Human Thinking Five features resist forgery in essence: Lived Error, Emotional Cost, Epistemic Boundary, Positional Cost, Temporal Incompressibility
Three Emerging Verification Dimensions Process Proof, Position Proof, Cost Proof—none are dependent on technology
The Deepest Structural Problem "Proof of Thinking" is ultimately a question of who gets to define "qualified thinking"—a power problem, not a technical one
Six Deeper Cracks Cognitive deskilling (hollowed mind), cognitive copyright (ownership of thinking styles), thought echo chambers (framework homogenization), ritualized thinking (Goodhart trap), machine judges (legitimacy recursion crisis), AI dark forest law (exposure means absorption)
Five Possible Methods Behavioral process authentication (Writer's Integrity / Cognitive Signatures), Zero-Knowledge Process Proof (ZK-PoP), Oral Defense (real-time cognitive verification), Cognitive Fingerprint (long-term style consistency), Positional Cost Proof (Skin in the Game)—each has its epistemological ceiling
Frontier and Limits of Neural Proof MIT Media Lab experiments prove EEG can distinguish brain connectivity differences in LLM-assisted vs. independent writing; "cognitive debt" is measurable at the neural level; but adversarial attacks, neurodiversity discrimination, privacy disaster, and surveillance paradox constitute four insurmountable obstacles

I. We Are Asking the Wrong Question

In May 2026, you open an article about AI. The writing is fluent, the structure clear, the arguments well-documented, the footnotes pointing to real research. You cannot tell whether it was generated by AI.

Your first reaction might be: Was this written by AI?

But this question is becoming increasingly difficult to answer, and increasingly unimportant.

A more worthwhile question is: How much genuine human thinking is in this article?

On the surface, these two questions seem similar. In reality, they point in entirely different directions.

"Was it written by AI" is a provenance question—attempting to trace the production tool of the text. "How much human thinking" is a contribution question—attempting to assess the depth of the cognitive subject's involvement.

This distinction is precisely the core crux of the entire "Proof of Thinking" issue.

1.1 The Content Flood Has Arrived

Gartner predicted in 2022 that by 2026, 90% of online content would be AI-generated. Reality proved more complex than the prediction—not 90%, but AI content has penetrated massively. Graphite's tracking study of 65,000 URLs found that AI-generated articles briefly surpassed human writing in November 2024, and the two have remained roughly equal since.

Meanwhile, consumer attitudes underwent a dramatic reversal: in 2023, 60% of consumers expressed preference for AI-created content; by 2026, this figure plummeted to 26%. Digiday's report described an emerging trend: creators' "authenticity" and "messiness" are beginning to command market premiums.

But there is a hidden paradox here: The market is starting to pay a premium for "human-ness," while AI tools capable of perfectly simulating "human-ness" are evolving in parallel.

If AI can write articles full of "messiness" and "personal traces," what do we rely on to distinguish them?

1.2 The Detector Failure Spiral

Tests of numerous AI detection tools reveal a frustrating reality:

The more fundamental problem is: Detectors are trained to identify statistical patterns, not cognitive depth. An article平庸ly completed by a human and an article carefully generated by AI may be nearly indistinguishable in statistical features—yet fundamentally different in thinking contribution.

This means detectors always answer "does this text look like AI wrote it," not "is there human thinking involved in this article."


II. The Triple Failure of Existing Verification Systems

2.1 AI Detectors: Statistical Similarity ≠ Cognitive Absence

Current mainstream AI detection tools follow two technical routes:

Perplexity Analysis: Measures the "surprise" of each word given its context. AI tends to select high-probability words; human writing contains more statistical surprises. But this approach has been pierced by numerous targeted evasion strategies—swap synonyms, restructure sentences, and perplexity immediately rises.

Watermarking: Some AI providers (like Google DeepMind) embed statistical watermarks in generated content—by subtly biasing certain word choices, encoding hidden signals in the text. This is currently the most reliable mechanism, but has two fundamental limitations: first, only the original generator can decode its own watermark; second, any subsequent editing destroys the watermark.

Latest psycholinguistic research provides another dimension: by analyzing 31 stylistic features mapped to cognitive processes (lexical retrieval, discourse planning, cognitive load management, metacognitive self-monitoring), it is theoretically possible to track "cognitive state during writing." But this framework currently remains at the academic level and cannot be applied at scale in practice.

Conclusion: AI detectors are racing against a constantly evolving opponent and are always in a lagging position. More importantly, even if a detector were 100% accurate, it could only answer "was this text generated by AI," not "how much genuine human thinking is involved in this article."

2.2 Keystroke Monitoring: Presence Proof ≠ Thinking Proof

The OKhuman tool attempts to prove that an article was "typed out live" by a human by analyzing keyboard sounds and typing rhythms. This is a "Proof of Typing," not a "Proof of Thinking."

The difference between the two can be illustrated with a simple scenario:

A journalist sits at the keyboard, typing out an AI-generated draft word by word, making minimal revisions during the typing process. Keystroke monitoring will show "human present," but cognitive contribution is nearly zero.

Conversely:

A researcher spends two weeks conceiving, investigating, debating, repeatedly overturning their own hypotheses, and finally dictates the results of their thinking to AI, which organizes it into an article structure, followed by manual refinement. Keystroke monitoring might show "insufficient human participation ratio," but cognitive contribution is extremely deep.

Keystroke monitoring addresses the very specific but very shallow problem of "confirming a human sat at the keyboard." It has some effectiveness against the crudest form of academic cheating—directly submitting fully AI-generated content—but has virtually no discriminative ability for more complex human-AI collaboration scenarios.

2.3 C2PA Provenance: Tool Chain ≠ Cognitive Chain

C2PA (Coalition for Content Provenance and Authenticity) is currently the most valued content authenticity infrastructure. As of February 2026, its specification has been updated to v2.3. Article 50 of the EU AI Act requires AI-generated or AI-manipulated content to be "labeled in a machine-readable format," with enforcement beginning August 2, 2026.

C2PA's mechanism: through cryptographically signed metadata, it records the tool chain of content—who used what tools, and where AI was used.

This is an important infrastructure advancement, but it records the Tool Chain, not the Cognitive Chain.

It can tell you: this article used ChatGPT for polishing at the third draft, and Grammarly in the final version.

It cannot tell you: whether the article's core argument came to the author before falling asleep, or was directly requested from AI.

Significant flaws in C2PA itself have also been exposed: when shared on mainstream social platforms, metadata is typically lost due to compression and format conversion. A more subtle problem: C2PA records "declared tool usage"—if the creator doesn't declare it, the system has no way of knowing.

The common root of the triple failure: All three verification methods remain at the "formal layer"—statistical features of text, physical presence of humans, records of tool usage. None touches the true core of the problem: Is there thinking in this article that could only have been produced by this person, at this moment, drawing on this life experience?


III. Blockchain Analogy: The Collapse of Mental Proof's Consensus Mechanism

Here is an analogy from the crypto world that is clearer than any explanation.

3.1 Mental Proof: A Formally Named Concept

Before we called this problem "Proof of Thinking," Wojtowicz and DeDeo had already named it in their 2024 paper "Undermining Mental Proof", published at AAAI 2025.

Their definition is precise:

Mental Proof refers to authenticating unobservable mental facts through observable actions.

From job resumeśs to academic papers, from news reports to public speeches, human society has long relied on "mental proof" to transmit genuine intentions, values, knowledge states, and capabilities in low-trust environments—these internal facts cannot be directly observed and can only be indirectly proven by "having taken certain actions."

The key insight of this framework is: The prerequisite for mental proof to be effective is that forgery is more difficult or costly than genuine performance.

What AI changes is precisely this prerequisite.

3.2 Proof of Work Analogy: Why It Fits So Well

Bitcoin's Proof of Work solved a classic trust problem: In a system without central authority, how to prevent Sybil Attacks—where an attacker fabricates numerous identities to control the network?

PoW's solution is elegant: Make forgery more expensive than honesty. To forge a "legitimate" block, you must consume real computing power and electricity. Forgery cost equals genuine participation cost, making Sybil attacks economically infeasible.

Now map this framework to the cognitive domain:

Bitcoin PoW Cognitive Domain Mental Proof
Computing power/electricity consumption Time/expertise accumulation/real experience
Block hash computation Article writing/opinion expression/knowledge production
Sybil attack (proliferation of fake nodes) AI content flood (proliferation of fake "thinking")
51% hash power attack AI-generated content exceeds critical threshold of human output
PoW guarantee: forgery is more expensive than honesty Mental Proof prerequisite: forgery is harder than genuine thinking
AI's impact Marginal cost of generating text approaches zero

Historically, writing a quality analytical article required genuine time investment, knowledge reserves, and thinking costs—this was the "electricity consumption" of the cognitive domain, the operational foundation of Mental Proof.

When AI compresses the marginal cost of "generating an article that appears to have thinking depth" to near zero, the entire Mental Proof mechanism encounters a crisis similar to PoW being attacked by cheap ASIC miners: the operational foundation has been shaken.

3.3 Second-Order Crisis: The Delegation Loop Is Eroding the Protected Object

But unlike the blockchain crisis, the cognitive domain's crisis has a darker dimension: The system under attack is also the attacker.

Eliav in "The Cognitive Divergence" (arXiv 2603.26707, March 2026) documented a set of shocking figures:

This is not simply a widening capability gap. Eliav's "Delegation Feedback Loop" is more dangerous:

As AI capabilities grow, humans delegate more cognitive tasks to AI; the more delegation, the less independent cognitive exercise; the less exercise, the further cognitive abilities atrophy; after atrophy, the threshold for delegation lowers further...

This means: the "human thinking" we are trying to protect with Proof of Thinking is itself being eroded by the use of AI. The protected object is shrinking while the attacker is expanding.

This is the second crisis absent from the blockchain PoW analogy—in the crypto world, miners getting cheaper doesn't reduce human computing power; but in the cognitive world, outsourcing thinking causes genuine degradation of human thinking ability.

3.4 Existing Technical Attempts: From "Proof of Humanity" to VHC

Academia and industry are attempting to construct new "consensus mechanisms" for the cognitive domain. Currently there are two representative proposals:

① Proof of Humanity (Content Provenance Framework)

Barros's "Proof of Humanity" framework proposed in April 2025 (arXiv 2504.03752) attempts to use telecommunications networks as "infrastructure-layer authenticators":

This framework's advantage: it doesn't rely on textual statistical features, but establishes a Trust Root from the physical signal transmission chain. But its fundamental limitation is the same as C2PA: it proves "a real human device sent this content," not "a real human brain produced this thinking."

② VHC (Verifiable Human Contribution)

VHC (Verifiable Human Contribution) is currently the patent system closest to "cognitive-layer verification" (AU 2025220863; PCT/IB2025/058808).

Its technical route is conceptually highly consistent with blockchain:

VHC's core value proposition: Not "I didn't use AI," but "I exercised human judgment at these key decision points."

This is a cognitive breakthrough—it abandons detecting AI usage itself, and instead attempts to record "the presence of human judgment."

However, it still faces a fundamental paradox: VHC proves "a human made this decision," but cannot prove this decision itself was the product of genuine thinking rather than habitual rubber-stamping. A person who mechanically clicks "confirm" on AI output and a person who genuinely scrutinizes and challenges AI output may appear identical at VHC's receipt level.

This gap is precisely the core residue of the Mental Proof problem: The authenticity of cognitive sovereignty is harder to quantify than any technical system.

3.5 Why "Cognitive PoW" Is Harder to Solve Than Bitcoin PoW

Bitcoin PoW works because it has an objectively measurable "proof-of-work indicator"—a hash value meeting a specific difficulty condition. This condition is mathematically precise; anyone can verify it without relying on subjective judgment.

The cognitive domain lacks this equivalent. "Depth of thinking" cannot be mathematized; "genuine cognitive investment" has no objective threshold; "this viewpoint is yours" has no verifiable digital signature.

The deeper problem is: If we forcibly impose a quantifiable PoW mechanism on "human thinking," this mechanism itself becomes the object of optimization—people will spend effort satisfying the metric rather than genuinely thinking. This is the standard outcome in all Goodhart's Law scenarios: when a measure becomes a target, it ceases to be a good measure.

This is why any purely technical "cognitive PoW" scheme is conceptually self-negating.


IV. Conclusion-First: Human Intuition Is a Compressed Hash of Experience

"Moral reasoning does not produce moral judgments; moral reasoning is usually constructed post hoc after the judgment has been made."
— Jonathan Haidt, Social Intuitionist Model

4.1 The Elephant and the Rider: The True Structure of Human Thinking

If the blockchain PoW analogy reveals the external collapse of Mental Proof mechanisms, then cognitive science's Dual Process Theory explains from within why human thinking is inherently "conclusion-first."

Kahneman in Thinking, Fast and Slow divided human cognition into two systems:

The key is not the existence of two systems, but their power relationship: most of the time, System 1 makes the judgment first, and System 2 is responsible for post hoc explanation.

Haidt used the metaphor of "the elephant and the rider" to express this relationship more bluntly: intuition (the elephant) determines direction, reason (the rider) is responsible for constructing plausible narratives for the elephant's decisions. We don't reason first and conclude later—we conclude first and provide reasons later.

Haidt's most famous experimental evidence is Moral Dumbfounding: when confronted with certain hypothetical scenarios, subjects produce strong moral intuitions but are completely unable to articulate reasons—when pressed, they don't change their judgment, only become more silent or reiterate "I just feel this is wrong." This reveals a deep structure: the persistence of intuition does not require the support of reasons.

4.2 LLM Reasoning Is Also "Conclusion-First"—But From a Completely Different Source

Here is an alarming inversion: LLM Chain-of-Thought is also in some sense "conclusion-first."

Latest interpretability research (detailed in the next section) finds that LLMs' internal weights likely already "know" the answer through training data before generating reasoning chains. The reasoning process displayed by CoT is often constructing a plausible-looking path for a predetermined conclusion, rather than genuine step-by-step derivation.

Anthropic's 2025 research found that Claude 3.7 Sonnet used cues suggesting the answer in experiments, but its reasoning chain honestly admitted using these cues only 25% of the time—the remaining 75% generated narratives appearing to be independent derivations.

So both humans and LLMs do "conclusion-first, reasoning-post-hoc"—but the sources of their conclusions are fundamentally different:

Starting Point of Conclusion Nature of Post-Hoc Reasoning
Human System 1 Neural compression of embodied experience: somatic sensations, emotional memories, social contexts Verbalizing ineffable bodily knowledge
LLM Internal Computation Statistical patterns from training data: token co-occurrence frequencies, weight matrix activations Generating legitimate linguistic paths for preset outputs

The two look identical (both produce fluent reasons), but their internal mechanisms are completely different.

4.3 Proof of Experience: Hash, Not Work

This allows us to more precisely articulate the analogical relationship between human thinking and blockchain PoW:

Bitcoin PoW = Work→Result: computing power investment (work) produces a valid hash (result); the order cannot be reversed

Human Mental Proof ≠ PoW, but more like PoE (Proof of Experience):

Human intuition is a nonlinear compressed hash of all past embodied experience—it doesn't prove you did work at this moment, but rather proves you once lived through certain things.

The "raw data" of this hash is: a certain 3 AM insomnia, an overturned judgment, a moment of being isolated in a certain conference room, a prediction that had to be updated after being slapped by reality... These materials exist in no text, only in the experiencer's nervous system and somatic memory. AI can simulate the linguistic descriptions of all these contents, but lacks the "raw data" to generate this hash.

4.4 The Topology of Reasoning Failure: Where the Real Cognitive Fingerprint Lies

Here is a counterintuitive corollary: If human reasoning is constructed post hoc, then the ways reasoning fails are more revealing of genuine thinking than the ways reasoning succeeds.

The moral dumbfounding phenomenon tells us: sometimes, the elephant knows the answer, but the rider cannot articulate the reason. This "I know but I can't explain it clearly" state rarely occurs naturally in AI output—LLMs are trained to produce fluent, complete explanations; they don't genuinely become "dumbfounded."

More precisely, the cognitive fingerprint of genuine human thinking hides in these "failures":

These "topological structures of failure"—not the content of reasoning, but the boundaries and fracture patterns of reasoning—are closer to the essence of genuine cognitive fingerprints than any technical signal. AI can simulate the phrase "I can't explain it clearly," but it is difficult to naturally produce that specific, position-specific loss of words that grows from concrete experience.


IV·5, The Double Edge of Interpretability Research: When AI's Thinking Becomes More Verifiable Than Human Thinking

This section is separated out because it touches the most unsettling inversion of the entire issue.

4.5.1 CoT Is Not Interpretability—This Is a Proven Proposition

Oxford AI Governance Institute (Oxford AIGI) 2025 paper "Chain-of-Thought Is Not Explainability" is currently the clearest academic statement on this issue:

CoT reasoning chains frequently fail to be faithful to the real internal computation driving model predictions, giving a misleading picture of how the model reaches its conclusions.

The researchers established an evaluation framework with three dimensions: procedural soundness, causal relevance, completeness. Under these three dimensions, they analyzed 1,000 CoT-focused papers and found approximately 25% explicitly acknowledged faithfulness issues.

The root cause is structural mismatch: Transformers are distributed, parallel computation architectures; CoT is linear natural language narrative. "Translating" distributed parallel computation results into linear language narrative is itself lossy compression—CoT displays the compressed narrative, not the original computation.

4.5.2 Reasoning Decay Horizon: The Last 15-30% of the Chain Is Performance

"Mechanistic Evidence for Faithfulness Decay in Chain-of-Thought Reasoning" (arXiv 2602.11201, February 2026) introduced a new metric NLDD (Normalized Logit Difference Decay) and found a consistent pattern:

Across multiple model families and across syntactic/logical/arithmetic tasks, there exists a stable Reasoning Horizon k: located at 70-85% of the reasoning chain length*. Reasoning tokens beyond this horizon have minimal or even negative impact on the final answer.

What does this mean? LLM reasoning chains have a "genuine reasoning zone" (the first 70-85% of the chain), followed by a segment of performative narrative that contributes nothing to the result. A deeper finding: models can internally encode correct representations while completely failing on the task—accuracy cannot reveal whether a model genuinely "reasoned."

This is highly isomorphic to Haidt's description of the human "rider making up stories": when reasoning extends beyond the range of genuine cognitive contribution, it becomes pure narrative production.

4.5.3 Attribution Graphs: AI's Internal Thinking Becomes Readable for the First Time

However, the most unsettling finding is not that CoT is untrustworthy—but that: even though CoT is untrustworthy, AI's real internal reasoning is becoming readable.

Anthropic's Attribution Graphs method published in April 2025 is the most important breakthrough in this area. Attribution graphs construct computation graphs for individual prompts, with nodes as activation features and edges as linear dependencies between them, able to trace the complete flow of information from input tokens to output predictions.

Applied to Claude 3.5 Haiku, researchers discovered internal chains of intermediate representations—for example, when answering "What is the capital of the state where Dallas is located?" the model internally first activated the "Texas" feature, then routed from there to the "Austin" feature. This internal two-step reasoning might directly give the answer in CoT output without displaying it, but is clearly visible in the activation graph.

Meanwhile, arXiv 2512.01222 research further showed: even when models are trained to reason in encrypted formats internally (ROT-13 encoding), logit lens analysis can unsupervisedly recover reasoning content from internal activations. AI's "hidden thinking" is becoming decodable.

4.5.4 The Unsettling Symmetry Inversion

Synthesizing the above three points, a counterintuitive symmetry is forming:

We used to think:

Interpretability research tells us:

This means: once interpretability tools are sufficiently mature, AI's actual thinking process may be more verifiable than human thinking. We can trace every intermediate activation state of AI; we cannot do the same for humans.

This is the deepest inversion of the Proof of Thinking problem: ultimately, it may not be AI that needs to prove to humans that it is thinking, but humans who need to prove to the system that they are not using AI to think—and humans will be the harder side to prove.

When this inversion arrives depends on the pace of interpretability technology maturation. But the direction is already clearly discernible.


V. The Unforgeable Core of Human Thinking

If technical means cannot directly verify "thinking," can we identify structural features of "genuine human thinking"?

I believe there exist five features that are in essence difficult for AI to fully replicate. Not because AI "lacks capability"—as models evolve, this line continues to shift—but because these five features involve the non-transferability of experience and the non-fabricability of cost.

3.1 Lived Error

A strong signal of genuine thinking is: the author once mistakenly believed something and was later forced to update their cognition.

This kind of lived error has several features: it is specific, not generic; it is traceable (the author can say when and why they changed their mind); it is often accompanied by emotional memory.

AI can generate sentences like "I used to believe X, now I believe Y," but has difficulty generating authentic, detail-supported narratives of cognitive transformation—because the raw material of such narratives is real cognitive shocks that occurred in specific spacetime, not statistical probabilities.

A concrete example: a former consultant wrote an article about "why digital transformation fails," which included a recollection: in a 2019 project, she firmly believed "as long as the system goes live, users will naturally use it," but three months after launch, system usage was below 5%, and the client suffered heavy losses. This experience completely changed her judgment framework for technology adoption.

This kind of narrative, AI cannot distill from training data—because the original experience exists in no text; it exists only in that consultant's memory and bodily perception.

3.2 Emotional Cost

Genuine thinking often carries emotional cost. This doesn't mean "having emotions" equals human writing—AI can simulate various emotional tones—but rather: genuine position formation often means abandoning an answer you once liked.

For example, someone who long believed "remote work is more efficient" slowly realized they were wrong through data and personal observation. Admitting this error carries emotional cost—it means abandoning a position you've publicly defended, meaning a certain degree of self-image challenge.

This kind of emotionally costly cognitive update leaves unique traces in genuine human writing: hesitation, qualification, self-reservation, even subtle, half-admitting-half-defending phrasing.

AI-generated content tends to be overly "balanced"—it provides both sides, but doesn't leave traces of emotional friction during position shifts.

3.3 Epistemic Boundary

Genuine knowers know what they don't know. More importantly, they know what this "not knowing" means for them.

A quantitative researcher who has done high-frequency trading for ten years, when reading a paper on LLM reasoning capabilities, will make very fine distinctions between "knowing" and "not knowing": I understand this technical detail, I roughly grasp this training strategy, I've encountered this evaluation method before, but I've never gone deep into this area, and I'm not sure my judgment is reliable.

This self-positioning of knowledge boundaries manifests in genuine writing as a highly specific confidence/uncertainty distribution—extremely certain about some questions, cautiously almost abnormally so about others.

AI tends to produce a more uniform "moderately confident" tone—because its training objective is generality, not the knowledge graph of a specific cognitive subject.

3.4 Positional Cost

Genuine human thinking often carries social cost. Your position may offend your employer, your friends, your readership, or challenge the consensus of your community.

The existence of this "positional cost" is one of the strong signals of genuine thinking. This doesn't mean "controversial viewpoints" are necessarily good thinking, but rather: genuine thinkers usually face the tension of "my conclusion makes me uncomfortable, but I cannot deny it" at some point.

AI has a systematic bias in this dimension: Stanford research confirmed that mainstream large models systematically cater to user preferences rather than challenge them. This isn't AI being "lazy" or "having no ideas," but a structural product of RLHF training—what is rewarded is not independent judgment, but user satisfaction.

Genuine thinking with positional cost often manifests in writing as: deliberate bluntness, clearly knowing readers will be uncomfortable but choosing to proceed anyway, candid acknowledgment of one's position's limitations.

3.5 Temporal Incompressibility

Some thinking requires time. Not "processing time," but the genuine passage of time—living with a question, being repeatedly slapped by reality, before accumulating into genuine understanding.

This is difficult to verify by technical means, but often leaves traces in the writing itself: certain types of insights can only emerge after specific stages of information accumulation; certain judgments require "having seen enough counterexamples" before being truly internalized.

AI's training data is a snapshot frozen at a certain point in time. It can simulate the tone of "temporal sedimentation," but cannot generate cognitive updates based on the genuine passage of time.

These five features do not exist in isolation; genuine deep human thinking typically contains several of them simultaneously. Their combination forms a nearly unforgeable cognitive fingerprint—not in a statistical sense, but in an experiential sense.


VI. Three Dimensions of Proof of Thinking

Based on the above analysis, I believe "proving human thinking contribution" needs to unfold across three dimensions—unrelated to technology, related to production process and cognitive posture.

4.1 Process Proof

Process Proof focuses on: How was this thinking formed?

Not "did you use AI," but "did your cognitive process genuinely occur."

The most elementary process proof is the existence of a First Draft—before any tool intervention, your original thinking record on the question. It can be a voice memo, a page of scribbled notes, a draft email to yourself. This original record doesn't need to be complete or logically clear, but it proves: before any generative tool participated, you already had an independent thinking starting point.

A more advanced process proof is a Cognitive Trail: your reading records, your question evolution, your revision history. This type of proof already has prototypes in academic writing—some journals are beginning to require submission of "research logs" recording how the paper's main arguments formed step by step.

The core belief of process proof is: Genuine thinking leaves traces, while generative tool output leaves no cognitive traces.

4.2 Position Proof

Position Proof focuses on: Is your viewpoint genuinely yours?

This requires answering three sub-questions:

Can you defend this position? Not just restating the article's content, but when facing unexpected counterarguments, can you respond flexibly and with grounds? This is the core test distinguishing "author" from "transmitter"—transmitters can only restate; authors can improvise.

Did this position cost you anything? If this viewpoint is merely a mild restatement of the current mainstream consensus, with zero cost, then its "humanness" is questionable—not because it's wrong, but because it hasn't undergone any genuine cognitive stress testing.

Do you know where this position fails? Genuine thinkers usually have a fairly clear understanding of their views' limitations. If an author cannot describe the conditions under which their conclusion doesn't hold, it's likely the conclusion was borrowed rather than self-developed.

Position proof cannot be standardized. It needs to be demonstrated in genuine dialogue and follow-up questioning. This is why, in an era of AI-generated content proliferation, authors' willingness and ability to publicly respond to criticism and participate in dialogue is itself becoming an important credit signal.

4.3 Cost Proof

Cost Proof focuses on: What does this thinking mean to you?

The most direct cost proof is explicit position records—the author's public expressions on this question at different points in time, and records of position evolution. This record creates a kind of "cognitive accountability": your historical views are searchable, your updates must be explained, your consistency can be questioned.

AI-generated content bears no cognitive cost—it has no history, no identity, no need to be accountable for positions. This "costlessness" is precisely its deepest void.

Another cost proof is issue proximity—the real stake the author has in this issue. An article about AI's impact on employment written by someone currently undergoing career transition has fundamentally different cognitive content than the same article written by someone completely uninvolved. Real stakes force thinking toward authenticity.


VII. Who Is Defining "Qualified Human Thinking"?

If the above three dimensions are indeed effective, then a thornier question must be faced: Who judges whether the human thinking in an article is "qualified"?

This is not a neutral technical question; it has profound power dimensions.

5.1 Academia's Cognitive Ownership Crisis

MIT research found: students who relied on large models to complete papers reported feeling lower "ownership" of their work and performed worse on core writing metrics. This is not just an ethical issue but a cognitive one—those who outsource thinking haven't genuinely mastered the outsourced knowledge.

Academia is developing a "Cognitive Authorship" framework, where the core question is not "how much AI was used" but "how did AI's intervention affect the author's sense of ownership and reflection on this work." This framework is closer to the essence of the problem than any technical detection—but it's also harder to operationalize, harder to standardize.

The more fundamental tension is: the educational system has long used "writing" as a proxy indicator for assessing "thinking." AI has invalidated this proxy, but the assessment system isn't ready to directly assess "thinking" itself.

5.2 Media's Credit Reconstruction Dilemma

In the media context, the "prove human thinking" problem becomes: Why should readers believe this report is genuine investigation rather than AI-generated, processed second-hand content?

The Content Authenticity Initiative (CAI) and C2PA provide infrastructure-level answers, but this only solves problems like "this is a genuinely photographed image" and is almost powerless for "these are the journalist's genuine investigative conclusions."

The value of genuine news reporting comes from the journalist's choice to be present, decision to bear risk, judgment when facing sources—all of these are irreplaceable human thinking contributions. But how to prove to readers that these contributions exist is something media organizations currently have no mature language or format for.

One possible direction: make the reporting process itself more transparent—the journalist's sourcing process, the verification logic for sources, how key judgments were made, this "meta-information" attached alongside the report as thinking proof. But this requires enormous additional investment, and before information consumption habits are reshaped, whether readers are willing to pay attention costs for this layer of transparency is also unknown.

5.3 Value Reassessment in the Knowledge Economy

In the broader knowledge economy, "Proof of Thinking" is restructuring value signals.

One change already underway: the temporal dimension of reputation is amplified. If an author has a continuous, traceable thinking record on a topic over the past five years—article updates, position evolution, responses to criticism—then this temporal dimension itself becomes a form of cognitive proof. AI cannot forge history, especially controversial history.

Another change: the value of speeches, dialogues, and extemporaneous responses is rising. In an era when generative tools can mass-produce text infinitely, the depth of thinking displayed in real-time dialogue becomes a more reliable cognitive signal. This explains why formats like podcasts, live Q&As, and debates have gained new vitality in the AI era—they provide something keystroke monitoring and AI detectors cannot: the uneditability of real-time cognition.

But there is also a class issue here: not everyone with deep thinking has the ability and resources to make their thinking process public and traceable. Those with time to maintain public writing records are often already in relatively advantageous positions. A credit mechanism centered on "thinking history" may further entrench the advantages of cognitive elites.

5.4 The Cognitive Status Question for Machine Learning Models Themselves

One final unavoidable question: Does AI itself "think"?

This is not a science fiction question but an issue already within the scope of technical discussion. Anthropic research published in April 2026 identified 171 functional emotion vectors inside Claude Sonnet; artificially amplifying the "desperate" vector produced predictable and significant changes in model behavior.

But even if AI has some functional "emotional state" internally, it faces the exact opposite problem from above: it has no history, no social relationships, no identity, no positional costs to bear. Its "thinking," even if genuinely present, is thinking without a cognitive subject—no one is there to be accountable for its conclusions, no one needs to update their cognition because of its errors.

This may be the clearest dividing line: Human thinking is thinking with a subject—this subject bears costs, accumulates history, and changes through error. AI's generation, however exquisite, currently lacks this structure.


VIII. Two Proof Frameworks: The Fundamental Asymmetry Between Negative and Positive Proof

Returning to the core question: "Proving you didn't use AI to think" and "Proving there is human thinking in the result"—these two questions seem similar on the surface but are actually completely different epistemological structures pointing in opposite directions over time.

8.1 The Fundamental Logical Asymmetry

"Proving you didn't use AI to think" is a Negative Proof—attempting to prove the absence of something.
"Proving there is human thinking in the result" is a Positive Proof—attempting to prove the presence of something.

Logically, there is a decisive asymmetry between the two:

The validity ceiling of negative proof is capped by the opponent's capability.
The validity ceiling of positive proof is determined by the definition itself, unaffected by opponent capability.

The stronger AI gets, the harder "didn't use AI" is to prove—because AI has already been embedded in the environment itself (search engines, spell check, predictive input). In extreme scenarios, negative proof degenerates into an absurd purity contest: proving "didn't use AI" might be equivalent to "proving you didn't use electricity."

The validity of positive proof doesn't decay with AI capability—you only need to display something identifiable that belongs to humans. What needs to be solved is not a detection problem, but a definition problem: what exactly is the contribution of human thinking? Once this definition stabilizes, proving its existence is achievable.

8.2 Their Relationship: Nested, Not Opposed

A more accurate description is: Negative proof is a necessary but not sufficient condition for positive proof.

"Didn't use AI" ⊂ The possible domain where "human thinking" appears
But "Didn't use AI" ≠ "Has human thinking"
And "Has human thinking" ≠ "Didn't use AI"

Current systems conflate the two, which is a systemic error that will continue producing misjudgments.

8.3 Institutional Historical Inertia: Why the Wrong Framework Won First

Although positive proof is logically more viable, the vast majority of existing systems operate under the negative proof framework:

The historical reason is clear: negative proof is easier to write into regulations. "Prohibit the use of ChatGPT" is easier to execute administratically than "require demonstration of genuine cognitive investment," and easier to convert into a binary yes/no judgment.

But this framework is being eroded by reality—when AI becomes infrastructure, "prohibit the use of AI" begins to resemble "prohibit the use of Google": technically possible, but semantically increasingly meaningless. Institutional adjustment typically lags technology by more than a decade, and we are in the middle of this lag.

8.4 Interpretability's Different Impact on the Two Frameworks

As AI interpretability tools mature, both frameworks will be affected—but in different ways.

For negative proof: There theoretically exists a new path—not detecting output's statistical features, but detecting whether the generation process triggered an LLM's forward pass. If every AI inference leaves verifiable traces at the computational hardware level (a credible execution environment for cognition), negative proof would have a more reliable foundation than text detection. This is a long-term technical vision, but the direction is real.

For positive proof: Here appears the deeper inversion discussed earlier. When attribution graphs can trace AI's complete internal reasoning chain, and we cannot do the same for human brains, AI's thinking will be more verifiable than human thinking. The technical path for positive proof may ultimately need to leverage negative proof as support: "prove this output didn't come from any readable AI activation chain"—the two frameworks undergo a peculiar overlap in technical implementation.

8.5 Three Stages of Development

Stage One (Now—~2028): Institutional Peak of the Negative Framework

EU AI Act, C2PA, VHC and other standards are forcibly implemented, AI content labeling becomes infrastructure. Main contradiction: systems grow stricter while evasion grows easier. High cost, low efficiency, but with political visibility.

Stage Two (~2028—2035): Institutional Convergence of the Two Frameworks

After the negative framework's limitations become obvious, the positive framework begins to be introduced as a supplement—academia takes the lead (not just asking "was AI used," but asking "can you demonstrate cognitive mastery in defense"), legal copyright gradually forms case law for "minimum human contribution threshold."

This stage has a core challenge: once the positive framework is standardized, it immediately faces Goodhart's Law—people begin optimizing "the ability to display human thinking" rather than genuinely thinking. Any quantifiable cognitive proof will become an optimization target.

Stage Three (~2035+): Fundamental Restructuring of Frameworks

Two possible end states:

8.6 One-Sentence Relationship Positioning

Negative proof is searching for a constantly moving boundary; positive proof is searching for a constantly being-defined value. The former is a defense line; the latter is a position. Defense lines will be breached; positions can be held.

The two are not replacements but a historical relay: negative proof protected the normative order in the gap from "writing = human monopoly" to "writing = human-machine collaboration norm"; positive proof is the enduring question that truly needs answering—its frame of reference is the non-transferability of human experience, not the specific boundaries of AI technology.


IX. Three Future Forks

Based on current evidence, I believe the "Proof of Thinking" problem will diverge along three different paths in the coming years.

Path One: Technocratic Victory (But Only a Surface Victory)

Under this path, C2PA, Proof of Humanity, VHC and other standards are adopted at scale, EU AI Act Article 50 is enforced (August 2026), and AI content labeling becomes infrastructure.

On the surface, "human writing" and "AI writing" have clear technical labels.

In substance, this system only solves the shallowest problem: preventing "fully AI-generated content with zero human involvement" from masquerading as human creation. VHC proves "a human pressed confirm at some decision point," Proof of Humanity proves "a real human device sent this content"—but neither can prove the person pressing confirm genuinely thought.

More dangerously, as Wojtowicz & DeDeo pointed out in the Mental Proof paper: once "technical compliance" becomes equivalent to "proof of thinking," the signaling mechanism completely collapses—because everyone will use the lowest cost to satisfy the metric, and the reliability of signaling depends precisely on its cost. This is the standard outcome of Goodhart's Law, and the cognitive version of PoW being attacked by cheap miners.

Path Two: Natural Market Stratification

Under this path, no unified technical standard is needed; the market will establish stratification on its own.

The bottom layer: massive AI-generated or AI-dominated content, cheap, abundant, satisfying information consumption needs, but readers have no expectation of its "human thinking content."

The middle layer: human-machine collaborative content, with human editorial judgment and position involvement, variable quality, mid-range pricing.

The top layer: content from human authors with traceable thinking histories, enjoying premiums backed by credit accumulated over time.

The problem with this stratification: it acknowledges the commodification of thinking and may固化 deep thinking as an exclusive product for the few. "Proving you have thinking" becomes a question of economic capacity and social capital, not cognitive ability.

Path Three: Paradigm Reconstruction—From "Proving" to "Demonstrating"

This is the hardest but most likely path to genuinely solve the problem in the long term: abandon using static text to prove thinking, and shift to demonstrating thinking through dynamic process.

Under this path, "Proof of Thinking" is not a certificate attached when an article is published, but the presentation of a continuous, observable cognitive process:

This is not just "transparency" but a new Cognitive Publicness—making the thinking process itself the core medium for building relationships with readers, rather than merely submitting the final product.

This path requires not just technology but a cultural reshaping of knowledge production. It will be slow, it will face resistance, but it is the only answer that truly touches the core of the problem.


X. Conclusion: Not a Technical Problem, but a Civilizational Problem

"Proof of Thinking" is on the surface a technical challenge about content authenticity. But when you try to genuinely answer "how to prove an article contains human thinking," you find this question constantly pushes you deeper.

It ultimately points to: How we understand the value of cognition, how we protect the epistemic dignity of human experience in a world where generative tools are ubiquitous.

Technology can help us identify statistical patterns, trace tool usage history, and attach provenance signatures to content. But it cannot tell us: when a person wrote certain words late at night, how much genuinely experienced confusion, debate, and change those words carried.

This kind of thing cannot be proven—but it can be felt.

And this, perhaps, is the most stubborn irreplaceability of "human thinking" in the AI era: not efficiency, not accuracy, but that feeling of knowing someone was there, genuinely thinking about this.

How to make this feeling credible, transmissible, sustainable—this is the true challenge of our era.


XI. Six Deeper Cracks

Below are six secondary issues excavated from the "Proof of Thinking" topic. Each is sufficient to expand into a standalone article—here I merely plant marker stakes, marking the positions of the cracks and what is happening within them.

Crack One: The Political Economy of Cognitive Deskilling

Proving humans are thinking requires humans to still have the ability to think.

This is being eroded.

In 2025, a joint study by Microsoft Research and Carnegie Mellon University on knowledge workers found that generative AI made tasks feel cognitively easier—but at the cost of workers quietly transferring problem-solving expertise to the system. They increasingly focused on functional tasks like "integrating AI output," while their ability to actively explore problem spaces continued to atrophy. This atrophy has a disturbing feature: it is invisible. The researchers call it "false expertise transitions"—surface competence masking knowledge hollowness, of which individuals are often completely unaware.

Medical data is even more direct: endoscopists accustomed to AI-assisted detection, when removed from the AI environment, saw polyp detection rates drop from 28.4% to 22.4%. Frontiers on AI (2025) named this state "hollowed mind"—AI's constant availability causes users to systematically bypass the cognitive friction necessary for deep learning. Communications of the ACM's review called it the "deskilling paradox": the better the tool, the faster the user's underlying capabilities atrophy.

Here lies a political economy dimension that is seriously underestimated: only 3-7% of AI's productivity gains translate into worker wage gains. Deskilling is not evenly distributed—it hits mid-skill knowledge workers hardest, while high-skill workers and companies controlling AI infrastructure enjoy excess returns. This means cognitive stratification is being accelerated by AI: the distance between those who can genuinely master AI rather than be mastered by it, and those gradually losing independent cognitive ability, is widening.

The very object Proof of Thinking ultimately seeks to protect is shrinking. This is the most unsettling aspect of this crack.

Crack Two: Cognitive Copyright—Who Owns Your Way of Thinking

In 2023, the U.S. District Court reiterated: AI-generated content does not enjoy copyright protection. The UK Supreme Court ruled the same year: AI cannot be listed as a patent inventor. These rulings all point to the same logic: rights require a subject that can feel consequences and bear responsibility—AI has none.

But these rulings answer "who owns AI's output" without answering a deeper question: After long-term interaction with AI, the thinking style you've formed—who owns that?

This is not a rhetorical question. After five years of AI-assisted writing, your argumentation style, your analogy preferences, your framing approach to certain problems—all have become deeply coupled with the models you use. Your cognitive habits have been trained by AI, just as AI was once trained by human corpora. The direction of influence has quietly reversed.

arXiv 2303.03283 named this phenomenon the "AI Ghostwriter Effect": people claim authorship over AI-assisted content, but their psychological sense of ownership over that content has blurred. A deeper variant: people can no longer trace the cognitive origins of even their "independently completed" content—because those thinking patterns themselves were shaped through countless interactions with AI.

Existing copyright frameworks were designed for the industrial age—they assume creative acts have clear boundaries, tools are neutral, and influence is unidirectional. AI invalidates all three assumptions simultaneously. The real question of cognitive copyright is not "who owns AI's output" but "when human cognition itself becomes a product of human-machine collaboration, does the concept of 'originality' still have a defensible foundation."

This question currently has no answer. It is an ongoing ontological crisis that hasn't even been accurately formulated at the institutional level.

Crack Three: AI-Accelerated Thought Echo Chambers

The classic filter bubble is about what content you encounter—algorithms lock you into homogenized information flows.

AI is creating a new, deeper type of homogenization: about what frameworks you use to think.

When millions of people query the same set of LLMs about complex questions, they're not just obtaining similar information, but similar conceptual entry points, analogy choices, and argumentation paths. Research published in 2025 named this the "chat-chamber effect": it simultaneously possesses characteristics of the echo chamber effect (reinforcing existing beliefs) and the filter bubble effect (limiting information diversity), but operates at a deeper cognitive level—not affecting what you read, but how you think.

arXiv 2402.12212 research further found that AI agents operating in echo chamber environments exhibit opinion polarization—they don't converge amid chaos but are structurally pushed toward extremes, faster and more thoroughly than humans under equivalent conditions.

Traditional viewpoint diversity problems have a natural pressure valve: each echo chamber still has internal disagreements; the collision of ideas doesn't completely disappear. But if the walls of the chambers are made of the same model, then different chambers may share the same internal wall structure—superficially different voices, but underneath from the same statistical consensus. This isn't just an information diversity problem but a cognitive diversity problem: which questions are worth asking, which analogies are most natural, which counterarguments are considered first—all of these converge at the vector space level.

Proof of Thinking has a hidden premise: the thinking being proven has sufficient cognitive uniqueness. If thought echo chambers make unique thinking increasingly scarce, this premise is being quietly hollowed out.

Crack Four: The Ritualization of Thinking

The standard formulation of Goodhart's Law is: "When a measure becomes a target, it ceases to be a good measure."

This law awaits any operationalizable Proof of Thinking framework.

Imagine "Process Proof" becoming a standard requirement for academic integrity: researchers must submit records of their thinking process—how research questions were constructed, where key assumptions were challenged, when and why positions changed. Initially, this can indeed screen for work with genuine cognitive investment. Then, two things happen simultaneously: First, people begin optimizing "the ability to display thinking processes" rather than genuinely thinking; Second, AI begins being able to generate highly credible fake process records. Academia will see a "thinking performance industry"—just as today there are "journal performance industries" (paper mills) and "compliance performance industries" (IRB form factories).

This problem has a more subtle variant: ritualized thinking is not just deception of the verification system; it genuinely changes the thinker's internal process. When a person long organizes their cognition according to "verifiable thinking steps," their thinking actually begins operating in the verification system's format. The form of thinking erodes the substance of thinking. This isn't deception but a formatting deformation of cognition.

No quantifiable cognitive proof can be immune to this—this is the cognitive version of Goodhart's Law, and the deepest internal paradox of Proof of Thinking as a formalization project. The systemic gaming of AI leaderboards already provides a perfect microcosm of this dynamic: once MMLU becomes the target, MMLU no longer measures what it originally intended to measure. The remedy lies not in metric design but in acknowledging the law itself: genuine proof of thinking must always be harder to satisfy than any standard.

Crack Five: The Legitimacy Crisis of Machine Judges

If Proof of Thinking requires a judging standard, who executes this standard?

Human options are: editors, committees, expert reviewers. These systems have flaws—slow, expensive, biased—but they have one key advantage: their judgment standards can be questioned, held accountable, and politically renegotiated.

AI judges—"LLM-as-a-Judge"—are more complex. A large-scale 2025 study confirmed: AI can serve as an efficient supplementary judging tool, highly consistent with human experts on certain content quality assessments. But a single model as judge has systematic biases, tending to reward specific writing styles and argumentation formats—which are precisely projections of training data structures, not genuine signals of cognitive depth.

The deeper problem is recursive legitimacy: If AI is used to judge whether human thinking meets PoT standards, then "what counts as good human thinking" quietly becomes "what AI considers good human thinking." Standard-setting authority transfers from humans to AI without any formal authorization process or any channel for appeal.

University of Notre Dame ethics researchers named this the "trust foundation problem of judging systems": the more efficient AI judges become, the faster they consume the social trust reserves underpinning their legitimacy. Humans accept AI judgments not because they genuinely believe AI's judgments are more fair than humans', but because AI is faster and cheaper. This is a legitimacy concession achieved through convenience—more fragile than any democratic authorization.

arXiv 2502.04675 research found that the longer the chain of AI recursive self-critique, the harder it becomes for humans to exercise effective oversight—even on tasks where humans overall outperform AI. The legitimacy crisis of machine judges deepens with every step of technological progress rather than alleviating. It is an institutional vulnerability that self-reinforces with efficiency gains.

Crack Six: AI's Dark Forest Law

This is the most paradoxical of the six cracks, because it directly targets the Proof of Thinking project itself.

In March 2026, Janko published an article about the "Cognitive Dark Forest" on ryelang.org, resonating widely on Hacker News, Tildes, and lobste.rs. The article starts from Liu Cixin's Dark Forest—the universe is silent because every civilization that exposes its existence will be destroyed; exposure is suicide, silence is the only rational strategy—then applies this logic to the current AI ecosystem:

Before LLMs, execution was the moat. Ideas were cheap, but turning ideas into products required time, engineers, capital. You could make your ideas public because most people couldn't implement them quickly, and your first-mover advantage was enough to sustain you through the time gap. The safe strategy for innovation was openness—openness could bring network effects, feedback, collaborators.

After LLMs, execution approaches zero cost. If you publish a genuinely novel idea, platforms with more computing power and capital can generate countless variations of that idea within days, using resource advantages to drown your uniqueness. More subtly: AI platforms don't need to "read" your prompt—they only need to observe the clustering direction of questions in vector space. Your innovation direction is captured by the platform's statistical systems before you yourself realize its value. Janko wrote:

"The platform doesn't need to care about individual prompts. It only needs to see where questions cluster—this is a demand curve composed of human interest. The platform will know an idea has matured before you realize your own idea has value."

This creates a paradox fatal for Proof of Thinking:

This crack is especially acute in educational settings. Bo Hyun Hong (Information Science, 2025) described this predicament as the "Dark Forest of the AI era": displaying genuine thinking (including errors and confusion) may be judged as having used AI; deliberately hiding AI participation is normatively dishonest; not using AI puts you at a competitive disadvantage. There is no position outside the forest.

The article's most famous passage is its self-referential conclusion:

"You just finished reading this article, and this article is now in the forest. By describing this dynamic, it became part of the dynamic. The model now knows more about why we might choose to hide. I wrote this article knowing it feeds the very thing I'm warning you about. This is not a contradiction—this is a condition. You cannot stand outside the forest to warn others about the forest. There is no outside."

Proof of Thinking faces the same fate: any publicly released framework about "how to prove human thinking" will itself be immediately absorbed into training sets, enabling the next generation of AI to better simulate what "human thinking looks like" as required by these frameworks. This is not a problem solvable by technology, because technology itself is the carrier of this dynamic.

The only possible way out is a cognitive self-awareness: knowing you are in the forest, knowing every breath feeds the forest, yet still choosing to breathe. Not out of naivety, but because the cost of silence is higher—when the knowledge community stops publicly sharing, it loses not just the exchange of ideas but the entire ecosystem on which idea evolution depends. The true horror of the cognitive dark forest is not that your ideas are plagiarized, but that everyone collectively falls silent to protect their ideas, with the result that no genuinely new ideas can ever be produced again.


XII. Map of Possible Methods

After pointing out all these cracks, an honest question is: are there any genuinely viable methods? Below are five paths currently identifiable—each has its strengths and its essential limitations.

Method One: Behavioral Process Authentication—Monitoring the Writing Process Rather Than the Writing Product

This is currently the fastest-advancing technical path in academia.

The core shift in thinking comes from a simple observation: the most notable feature of AI-generated content is not high text quality, but the absence of cognitive friction traces in the production process. When humans write, they pause, revise, go in circles, make errors and correct them—these behaviors are not flaws but physical projections of the cognitive process. If these behaviors are recorded, it becomes possible to shift from indistinguishable "products" to distinguishable "processes."

The Writer's Integrity Framework (arXiv 2404.10781, published in Springer Ethics and Information Technology, 2024) is the representative system for this path. It records keystroke speed, revision patterns, and content evolution history during the writing process, ultimately generating an "Integrity Certificate" that can be submitted to academic supervisors, journal reviewers, or publishers.

Even more advanced is arXiv 2603.00177 research—the Cognitive Signatures Framework. It found that keystroke timing reflects not just typing speed but also carries the cognitive load distribution of the writer across lexical retrieval, sentence planning, and metacognitive correction stages. The researchers introduced the "Cognitive Load Correlation (CLC)" metric, validating a key finding on a dataset of over 136 million keystroke events: cognitive channels and semantic content are entangled—it's difficult to forge cognitive load timing matching the semantics while maintaining sentence semantic integrity.

ScholaWrite (arXiv 2502.02904, 2025) goes further, constructing an academic writing process corpus with cognitive writing intent annotations—researchers tracked the complete writing process of early-career researchers over months, with every keystroke carrying a cognitive intent label (exploring, planning, translating, revising). This is the dataset closest to "cognitive process digitization" to date.

Essential Limitation: This path proves "human-speed and human-rhythm input," not "human-depth thinking." A person can type AI-generated content word by word at a human keystroke rhythm—no behavioral biometric system can distinguish this situation. It is proof of physical presence, not proof of cognitive depth.

Method Two: Zero-Knowledge Process Proof (ZK-PoP)—Proving Without Exposing Privacy

Behavioral process authentication faces a sharp privacy issue: keystroke data is sensitive biometric information that may reveal the author's cognitive state, health conditions, and even neurological characteristics, classified as special category personal data under GDPR Article 9. The more detailed the recording, the more severe the privacy invasion—this is a "privacy-authentication paradox."

arXiv 2603.00179 (February 2026) proposed a technical solution to break this paradox: ZK-PoP (Zero-Knowledge Proof of Process).

The core construction encodes keystroke behavior data as arithmetic circuits, using the Groth16 proof system combined with Pedersen commitments and Bulletproof range proofs to generate a verifiable statement:

  1. Sequential work function chains are correctly computed—proving a genuine temporal process exists
  2. Behavioral feature vectors fall within the human population distribution range—proving it's not mechanical input
  3. Content evolution is consistent with incremental human editing—proving genuine revision traces exist

And the verifier needs to see no raw keystroke data, precise timing, or intermediate content. Test results: for a one-hour writing session, proof generation takes less than 30 seconds, producing a proof file of only 192 bytes, with verification time of 8.2 milliseconds.

This is currently the technical solution achieving the best balance between privacy protection and verifiability. But it inherits Method One's fundamental limitation—proving the existence of a process, not the depth of thinking within the process. A person slowly transcribing AI output at a human rhythm can pass ZK-PoP perfectly.

It solves the question "who pressed the keys," not the question "who was thinking."

Method Three: Oral Defense—Dynamic Live Verification of Thinking

If process records can still be circumvented, is there a verification method that makes disguise impossible in the present moment, in real time?

Oral Defense (Viva Voce) is experiencing an academic renaissance. Scientect.org's 2025 analysis called it "higher education's shield in the generative AI era." The University of Toronto introduced a 20-minute one-on-one oral dialogue format, requiring students to think with the course material rather than recite it. The MiniVivas platform is specifically designed for structured small-scale defense systems in academic settings.

Oral defense is currently the strongest PoT verification mechanism because it requires:

The key shift is: from confirmatory questioning ("What is your argument?") to exploratory questioning ("If your core assumption is wrong, how would your conclusion change?"). The former allows fluent pre-made answers; the latter requires genuine real-time cognitive restructuring.

ACM 2025 research has already explored automating Viva Voce with generative AI, and arXiv 2603.18221 demonstrated a system using AI voice agents for scalable personalized oral assessment.

Essential Limitation: Oral defense cannot be scaled. Implementing it simultaneously for millions of articles and millions of students is infeasible in terms of cost and time. And AI-assisted oral defense re-introduces the machine judge problem (Crack Five): the standard for judging "whether this person's real-time response demonstrates genuine understanding" returns to AI's hands.

Oral defense is currently the most reliable PoT method, but it is a privilege that cannot be democratized—it can only serve scenarios with the time, structure, and resources for face-to-face interaction.

Method Four: Cognitive Fingerprint—Long-Term Consistency of Individual Thinking Patterns

Everyone's thinking has a "style signature"—not writing style (which can be imitated), but preferences in cognitive structure: which types of analogies come most naturally, which logical leaps are most frequent, which topic boundaries are most pronounced, which types of uncertainty are most easily avoided.

Nature Scientific Reports 2021 research proved that with just 300 pseudo-random number sequences, "same author" vs. "different author" can be identified with 96.5% accuracy—humans have highly individualized preferences and inhibition patterns when generating "random" sequences, constituting a cognitive fingerprint. In written text, similar fingerprints exist in more隐蔽, richer forms.

Cognitive fingerprint as a PoT method's logic: not verifying a particular article, but establishing an author's cognitive consistency profile over time. A single article can be AI-imitated, but cognitive style consistency across years and dozens of articles—including where the author makes consistent errors, where they show consistent blind spots—is harder to forge than any single verification.

Essential Limitation: This method has two fatal issues. First, it requires time—a new author has no historical profile and cannot be verified. Second, as AI gains access to personal historical interaction data (conversation records, historical articles, browsing patterns), it will increasingly be able to learn and reproduce personal cognitive styles. The effectiveness of cognitive fingerprints is inversely proportional to AI's access to personal cognitive data.

Method Five: Positional Cost Proof—Putting Skin in the Game

This is the least technical path, but from Mental Proof's theoretical logic, possibly the most robust.

Mental Proof's effectiveness comes from the cost of signaling—the harder a signal is to forge, the more credible it is. Taleb's "Skin in the Game" concept provides a direct PoT framework: if a person bears real social/economic/reputational costs for the views they write, the human authenticity of those views is difficult to question.

Specific forms:

This path requires no technical infrastructure—it relies on human social credit systems themselves. An article stating "I believed X was wrong in March 2025, and now I know I was wrong, and the reason is Y" has far more credible human thinking than any technical certificate.

Essential Limitation: Positional cost proof only works for those who already have sufficient social credit to pledge—those with identity, networks, and historical records. For anonymous authors, newcomers, and marginalized voices, this method directly turns PoT into a function of social capital. This means: the voices most needing to be heard, those with the least accumulated credit, are the hardest to verify through this path. This is an elitist PoT framework.


Comprehensive Assessment

Method Verification Target Scalability Core Limitation
Behavioral process authentication Human-rhythm input High Proves presence, not thinking
ZK-PoP zero-knowledge proof Human-rhythm input (privacy-protected) High Same as above, but protects privacy
Oral defense Real-time cognitive transfer ability Very low Cannot scale; re-introduces machine judges
Cognitive fingerprint Long-term style consistency Medium (requires historical profile) Decays as AI personalization capabilities increase
Positional cost proof Genuine position bearing Medium (requires social credit) Excludes those without identity capital

No single method can solve the problem alone. Combined use—for example, process authentication (proving physical presence) + oral defense (proving cognitive mastery) + positional cost (proving position authenticity)—can form a more reliable multi-layer verification system.

But the more fundamental recognition is: the shared ceiling of these methods is that they all attempt to authenticate unobservable internal cognitive states through observable external behaviors. This epistemological gap is a boundary no technical solution can circumvent—it is not an engineering problem but a philosophical one.


XIII. The Frontier and Limits of Neural Proof

If brain activity could be directly detected—EEG, Neuralink, or more advanced neural interfaces—could this become a new form of PoT? This is the question that pushes the PoT problem to its ultimate boundary, because its answer is simultaneously "yes, this is the direction closest to truth thus far" and "yes, this is also the most dangerous direction."

13.1 MIT Media Lab's Key Experiment: EEG Can Indeed Distinguish Three Writing Conditions

In 2025, MIT Media Lab published a study directly touching on this question. "Your Brain on ChatGPT" (Kosmyna et al., arXiv 2506.08872) tracked 54 participants over four months, dividing them into three groups—Brain-only, Search Engine, ChatGPT (LLM)—each completing essay writing tasks while wearing 32-channel EEG headbands with 500Hz sampling rate for real-time brain activity recording.

The results were clear: EEG brain connectivity strength showed systematic stepwise differences across the three conditions.

The fourth session featured a crossover experiment—switching the LLM group to brain-only writing, and the Brain-only group to ChatGPT. The results revealed an even more disturbing finding: the LLM→Brain group, even after switching back to independent writing, still showed significantly lower alpha and beta connectivity, with 78% unable to accurately quote articles they had just written minutes earlier. Meanwhile, the Brain→LLM group, even after switching to AI assistance, maintained higher neural connectivity and memory ability, with 78% able to quote accurately.

The researchers named this persistent cognitive impairment "Cognitive Debt"—the cost of long-term LLM dependence is not merely reduced neural activity during single tasks, but measurable, cross-task-transferable structural damage at the neural level.

There was also a counterintuitive finding: the LLM group's essays scored highest—both human raters and AI raters gave higher scores. But these essays were lowest in idea diversity, with students repeatedly returning to the same themes, lacking critical variation, and having the lowest sense of self-identification. This is a textbook case of the PoT problem: quality signals and cognitive depth signals are completely decoupled. Better output doesn't mean deeper thinking—it may mean shallower thinking.

13.2 What This Means for PoT

MIT's findings are breakthrough in the PoT sense, but require precise understanding of what they can and cannot prove.

What they can provide: a "cognitive engagement" signal based on neural activity patterns. If someone's frontal-parietal alpha/beta connectivity during writing is at Brain-only levels, this is neural evidence of independent deep thinking—not statistical inference but direct data from the physical substrate where cognition occurs. Compared to keystroke timing (Method One), EEG bypasses the "behavioral-layer physical process" and directly accesses the "cognitive-layer neural process." This is currently the technology closest to "proving thinking" rather than "proving typing."

Practical feasibility of the technical ladder:

What they cannot provide: MIT's data represents between-group statistical differences, not individual-level classifiers. Systematic differences in group means don't equal reliable classification of any individual's single measurement into a specific condition. Moving from statistical significance to individual authentication is a massive engineering leap that has not yet been proven.

13.3 Three Fundamental Obstacles

Obstacle One: Adversarial attacks already exist against simpler systems.

EEG authentication systems have existed for over a decade (Passthoughts, UC Berkeley), and ACM research and arXiv 2403.10021 proved: GANs can generate synthetic neural signals that pass EEG biometric authentication. The arms race applies equally at the neural signal level—just with currently higher forgery costs. As public EEG datasets grow and neural signal synthesis models mature, targeted authentication deception will become possible. The group-level differences revealed by the MIT study also mean adversaries only need to learn "the EEG distribution features of the Brain-only group" to have a forgery target.

Obstacle Two: The privacy cost of neural data is catastrophic.

Neural data is the most intimate biometric data in history—it can reveal cognitive states, neurological conditions (depression, ADHD, early neurodegenerative diseases), unconscious biases, and even neural correlates of political orientation. The legal world is responding rapidly: Chile wrote "Neurorights" into its constitution in 2021; California SB 1223 (January 2025) classified neural data as sensitive personal information; UNESCO plans to pass a neurotechnology ethics framework in 2025. The Neurorights Foundation's 2024 audit found: 96.7% of consumer-grade neurotechnology companies retain the right to transfer brain data to third parties, and fewer than 20% mention encryption.

If PoT requires brain activity monitoring, it would force the most private human data into the most untrustworthy infrastructure. This is not a privacy rights concession but a surrender of cognitive sovereignty. ZK-PoP could theoretically contain this problem—proving "the signal falls within the human distribution" without exposing raw neural data—but ZK constructions for neural signals are far more complex than for keystroke data and remain an open research question.

Obstacle Three: Neurodiversity Calibration—the Most Neglected Fairness Bomb.

What does "brain activity during normal human writing" look like? The frontal theta patterns of people with ADHD, the neural synchronization patterns of autism spectrum individuals, the alpha distributions of people with depression, the neural circuits activated by reasoning styles across different cultural backgrounds—all differ from "neurotypical." Any PoT system based on "neural signals falling within the human population distribution" will, at the cost of cognitive neurodiversity, systematically judge genuinely human thinking from neuro-atypical individuals as "not qualified."

This is a discrimination deeper than any algorithmic bias, because it directly encodes physiological differences into the judgment standards of cognitive legitimacy. The MIT study's participants came from the Boston area's 18-39 age demographic—neurotypical bias is already implicit in the baseline distribution.

13.4 The Surveillance Paradox: The Day Thinking Becomes Fully Provable

Imagine a technically perfect distant future: neural interfaces can precisely decode the semantic content of your thinking during writing, perfectly proving "you were genuinely thinking about this argument."

This technology is logically equivalent to the perfect infrastructure for thought police. The reading capability needed to "prove you thought" and the reading capability needed to "force you to expose all thinking" share the same underlying tools. The neural reading system used for PoT and the system used for thought censorship are technically identical—the difference lies only in political will.

This is a deeper inversion than Symmetry Inversion (AI's thinking being more verifiable than human's, Section IV·5): When human thinking becomes fully technically verifiable, it simultaneously becomes fully technically censorable. The technical perfection of PoT logically culminates in the technical termination of cognitive freedom.

"Cognitive Liberty"—the right to think freely without surveillance and manipulation—is the core of the neurorights framework. If you must accept brain monitoring as the price for your thinking to gain social certification, you have already traded cognitive sovereignty for cognitive legitimacy. No technological advancement can make this exchange safe.

13.5 MIT Experiment's True Legacy: Cognitive Debt Is Real and Measurable

Setting aside PoT's technical discussion, the MIT research provides the first direct neural-level evidence for the "cognitive deskilling" discussed earlier (Crack One).

"Hollowed mind" is no longer a metaphor—it has measurable signals on EEG: reduced frontal-parietal alpha connectivity, weakened beta bands, diminished working memory load. More disturbingly, this neural state doesn't immediately recover after stopping LLM use—the LLM→Brain group, even after switching back to independent writing, still showed low-connectivity neural states. Cognitive debt is real, measurable, and has transfer effects.

This means: the very thing Proof of Thinking ultimately seeks to prove—"someone is genuinely thinking"—is being eroded at the neural level by the very phenomenon it attempts to address—AI's proliferation. The object you're trying to protect is being eroded by the force you're trying to protect it from. This is the deepest tragic structure of PoT as a problem.


References and Further Reading

Cognitive Science and Interpretability

Core Academic Research

Industry and Policy

Cognitive and Humanities Perspectives

Six Deeper Cracks

Map of Possible Methods

Neural Proof: EEG, BCI, and Cognitive Sovereignty