In a sentence: Sutton and Rafiee formally put cognitive science's enactive cognition on the RL agenda—arguing that "perception itself is a skillful action," and honestly acknowledging that this proposition has not yet been operationalized, offering only a discussable, falsifiable research agenda.
The final question in the paper—"Does a software agent with tools and APIs count as embodied?"—is almost an ontological question tailor-made for the MCP / computer-use era. It and concurrent engineering topics (agents generating their own experience, compressing skills into trainable skill documents) are the theoretical and engineering faces of the same proposition: one provides the cognitive essence for skill, the other makes skill into an artifact.
Enactive (embodied) cognition is not just a slogan; the paper breaks it down into four separately discussable concepts, comparing each against the current state of AI implementation—clarifying "where the mainstream has reached, and where enactive wants to push."
Cognition is grounded in ongoing interaction; "the world is its own best model" (Brooks). Rule systems have no experience; supervised learning learns fixed datasets in one pass; RL puts experience back at the core (self-collected data). Echoes Silver & Sutton's "The Era of Experience" and the Big World Hypothesis.
Perception is mastering sensorimotor couplings; to perceive is to act (Noë / Merleau-Ponty's intentional arc · maximal grip). The mainstream still treats perception as "passive extraction prior to action"; video generation models can continue patterns, but cannot skillfully intervene when patterns break.
Autopoiesis self-maintenance → normativity emerges from self-preservation. Supervised learning does not self-evaluate; standards are externally given; RL uses reward to self-evaluate entire trajectories, but reward is still externally specified; intrinsic motivation / hindsight learning are moving closer.
Body morphology determines possible couplings and affordances; it is a constitutive condition of cognition. The mainstream makes it "pattern recognition on static datasets"; embodied RL treats the body as an external constraint; soft robotics / morphological computation proves "the body computes," yet remains marginal.
The four concepts form a progressive scale: from "having experience or not" to "whether experience is self-generated and inseparable from the body." RL is already firmly established in the first slot; the further along, the more open the territory.
To perceive is not to receive the world —
it is a skillful way of acting in it. — THE ENACTIVE THESIS, AS RAFIEE & SUTTON FRAME IT FOR RL
The paper's most restrained yet crucial judgment is this: the relationship between RL and enactive is one of structural resonance, not equivalence. Three resonances genuinely exist—self-generated experience, action-centricity, and temporally extended reward evaluation; but three gaps are equally real:
RL's reward is externally given, whereas enactive requires normativity to emerge endogenously from the agent's self-maintenance.
In RL, action and perception are still two separable modules; enactive demands that the two mutually constitute each other and are irreducible in principle.
The mainstream treats the body as an external constraint or engineering detail; enactive views the body as a constitutive condition of cognition.
The paper admits: this proposition has not yet been operationalized. It does not pretend to give answers; instead, it puts four not-yet-quantifiable questions on the table—which is exactly what an honest position paper should do:
The fourth question is almost an ontological question tailor-made for the MCP / computer-use era: when an agent's "body" is the set of tools and API boundaries it can invoke, the word "embodiment" needs to be redefined.
This position paper provides a unified theoretical coordinate for the recurring theme of "agents generating their own experience" (Codex for Knowledge Work, CooperBench, situational awareness in the evaluation era). While the engineering world is busy making agents self-collect data and self-evaluate, this paper asks: what do these actions mean in cognitive science, and what steps are still missing.
It forms a beautiful contrast with engineering practices like "compressing skillful procedures into trainable skill documents": one makes skill into an artifact, the other provides the cognitive essence for skill—precisely the engineering and theoretical faces of the same proposition. The former argues that "perception / cognition is itself skillful engagement," while the latter compresses this engagement into reusable documents.
It continues the thread of Sutton's "The Era of Experience," extending a slogan into a discussable, falsifiable research agenda. For AI engineers, its value lies not in being usable today—but in pointing out the three hurdles current architectures have yet to cross: "reward externality," "action-perception separability," and "embodiment definition," and formally handing the question of "whether software agents count as embodied" to the MCP / computer-use era.
📌 Window note: This article is compiled from arXiv:2605.24238v1 (2026-05-22), a manually supplemented deep-dive from AI Buzzwords EP.88 "Topic II." It connects with Topic I SkillOpt (how agents learn), Topic III Microsoft Build 2026 (who manages agents / where they run), and Topic IV Palantir AIPCon 10 (how agents land in industries) to form the "Agent Control Plane" main thread—this piece answers the most ontological question among them: what exactly is an agent.
First published 2026-06-05