Skip to content
← DeepDive Agents & Models · 中文
DEEPDIVE / [MODEL] · Enactive Cognition
DD · 0016 · 2026-06-05
RAFIEE & SUTTON · Position Paper ARXIV:2605.24238 · MAY 22 Perception as Skillful Action

Sutton Puts Embodied Cognition on the RL Table

A position paper co-authored by RL founding father Sutton and Rafiee, arguing for systematically introducing cognitive science's enactive (embodied) cognition into mainstream AI and RL.
Perception is not "the brain passively receiving input → processing → issuing commands," but rather "a skillful activity of perceiving through action, and understanding how one's own actions shape experience."
This piece breaks down its four key concepts, structural resonance and non-equivalence with RL, and the ontological question it leaves for the MCP / computer-use era.
Paper
05·22
arXiv · cs.AI · Position Paper
Key Concepts
4
Experience · Action-Perception · Autonomy · Embodiment
RL × Enactive
Resonance
Structural resonance, but not equivalence
Left for Others
4 Qs
Un-operationalized open questions
Agent Embodied Actor Environment The world is its own best model Action Perception Sensorimotor coupling · Intentional arc · Maximal grip Experience Self-generated data Action-Perception Inseparable Autonomy Normativity / autopoiesis Embodiment Body as cognitive condition RL structural resonance (self-generated experience · action-centric · reward evaluation) but not equivalence
SVG · The core of enactive: perception is not passive reception, but "perceiving through action"—action and perception mutually constitute each other in agent↔environment interaction; the paper argues for introducing the four concepts above into RL (diagram created based on the paper's arguments).

In a sentence: Sutton and Rafiee formally put cognitive science's enactive cognition on the RL agenda—arguing that "perception itself is a skillful action," and honestly acknowledging that this proposition has not yet been operationalized, offering only a discussable, falsifiable research agenda.

Counter-consensus insight

The final question in the paper—"Does a software agent with tools and APIs count as embodied?"—is almost an ontological question tailor-made for the MCP / computer-use era. It and concurrent engineering topics (agents generating their own experience, compressing skills into trainable skill documents) are the theoretical and engineering faces of the same proposition: one provides the cognitive essence for skill, the other makes skill into an artifact.

§ 01 / Four Concepts

Four Key Concepts
× Current AI Comparison

Enactive (embodied) cognition is not just a slogan; the paper breaks it down into four separately discussable concepts, comparing each against the current state of AI implementation—clarifying "where the mainstream has reached, and where enactive wants to push."

01 · EXPERIENCE

Cognition is grounded in ongoing interaction; "the world is its own best model" (Brooks). Rule systems have no experience; supervised learning learns fixed datasets in one pass; RL puts experience back at the core (self-collected data). Echoes Silver & Sutton's "The Era of Experience" and the Big World Hypothesis.

02 · ACTION-PERCEPTION Inseparability

Perception is mastering sensorimotor couplings; to perceive is to act (Noë / Merleau-Ponty's intentional arc · maximal grip). The mainstream still treats perception as "passive extraction prior to action"; video generation models can continue patterns, but cannot skillfully intervene when patterns break.

03 · AUTONOMY

Autopoiesis self-maintenance → normativity emerges from self-preservation. Supervised learning does not self-evaluate; standards are externally given; RL uses reward to self-evaluate entire trajectories, but reward is still externally specified; intrinsic motivation / hindsight learning are moving closer.

04 · EMBODIMENT

Body morphology determines possible couplings and affordances; it is a constitutive condition of cognition. The mainstream makes it "pattern recognition on static datasets"; embodied RL treats the body as an external constraint; soft robotics / morphological computation proves "the body computes," yet remains marginal.

The four concepts form a progressive scale: from "having experience or not" to "whether experience is self-generated and inseparable from the body." RL is already firmly established in the first slot; the further along, the more open the territory.

To perceive is not to receive the world —
it is a skillful way of acting in it.
— THE ENACTIVE THESIS, AS RAFIEE & SUTTON FRAME IT FOR RL
§ 02 / Tension

Structural Resonance
But Not Equivalence

The paper's most restrained yet crucial judgment is this: the relationship between RL and enactive is one of structural resonance, not equivalence. Three resonances genuinely exist—self-generated experience, action-centricity, and temporally extended reward evaluation; but three gaps are equally real:

Evaluation Still Externally Specified

RL's reward is externally given, whereas enactive requires normativity to emerge endogenously from the agent's self-maintenance.

Action-Perception Not Truly Inseparable

In RL, action and perception are still two separable modules; enactive demands that the two mutually constitute each other and are irreducible in principle.

Embodiment Treated as Implementation Detail

The mainstream treats the body as an external constraint or engineering detail; enactive views the body as a constitutive condition of cognition.

The paper admits: this proposition has not yet been operationalized. It does not pretend to give answers; instead, it puts four not-yet-quantifiable questions on the table—which is exactly what an honest position paper should do:

The fourth question is almost an ontological question tailor-made for the MCP / computer-use era: when an agent's "body" is the set of tools and API boundaries it can invoke, the word "embodiment" needs to be redefined.

§ 03 / So What

Why This Thread
Is Worth Connecting

This position paper provides a unified theoretical coordinate for the recurring theme of "agents generating their own experience" (Codex for Knowledge Work, CooperBench, situational awareness in the evaluation era). While the engineering world is busy making agents self-collect data and self-evaluate, this paper asks: what do these actions mean in cognitive science, and what steps are still missing.

It forms a beautiful contrast with engineering practices like "compressing skillful procedures into trainable skill documents": one makes skill into an artifact, the other provides the cognitive essence for skill—precisely the engineering and theoretical faces of the same proposition. The former argues that "perception / cognition is itself skillful engagement," while the latter compresses this engagement into reusable documents.

It continues the thread of Sutton's "The Era of Experience," extending a slogan into a discussable, falsifiable research agenda. For AI engineers, its value lies not in being usable today—but in pointing out the three hurdles current architectures have yet to cross: "reward externality," "action-perception separability," and "embodiment definition," and formally handing the question of "whether software agents count as embodied" to the MCP / computer-use era.

📌 Window note: This article is compiled from arXiv:2605.24238v1 (2026-05-22), a manually supplemented deep-dive from AI Buzzwords EP.88 "Topic II." It connects with Topic I SkillOpt (how agents learn), Topic III Microsoft Build 2026 (who manages agents / where they run), and Topic IV Palantir AIPCon 10 (how agents land in industries) to form the "Agent Control Plane" main thread—this piece answers the most ontological question among them: what exactly is an agent.

Revision history

First published 2026-06-05