The Token of Physical Intelligence

This paper proposes a hypothesis-driven perspective on the learning substrate of physical intelligence. While recent robot foundation models have largely focused on robot trajectories, actions, action tokenization, and multimodal observations, we ask whether human-generated control sequences produced during robot teleoperation constitute a more fundamental and potentially transferable unit of behavioral data. We distinguish human-side control sequences from robot-side trajectories, emphasizing the transformation from human intent and interface interaction to embodiment-specific robot execution. Inspired by the analogy between human-generated linguistic sequences and token-based language modeling, as well as prior work on native human interface action learning in digital environments, we investigate whether an analogous learning substrate exists for physical intelligence. The paper does not claim that human control sequences are mathematically equivalent to linguistic tokens. Instead, it formulates this as a testable research hypothesis and identifies a potential research direction around interface-independent representations of heterogeneous human control signals across robots, tasks, and embodiments. We review related work in behavioral foundation models, teleoperation learning, robot action tokenization, cross-embodiment learning, and human-generated action modeling, and propose experimental criteria for testing whether human control sequences provide advantages in transfer, compositionality, sample efficiency, and generalization.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-05
DOI
https://doi.org/10.5281/zenodo.23169030
Primary Topic
Reinforcement Learning in Robotics
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

The Token of Physical Intelligence

huajun he
Zenodo (CERN European Organization for Nuclear Research)
Reinforcement Learning in Robotics
article

The Token of Physical Intelligence

huajun he
article en

Abstract

This paper proposes a hypothesis-driven perspective on the learning substrate of physical intelligence. While recent robot foundation models have largely focused on robot trajectories, actions, action tokenization, and multimodal observations, we ask whether human-generated control sequences produced during robot teleoperation constitute a more fundamental and potentially transferable unit of behavioral data. We distinguish human-side control sequences from robot-side trajectories, emphasizing the transformation from human intent and interface interaction to embodiment-specific robot execution. Inspired by the analogy between human-generated linguistic sequences and token-based language modeling, as well as prior work on native human interface action learning in digital environments, we investigate whether an analogous learning substrate exists for physical intelligence. The paper does not claim that human control sequences are mathematically equivalent to linguistic tokens. Instead, it formulates this as a testable research hypothesis and identifies a potential research direction around interface-independent representations of heterogeneous human control signals across robots, tasks, and embodiments. We review related work in behavioral foundation models, teleoperation learning, robot action tokenization, cross-embodiment learning, and human-generated action modeling, and propose experimental criteria for testing whether human control sequences provide advantages in transfer, compositionality, sample efficiency, and generalization.

Zenodo (CERN European Organization for Nuclear Research)
Openalex Percentile: Top 10%
Reinforcement Learning in Robotics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.