The Token of Physical Intelligence
This paper proposes a hypothesis-driven perspective on the learning substrate of physical intelligence. While recent robot foundation models have largely focused on robot trajectories, actions, action tokenization, and multimodal observations, we ask whether human-generated control sequences produced during robot teleoperation constitute a more fundamental and potentially transferable unit of behavioral data. We distinguish human-side control sequences from robot-side trajectories, emphasizing the transformation from human intent and interface interaction to embodiment-specific robot execution. Inspired by the analogy between human-generated linguistic sequences and token-based language modeling, as well as prior work on native human interface action learning in digital environments, we investigate whether an analogous learning substrate exists for physical intelligence. The paper does not claim that human control sequences are mathematically equivalent to linguistic tokens. Instead, it formulates this as a testable research hypothesis and identifies a potential research direction around interface-independent representations of heterogeneous human control signals across robots, tasks, and embodiments. We review related work in behavioral foundation models, teleoperation learning, robot action tokenization, cross-embodiment learning, and human-generated action modeling, and propose experimental criteria for testing whether human control sequences provide advantages in transfer, compositionality, sample efficiency, and generalization.
Authors
- huajun he
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-05
- DOI
- https://doi.org/10.5281/zenodo.23169030
- Primary Topic
- Reinforcement Learning in Robotics
- Type
- article
- Field-Weighted Citation Impact
- 0.00