Reflex Babbling: Learning to Stand Up from Scratch on a Low-Cost Quadruped in 76 Minutes of Real-World Trials

This working paper reports a consumer quadruped robot that learns to stand up from a limp, belly-down pose at run time, entirely on its own hardware. The robot is an XGO-Lite V2: about 0.55 kg, 2.3 kg·cm hobby servos, closed servo firmware, and a Raspberry Pi CM4 that runs both the controller and the learner. All learning happens on the robot from a zero-initialized controller. No policies, weights or demonstrations from simulation are used. A MuJoCo model of the robot serves only as a pre-deployment check that the setup cannot damage the hardware and that the algorithm can learn the task at all. The controller is a 22-parameter linear reflex. Every 30 ms it maps three biologically motivated signals to joint targets: vestibular tilt, a body-schema estimate of body height, and a tendon-like joint load signal. CMA-ES, running on the CM4, tunes the reflex from a fitness computed only from the robot's own sensors. Results: After 22 generations (286 episodes, 76 minutes of episode time), the learned reflex stood up in 10 of 10 evaluation trials, to 109.1 ± 0.3 mm with 3.0 ± 0.1° tilt. That is 13 to 15 mm taller than the vendor's built-in stand, with less servo load. An independent camera-based check with an AprilTag confirmed the body-schema height estimate to within 5.4 ± 1.0 mm. The paper also documents the failure modes of unattended learning on uncooperative hardware, and the safeguards that fixed them: an undocumented firmware sensor stream; an IMU whose zero changes at every power-up; a robot stranded on its side that fed the optimizer identical scores; recovery of a corrupted learning state by deterministic on-device replay. Finally, it analyses how the system aligns with the Alberta Plan for AI research, and why episodic evolution strategies are used for now. It proposes a roadmap: an options architecture of learned motion primitives, a body-first curriculum, a homeostatic energy drive, and a hardware-in-the-loop stage. Compute and memory budgets for these are measured on the CM4.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-05
DOI
https://doi.org/10.5281/zenodo.23172894
Primary Topic
Robotic Locomotion and Control
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Reflex Babbling: Learning to Stand Up from Scratch on a Low-Cost Quadruped in 76 Minutes of Real-World Trials

Bernd Porr
Zenodo (CERN European Organization for Nuclear Research)
Robotic Locomotion and Control
article

Reflex Babbling: Learning to Stand Up from Scratch on a Low-Cost Quadruped in 76 Minutes of Real-World Trials

Bernd Porr
article en

Abstract

This working paper reports a consumer quadruped robot that learns to stand up from a limp, belly-down pose at run time, entirely on its own hardware. The robot is an XGO-Lite V2: about 0.55 kg, 2.3 kg·cm hobby servos, closed servo firmware, and a Raspberry Pi CM4 that runs both the controller and the learner. All learning happens on the robot from a zero-initialized controller. No policies, weights or demonstrations from simulation are used. A MuJoCo model of the robot serves only as a pre-deployment check that the setup cannot damage the hardware and that the algorithm can learn the task at all. The controller is a 22-parameter linear reflex. Every 30 ms it maps three biologically motivated signals to joint targets: vestibular tilt, a body-schema estimate of body height, and a tendon-like joint load signal. CMA-ES, running on the CM4, tunes the reflex from a fitness computed only from the robot's own sensors. Results: After 22 generations (286 episodes, 76 minutes of episode time), the learned reflex stood up in 10 of 10 evaluation trials, to 109.1 ± 0.3 mm with 3.0 ± 0.1° tilt. That is 13 to 15 mm taller than the vendor's built-in stand, with less servo load. An independent camera-based check with an AprilTag confirmed the body-schema height estimate to within 5.4 ± 1.0 mm. The paper also documents the failure modes of unattended learning on uncooperative hardware, and the safeguards that fixed them: an undocumented firmware sensor stream; an IMU whose zero changes at every power-up; a robot stranded on its side that fed the optimizer identical scores; recovery of a corrupted learning state by deterministic on-device replay. Finally, it analyses how the system aligns with the Alberta Plan for AI research, and why episodic evolution strategies are used for now. It proposes a roadmap: an options architecture of learned motion primitives, a body-first curriculum, a homeostatic energy drive, and a hardware-in-the-loop stage. Compute and memory budgets for these are measured on the CM4.

Zenodo (CERN European Organization for Nuclear Research)
Openalex Percentile: Top 23%
Robotic Locomotion and Control
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Reflex Babbling: Learning to Stand Up from Scratch on a Low-Cost Quadruped in 76 Minutes of Real-World Trials — Bernd Porr · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS