Reward-Free Scooter Balance Control via Diffusion World Models with Goal-Conditioned Trajectory Generation

We present a reward-free control framework for balancing and steering a two-wheeled scooter using a diffusion-based world model. Rather than engineering a reward, we specify goals directly in observation space: target values (e.g., zero roll and zero yaw error) are pinned through a continuous mask, and classifier-free guidance amplifies the goal signal during trajectory generation. Because the mask is continuous at inference, goals can be traded off online (for instance, relaxing the balance constraint during sharp turns to allow necessary leaning) without retraining. The model is a FiLM-Mixer denoising network trained with V-prediction diffusion. At deployment, the controller runs in real time using a single diffusion step with warm-started predictions. We validate the approach on a full-sized Thormang3 humanoid operating a Gogoro Viva scooter in simulation, and deploy it on physical hardware. It matches a PPO baseline tuned with six reward components on balance, survival, and heading tracking while producing smoother commands, all without the per-task reward-shaping step. Diffusion training introduces its own loss-weight hyperparameters; unlike reward weights, however, these are task-agnostic. They govern the denoising procedure rather than the desired behavior, and are therefore set once and reused unchanged across goals rather than re-tuned for each new task. Because the model learns to predict trajectories rather than to maximize a reward, its training signal depends only on observed states and actions, not on reward labels. Real hardware recordings can therefore be folded directly into the same loss, providing a route toward closing the sim-to-real gap that reward-based methods such as PPO structurally cannot use.

Authors

Institutions

Publication Details

Journal
Machines
Published
2026-09-11
DOI
https://doi.org/10.3390/machines14091036
Primary Topic
Zebrafish Biomedical Research Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Reward-Free Scooter Balance Control via Diffusion World Models with Goal-Conditioned Trajectory Generation

Saeed Saeedvand, Ugo Roux, Jacky Baltes
Machines
Zebrafish Biomedical Research Applications
article

Reward-Free Scooter Balance Control via Diffusion World Models with Goal-Conditioned Trajectory Generation

Saeed Saeedvand, Ugo Roux, Jacky Baltes
article en

Abstract

We present a reward-free control framework for balancing and steering a two-wheeled scooter using a diffusion-based world model. Rather than engineering a reward, we specify goals directly in observation space: target values (e.g., zero roll and zero yaw error) are pinned through a continuous mask, and classifier-free guidance amplifies the goal signal during trajectory generation. Because the mask is continuous at inference, goals can be traded off online (for instance, relaxing the balance constraint during sharp turns to allow necessary leaning) without retraining. The model is a FiLM-Mixer denoising network trained with V-prediction diffusion. At deployment, the controller runs in real time using a single diffusion step with warm-started predictions. We validate the approach on a full-sized Thormang3 humanoid operating a Gogoro Viva scooter in simulation, and deploy it on physical hardware. It matches a PPO baseline tuned with six reward components on balance, survival, and heading tracking while producing smoother commands, all without the per-task reward-shaping step. Diffusion training introduces its own loss-weight hyperparameters; unlike reward weights, however, these are task-agnostic. They govern the denoising procedure rather than the desired behavior, and are therefore set once and reused unchanged across goals rather than re-tuned for each new task. Because the model learns to predict trajectories rather than to maximize a reward, its training signal depends only on observed states and actions, not on reward labels. Real hardware recordings can therefore be folded directly into the same loss, providing a route toward closing the sim-to-real gap that reward-based methods such as PPO structurally cannot use.

MachinesVol. 14(9)
National Taiwan Normal University (TW)
Openalex Percentile: Top 14%
Zebrafish Biomedical Research Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Reward-Free Scooter Balance Control via Diffusion World Models with Goal-Conditioned Trajectory Generation — Saeed Saeedvand, Ugo Roux, et al. · Machines (2026) | TGRS Research Map | TGRS