Efficient Learning from Synthetic Data and Immersive VR Televisualization for Robotic Avatars

In industrial settings, robots already enhance productivity, precision, and safety, taking on tasks that are repetitive or hazardous. However, adapting perception systems to novel tasks is often hampered by the scarceness of suitable in-domain training data. Generation of synthetic training data is possible, but the domain gap between synthetic training data and real world must be addressed. In this thesis, we start with 2D synthetic scene generation via a Cut & Paste paradigm, with which we achieved second place in the Amazon Robotics Challenge 2017. We then increase realism by generating physically plausible 3D scenes, also in efficient online manner. Models trained in this way achieve accuracy comparable to methods trained on hand-annotated real data. We investigate refinement of the generated training samples to make them even more realistic using a generative adversarial network. A contrastive loss and patch-based training ensure label consistency and prevent mode collapse. The refined training data results in model accuracy very close to real training data. As an application of synthetic training data, we train robust surface features using a contrastive loss between frames with different lighting conditions and arrangements. The resulting features can be fused onto 3D object meshes, which allows direct rendering into feature space. We successfully utilize this in a render-and-compare framework for 6D pose refinement. Finally, we also demonstrate how to leverage the immense knowledge and generalization capability present in large foundation models to provide regularization during the domain adaptation process. A contrastive loss structure serves as the adapter between the foundation model output and the specific task at hand. The method results in stable and hierarchical dense features and, as intended, improves accuracy in the target domain: Models trained using our method surpass those trained on real data. In addition to robotic perception, we also address teleoperation and avatar robotics. Advances in this field are opening new social applications, allowing people to project their skills and presence across distance. We present the NimbRo Avatar system, which won the M ANA Avatar XPRIZE competition, and describe several components that were key to this success. We detail the immersive VR visualization system, which features a robot head movable in 6D, resulting in correct parallax and disocclusions. Our spherical rendering method hides communication and movement latencies. We also propose a novel view synthesis method operating on dynamic scenes in real time, which can replace the costly 6D actuation of the head. The method achieved state-of-the-art results on the challenging RealEstate10k dataset.

Authors

Institutions

Publication Details

Journal
bonndoc (University of Bonn)
Published
2026-09-16
DOI
https://doi.org/10.48565/bonndoc-968
Primary Topic
Generative Adversarial Networks and Image Synthesis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Efficient Learning from Synthetic Data and Immersive VR Televisualization for Robotic Avatars

Max Schwarz
bonndoc (University of Bonn)
Generative Adversarial Networks and Image Synthesis
article

Efficient Learning from Synthetic Data and Immersive VR Televisualization for Robotic Avatars

Max Schwarz
article en

Abstract

In industrial settings, robots already enhance productivity, precision, and safety, taking on tasks that are repetitive or hazardous. However, adapting perception systems to novel tasks is often hampered by the scarceness of suitable in-domain training data. Generation of synthetic training data is possible, but the domain gap between synthetic training data and real world must be addressed. In this thesis, we start with 2D synthetic scene generation via a Cut & Paste paradigm, with which we achieved second place in the Amazon Robotics Challenge 2017. We then increase realism by generating physically plausible 3D scenes, also in efficient online manner. Models trained in this way achieve accuracy comparable to methods trained on hand-annotated real data. We investigate refinement of the generated training samples to make them even more realistic using a generative adversarial network. A contrastive loss and patch-based training ensure label consistency and prevent mode collapse. The refined training data results in model accuracy very close to real training data. As an application of synthetic training data, we train robust surface features using a contrastive loss between frames with different lighting conditions and arrangements. The resulting features can be fused onto 3D object meshes, which allows direct rendering into feature space. We successfully utilize this in a render-and-compare framework for 6D pose refinement. Finally, we also demonstrate how to leverage the immense knowledge and generalization capability present in large foundation models to provide regularization during the domain adaptation process. A contrastive loss structure serves as the adapter between the foundation model output and the specific task at hand. The method results in stable and hierarchical dense features and, as intended, improves accuracy in the target domain: Models trained using our method surpass those trained on real data. In addition to robotic perception, we also address teleoperation and avatar robotics. Advances in this field are opening new social applications, allowing people to project their skills and presence across distance. We present the NimbRo Avatar system, which won the M ANA Avatar XPRIZE competition, and describe several components that were key to this success. We detail the immersive VR visualization system, which features a robot head movable in 6D, resulting in correct parallax and disocclusions. Our spherical rendering method hides communication and movement latencies. We also propose a novel view synthesis method operating on dynamic scenes in real time, which can replace the costly 6D actuation of the head. The method achieved state-of-the-art results on the challenging RealEstate10k dataset.

bonndoc (University of Bonn)
University of Bonn (DE)
Decent work and economic growth
Openalex Percentile: Top 13%
Generative Adversarial Networks and Image Synthesis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.