Efficient Learning from Synthetic Data and Immersive VR Televisualization for Robotic Avatars
In industrial settings, robots already enhance productivity, precision, and safety, taking on tasks that are repetitive or hazardous. However, adapting perception systems to novel tasks is often hampered by the scarceness of suitable in-domain training data. Generation of synthetic training data is possible, but the domain gap between synthetic training data and real world must be addressed. In this thesis, we start with 2D synthetic scene generation via a Cut & Paste paradigm, with which we achieved second place in the Amazon Robotics Challenge 2017. We then increase realism by generating physically plausible 3D scenes, also in efficient online manner. Models trained in this way achieve accuracy comparable to methods trained on hand-annotated real data. We investigate refinement of the generated training samples to make them even more realistic using a generative adversarial network. A contrastive loss and patch-based training ensure label consistency and prevent mode collapse. The refined training data results in model accuracy very close to real training data. As an application of synthetic training data, we train robust surface features using a contrastive loss between frames with different lighting conditions and arrangements. The resulting features can be fused onto 3D object meshes, which allows direct rendering into feature space. We successfully utilize this in a render-and-compare framework for 6D pose refinement. Finally, we also demonstrate how to leverage the immense knowledge and generalization capability present in large foundation models to provide regularization during the domain adaptation process. A contrastive loss structure serves as the adapter between the foundation model output and the specific task at hand. The method results in stable and hierarchical dense features and, as intended, improves accuracy in the target domain: Models trained using our method surpass those trained on real data. In addition to robotic perception, we also address teleoperation and avatar robotics. Advances in this field are opening new social applications, allowing people to project their skills and presence across distance. We present the NimbRo Avatar system, which won the M ANA Avatar XPRIZE competition, and describe several components that were key to this success. We detail the immersive VR visualization system, which features a robot head movable in 6D, resulting in correct parallax and disocclusions. Our spherical rendering method hides communication and movement latencies. We also propose a novel view synthesis method operating on dynamic scenes in real time, which can replace the costly 6D actuation of the head. The method achieved state-of-the-art results on the challenging RealEstate10k dataset.
Authors
- Max Schwarz (ORCID: https://orcid.org/0000-0002-9942-6604)
Institutions
- University of Bonn (DE)
Publication Details
- Journal
- bonndoc (University of Bonn)
- Published
- 2026-09-16
- DOI
- https://doi.org/10.48565/bonndoc-968
- Primary Topic
- Generative Adversarial Networks and Image Synthesis
- Type
- article
- Field-Weighted Citation Impact
- 0.00