Neural Representations for Simultaneous Localization and Mapping

Autonomous robots need to know where they are and what surrounds them to operate reliably in the real world. This challenge of reconstructing the environment and localizing the robot within it is known as simultaneous localization and mapping (SLAM). Both tasks are grounded in a scene representation, an internal model of the robot's surroundings that also supports downstream applications such as navigation and simulation. This thesis addresses what form this representation should take and how it can be built and continuously updated from sensor measurements. Conventional scene representations such as point clouds, voxels, surfels, and meshes store geometry in explicit and discretized data structures, and have long formed the basis of modern robotic systems. However, their fixed-resolution discretization introduces information loss during fusion, and robots often need to maintain multiple redundant representations for different downstream tasks. Recent learning-based scene representations model scenes in a continuous and compact form, optimizing learnable parameters via differentiable functions to fit sensor observations. Yet many such methods are confined to small-scale offline operation, prone to catastrophic forgetting during long-term deployment, unable to accommodate loop closure corrections without remapping, and limited to a single sensor modality. To address these challenges, this thesis proposes a novel point-based implicit neural scene representation. It consists of sparse neural points, each with a local coordinate frame and an optimizable latent feature, decoded by globally shared neural networks into scene geometry, appearance, or object-level structure on demand. The representation is locally grounded, continuously learnable, scalable, and versatile. The thesis makes four main contributions that progressively extend this representation. The first contribution presents, to the best of our knowledge, the first full-fledged LiDAR SLAM system built entirely on a neural implicit representation. It integrates correspondence-free odometry, incremental map learning, and loop closure into a full pipeline, where loop closure corrections elastically deform the neural point map without remapping. The system runs at sensor frame rate and achieves state-of-the-art localization and surface reconstruction accuracy on established benchmarks while requiring less map memory than previous scene representations. While this system captures scene geometry, it cannot model appearance. The second contribution extends the representation to LiDAR-visual SLAM by jointly modeling a signed distance field and a Gaussian splatting radiance field in the same neural point map, where geometry and appearance mutually improve each other through a consistency loss. The system scales to kilometers-long trajectories while maintaining a globally consistent and compact map, achieving state-of-the-art photorealistic rendering and geometric reconstruction accuracy. Both contributions, however, treat the scene as a whole without modeling individual objects or their category specific structure. Thus, the third contribution addresses object-level modeling in cluttered horticultural environments, where occlusion makes per-instance shape estimation particularly challenging. It embeds learned shape priors into an occlusion-aware differentiable rendering pipeline that jointly estimates complete fruit geometry and pose from partial RGB-D observations, outperforming existing baselines in both shape completion and pose accuracy. The resulting models have been integrated into downstream robotic systems for autonomous fruit harvesting and safe leaf manipulation. All three contributions, however, operate within a single mapping session. To overcome this limitation, the fourth contribution addresses multi-session map fusion by casting multi-view point cloud registration as a single-stage 3D generation problem, bypassing exhaustive pairwise matching. A flow matching model with rigidity-enforcing sampling directly generates registered point clouds in a canonical frame from unposed multi-view inputs. Trained on large-scale cross-domain data, the model achieves top registration accuracy with the shortest runtime and scales favorably with view count. It also generalizes zero-shot across scene scales, sensor modalities, and low-overlap conditions, enabling downstream tasks such as offline SLAM and map merging. Together, these contributions establish neural point-based representations as a practical and unified foundation for SLAM and downstream tasks spanning geometry, appearance, object-level structure, and multi-view fusion. All methods achieve state-of-the-art results on established benchmarks, have been published in peer-reviewed venues, and are publicly available as open-source software.

Authors

Institutions

Publication Details

Journal
bonndoc (University of Bonn)
Published
2026-10-05
DOI
https://doi.org/10.48565/bonndoc-996
Primary Topic
Robotics and Sensor-Based Localization
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Neural Representations for Simultaneous Localization and Mapping

Yue Pan
bonndoc (University of Bonn)
Robotics and Sensor-Based Localization
article

Neural Representations for Simultaneous Localization and Mapping

Yue Pan
article en

Abstract

Autonomous robots need to know where they are and what surrounds them to operate reliably in the real world. This challenge of reconstructing the environment and localizing the robot within it is known as simultaneous localization and mapping (SLAM). Both tasks are grounded in a scene representation, an internal model of the robot's surroundings that also supports downstream applications such as navigation and simulation. This thesis addresses what form this representation should take and how it can be built and continuously updated from sensor measurements. Conventional scene representations such as point clouds, voxels, surfels, and meshes store geometry in explicit and discretized data structures, and have long formed the basis of modern robotic systems. However, their fixed-resolution discretization introduces information loss during fusion, and robots often need to maintain multiple redundant representations for different downstream tasks. Recent learning-based scene representations model scenes in a continuous and compact form, optimizing learnable parameters via differentiable functions to fit sensor observations. Yet many such methods are confined to small-scale offline operation, prone to catastrophic forgetting during long-term deployment, unable to accommodate loop closure corrections without remapping, and limited to a single sensor modality. To address these challenges, this thesis proposes a novel point-based implicit neural scene representation. It consists of sparse neural points, each with a local coordinate frame and an optimizable latent feature, decoded by globally shared neural networks into scene geometry, appearance, or object-level structure on demand. The representation is locally grounded, continuously learnable, scalable, and versatile. The thesis makes four main contributions that progressively extend this representation. The first contribution presents, to the best of our knowledge, the first full-fledged LiDAR SLAM system built entirely on a neural implicit representation. It integrates correspondence-free odometry, incremental map learning, and loop closure into a full pipeline, where loop closure corrections elastically deform the neural point map without remapping. The system runs at sensor frame rate and achieves state-of-the-art localization and surface reconstruction accuracy on established benchmarks while requiring less map memory than previous scene representations. While this system captures scene geometry, it cannot model appearance. The second contribution extends the representation to LiDAR-visual SLAM by jointly modeling a signed distance field and a Gaussian splatting radiance field in the same neural point map, where geometry and appearance mutually improve each other through a consistency loss. The system scales to kilometers-long trajectories while maintaining a globally consistent and compact map, achieving state-of-the-art photorealistic rendering and geometric reconstruction accuracy. Both contributions, however, treat the scene as a whole without modeling individual objects or their category specific structure. Thus, the third contribution addresses object-level modeling in cluttered horticultural environments, where occlusion makes per-instance shape estimation particularly challenging. It embeds learned shape priors into an occlusion-aware differentiable rendering pipeline that jointly estimates complete fruit geometry and pose from partial RGB-D observations, outperforming existing baselines in both shape completion and pose accuracy. The resulting models have been integrated into downstream robotic systems for autonomous fruit harvesting and safe leaf manipulation. All three contributions, however, operate within a single mapping session. To overcome this limitation, the fourth contribution addresses multi-session map fusion by casting multi-view point cloud registration as a single-stage 3D generation problem, bypassing exhaustive pairwise matching. A flow matching model with rigidity-enforcing sampling directly generates registered point clouds in a canonical frame from unposed multi-view inputs. Trained on large-scale cross-domain data, the model achieves top registration accuracy with the shortest runtime and scales favorably with view count. It also generalizes zero-shot across scene scales, sensor modalities, and low-overlap conditions, enabling downstream tasks such as offline SLAM and map merging. Together, these contributions establish neural point-based representations as a practical and unified foundation for SLAM and downstream tasks spanning geometry, appearance, object-level structure, and multi-view fusion. All methods achieve state-of-the-art results on established benchmarks, have been published in peer-reviewed venues, and are publicly available as open-source software.

bonndoc (University of Bonn)
University of Bonn (DE)
Openalex Percentile: Top 15%
Robotics and Sensor-Based Localization
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.