Memory-Efficient 3D Scene Understanding via Quantized Multi-View Features

Projecting high-dimensional representations derived with 2D foundation models onto 3D point clouds via multi-view aggregation has emerged as a powerful paradigm for 3D scene understanding. However, the conventional strategy of storing dense feature vectors - derived by 2D foundation models at each point - introduces a substantial memory overhead that scales linearly with the scene size, thereby limiting scalability and hindering practical deployment in large-scale settings. In this work a memory-efficient framework is proposed to address this limitation through feature quantization. Specifically, projected 2D features are clustered into a compact set of prototypical embeddings, enabling each 3D point to be represented by a single discrete index instead of a full high-dimensional descriptor. This representation drastically reduces memory requirements by orders of magnitude while preserving the semantic richness of the original features. The proposed quantization framework is validated on diverse point clouds, demonstrating that the representations retain strong performance in downstream tasks while significantly improving computational efficiency.

Authors

Institutions

Publication Details

Journal
ISPRS annals of the photogrammetry, remote sensing and spatial information sciences
Published
2026-09-28
DOI
https://doi.org/10.5194/isprs-annals-xii-4-w1-2026-9-2026
Primary Topic
3D Shape Modeling and Analysis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Memory-Efficient 3D Scene Understanding via Quantized Multi-View Features

Fabio Remondino, Ashkan Alami
ISPRS annals of the photogrammetry, remote sensing and spatial information sciences
3D Shape Modeling and Analysis
article

Memory-Efficient 3D Scene Understanding via Quantized Multi-View Features

Fabio Remondino, Ashkan Alami
article en

Abstract

Projecting high-dimensional representations derived with 2D foundation models onto 3D point clouds via multi-view aggregation has emerged as a powerful paradigm for 3D scene understanding. However, the conventional strategy of storing dense feature vectors - derived by 2D foundation models at each point - introduces a substantial memory overhead that scales linearly with the scene size, thereby limiting scalability and hindering practical deployment in large-scale settings. In this work a memory-efficient framework is proposed to address this limitation through feature quantization. Specifically, projected 2D features are clustered into a compact set of prototypical embeddings, enabling each 3D point to be represented by a single discrete index instead of a full high-dimensional descriptor. This representation drastically reduces memory requirements by orders of magnitude while preserving the semantic richness of the original features. The proposed quantization framework is validated on diverse point clouds, demonstrating that the representations retain strong performance in downstream tasks while significantly improving computational efficiency.

ISPRS annals of the photogrammetry, remote sensing and spatial information sciencesVol. XII-4/W1-2026(0)
University of Trento (IT), Fondazione Bruno Kessler (IT)
Openalex Percentile: Top 14%
3D Shape Modeling and Analysis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Memory-Efficient 3D Scene Understanding via Quantized Multi-View Features — Fabio Remondino, Ashkan Alami · ISPRS annals of the photogrammetry, remote sensing and spatial information sciences (2026) | TGRS Research Map | TGRS