Reduced-Precision Deployment Can Keep the First Mask and Still Change How Interactive Segmenters Correct

Edge deployments of promptable segmenters are checked by the quality of the first mask or by simulated click counts, yet an interactive tool is used through its corrections. We measure deployment on the correction itself: from the same image, click history and starting mask, we compare the next mask of a TensorRT engine with that of its PyTorch reference, splitting the change into the part of the starting error removed and the new error introduced. Across 6 segmenters on a Jetson Orin Nano and two benchmarks, the explicit INT8 engines of SAM 2.1-T, SAM 2.1-S, EdgeTAM and MobileSAM and the implicit INT8 engines of EdgeTAM, MobileSAM and EdgeSAM change both parts of the next correction, each under Holm control within its test family, and EfficientViT-SAM-L0 does so already at FP16 (+0.028 of the starting error removed and +0.056 introduced); the changes largely persist on objects whose first masks score alike, and simulated users need more clicks for most of these engines. For EfficientViT-SAM-L0 the change traces to 46 convolutions in the last two backbone stages: constraining only those to FP32 in an FP16-enabled engine removes 95% and 100% of the change on held-out data and keeps both mean differences of the repaired engine from the reference within ±0.01 of the starting error, at 20.9 ms against 13.8 ms for FP16. In the explicit INT8 engines most of the change arises in the image-encoder trunk: keeping the trunk at FP16, a rule chosen on two models, brings the corrections of EdgeTAM, SAM 2.1-S and SAM 2.1-T within the same margin of the reference’s and removes most of MobileSAM’s change, with a faster encoder than explicit INT8. In video propagation with SAM 2.1-T and EdgeTAM, the explicit INT8 engines’ masks diverge from the reference’s over the 30 frames after a correction beyond an implementation-only floor, with no detected change in mean region quality (a test that simulation on development sequences had given low power), and the trunk rule removes at least three quarters of that divergence; only SAM 2.1-T’s repaired engine then agrees with the reference within one point. A deployment check for an interactive segmenter should replay corrections, not only score the first output.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-05
DOI
https://doi.org/10.5281/zenodo.23156754
Primary Topic
Advanced Neural Network Applications
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Reduced-Precision Deployment Can Keep the First Mask and Still Change How Interactive Segmenters Correct

Ya-Fen Yeh, Guan-Yuan Chen
Zenodo (CERN European Organization for Nuclear Research)
Advanced Neural Network Applications
preprint

Reduced-Precision Deployment Can Keep the First Mask and Still Change How Interactive Segmenters Correct

Ya-Fen Yeh, Guan-Yuan Chen
preprint en

Abstract

Edge deployments of promptable segmenters are checked by the quality of the first mask or by simulated click counts, yet an interactive tool is used through its corrections. We measure deployment on the correction itself: from the same image, click history and starting mask, we compare the next mask of a TensorRT engine with that of its PyTorch reference, splitting the change into the part of the starting error removed and the new error introduced. Across 6 segmenters on a Jetson Orin Nano and two benchmarks, the explicit INT8 engines of SAM 2.1-T, SAM 2.1-S, EdgeTAM and MobileSAM and the implicit INT8 engines of EdgeTAM, MobileSAM and EdgeSAM change both parts of the next correction, each under Holm control within its test family, and EfficientViT-SAM-L0 does so already at FP16 (+0.028 of the starting error removed and +0.056 introduced); the changes largely persist on objects whose first masks score alike, and simulated users need more clicks for most of these engines. For EfficientViT-SAM-L0 the change traces to 46 convolutions in the last two backbone stages: constraining only those to FP32 in an FP16-enabled engine removes 95% and 100% of the change on held-out data and keeps both mean differences of the repaired engine from the reference within ±0.01 of the starting error, at 20.9 ms against 13.8 ms for FP16. In the explicit INT8 engines most of the change arises in the image-encoder trunk: keeping the trunk at FP16, a rule chosen on two models, brings the corrections of EdgeTAM, SAM 2.1-S and SAM 2.1-T within the same margin of the reference’s and removes most of MobileSAM’s change, with a faster encoder than explicit INT8. In video propagation with SAM 2.1-T and EdgeTAM, the explicit INT8 engines’ masks diverge from the reference’s over the 30 frames after a correction beyond an implementation-only floor, with no detected change in mean region quality (a test that simulation on development sequences had given low power), and the trunk rule removes at least three quarters of that divergence; only SAM 2.1-T’s repaired engine then agrees with the reference within one point. A deployment check for an interactive segmenter should replay corrections, not only score the first output.

Zenodo (CERN European Organization for Nuclear Research)
National Tsing Hua University (TW)
Advanced Neural Network Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.