INT8 Deployment Changes How Interactive Segmenters Correct, Even When the First Mask Is Just as Good

Promptable segmenters such as SAM 2 reach edge devices as reduced-precision inference engines, and the conversion is checked the way these models are evaluated: by the quality of the first mask, or by the number of clicks a simulated user needs. An interactive tool is used through its corrections, however, and neither check isolates them. We test deployment on the correction itself. From the same image, click history and starting mask, we compare the next mask of a TensorRT engine with that of the FP32 model and split the difference into starting error removed and new error introduced. On Berkeley and DAVIS (428 objects in 127 source images and videos), with SAM 2.1-T and EdgeTAM on a Jetson Orin Nano, FP16 engines change either quantity on average by at most about 1.1 × 10^−3 of the starting error. Explicit INT8 engines change both: EdgeTAM’s next correction removes an extra 7.3% of the starting error but introduces an extra 15.0%, and SAM 2.1-T’s removes an extra 5.5% and introduces an extra 6.7% (Holm-adjusted p ≤ 0.0012 over 12 tests). The changes keep their size on the objects whose first masks score within 0.01 IoU of the FP32 model’s, so first-output fidelity does not certify correction fidelity. In closed loop, EdgeTAM’s explicit INT8 engine changes the clicks a replayed user needs to reach 0.95 IoU by +2.62 and the rate of not reaching it within 20 clicks by +13.8 percentage points; SAM 2.1-T’s changes the clicks needed to reach 0.90 IoU by +1.29 when the model chooses where to ask. On this device the INT8 engines are about as fast as FP16 or slower (0.99–1.20 times its latency). Deployment checks for interactive segmentation should test the correction, not only the first output.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-29
DOI
https://doi.org/10.5281/zenodo.23033971
Primary Topic
Advanced Neural Network Applications
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

INT8 Deployment Changes How Interactive Segmenters Correct, Even When the First Mask Is Just as Good

Ya-Fen Yeh, Guan-Yuan Chen
Zenodo (CERN European Organization for Nuclear Research)
Advanced Neural Network Applications
preprint

INT8 Deployment Changes How Interactive Segmenters Correct, Even When the First Mask Is Just as Good

Ya-Fen Yeh, Guan-Yuan Chen
preprint en

Abstract

Promptable segmenters such as SAM 2 reach edge devices as reduced-precision inference engines, and the conversion is checked the way these models are evaluated: by the quality of the first mask, or by the number of clicks a simulated user needs. An interactive tool is used through its corrections, however, and neither check isolates them. We test deployment on the correction itself. From the same image, click history and starting mask, we compare the next mask of a TensorRT engine with that of the FP32 model and split the difference into starting error removed and new error introduced. On Berkeley and DAVIS (428 objects in 127 source images and videos), with SAM 2.1-T and EdgeTAM on a Jetson Orin Nano, FP16 engines change either quantity on average by at most about 1.1 × 10^−3 of the starting error. Explicit INT8 engines change both: EdgeTAM’s next correction removes an extra 7.3% of the starting error but introduces an extra 15.0%, and SAM 2.1-T’s removes an extra 5.5% and introduces an extra 6.7% (Holm-adjusted p ≤ 0.0012 over 12 tests). The changes keep their size on the objects whose first masks score within 0.01 IoU of the FP32 model’s, so first-output fidelity does not certify correction fidelity. In closed loop, EdgeTAM’s explicit INT8 engine changes the clicks a replayed user needs to reach 0.95 IoU by +2.62 and the rate of not reaching it within 20 clicks by +13.8 percentage points; SAM 2.1-T’s changes the clicks needed to reach 0.90 IoU by +1.29 when the model chooses where to ask. On this device the INT8 engines are about as fast as FP16 or slower (0.99–1.20 times its latency). Deployment checks for interactive segmentation should test the correction, not only the first output.

Zenodo (CERN European Organization for Nuclear Research)
National Tsing Hua University (TW)
Advanced Neural Network Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

INT8 Deployment Changes How Interactive Segmenters Correct, Even When the First Mask Is Just as Good — Ya-Fen Yeh, Guan-Yuan Chen · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS