Underwater Computer Vision for Ecologists: A Framework for Curating Image Training Datasets for Object Detection

As marine ecosystems experience accelerating change, there is an urgent need for efficient and scalable biodiversity monitoring tools. We present a 10-step framework for integrating computer vision (CV) tools into long-term underwater biodiversity monitoring, using a case study from coastal British Columbia. Over 9000 h of unbaited remote underwater video footage were collected from two kelp farms and reference sites between March 2022 and June 2023. The framework includes steps for creating an annotated training dataset using unsupervised and supervised CV tools, culminating in the training and validation of a YOLOv8 object detection model. This process produced over 241,000 annotations across 54 pseudo-taxonomic categories (representing both taxa and visually similar groups of fauna), with a focus on fish and gelatinous zooplankton groups. The final model achieved an overall F1 (2 × precision × recall/(precision + recall)) of 0.74 and a mean average precision at 0.5 intersection over union (mAP50) of 0.78. The model had the highest performance on fine-resolution taxa such as Phanerodon vacca (F1 = 0.88) and Aurelia labiata (F1 = 0.90), and lowest performance on pseudo-taxa with limited visual distinctiveness such as Actinopterygii (F1 = 0.60) and Cnidaria (F1 = 0.60). A re-training experiment using annotation thresholds between 25 training images to full dataset availability (~200–24,000 images per group) found that model performance was positively correlated with annotation effort, with F1 averaging 0.84 and mAP50 averaging 0.91 at the maximum training dataset size. Our results suggest that the model is most effective for abundant and visually distinctive taxa, while performance declines for groups with coarse taxonomic resolution. We recommend optimizing annotation effort by targeting genus- or species-level taxa, having at least one broad-level group to capture order-level abundances, and supplementing annotations of rare groups with common but morphologically similar groups, which may further improve model performance.

Authors

Institutions

Publication Details

Journal
Sensors
Published
2026-09-16
DOI
https://doi.org/10.3390/s26185869
Primary Topic
Water Quality Monitoring Technologies
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Underwater Computer Vision for Ecologists: A Framework for Curating Image Training Datasets for Object Detection

Declan McIntosh, Alexandra Branzan Albu, Francis Juanes, Talen Rimmer et al.
Sensors
Water Quality Monitoring Technologies
article

Underwater Computer Vision for Ecologists: A Framework for Curating Image Training Datasets for Object Detection

Declan McIntosh, Alexandra Branzan Albu, Francis Juanes, Talen Rimmer, Colin Bates, Tom Zhang
article en

Abstract

As marine ecosystems experience accelerating change, there is an urgent need for efficient and scalable biodiversity monitoring tools. We present a 10-step framework for integrating computer vision (CV) tools into long-term underwater biodiversity monitoring, using a case study from coastal British Columbia. Over 9000 h of unbaited remote underwater video footage were collected from two kelp farms and reference sites between March 2022 and June 2023. The framework includes steps for creating an annotated training dataset using unsupervised and supervised CV tools, culminating in the training and validation of a YOLOv8 object detection model. This process produced over 241,000 annotations across 54 pseudo-taxonomic categories (representing both taxa and visually similar groups of fauna), with a focus on fish and gelatinous zooplankton groups. The final model achieved an overall F1 (2 × precision × recall/(precision + recall)) of 0.74 and a mean average precision at 0.5 intersection over union (mAP50) of 0.78. The model had the highest performance on fine-resolution taxa such as Phanerodon vacca (F1 = 0.88) and Aurelia labiata (F1 = 0.90), and lowest performance on pseudo-taxa with limited visual distinctiveness such as Actinopterygii (F1 = 0.60) and Cnidaria (F1 = 0.60). A re-training experiment using annotation thresholds between 25 training images to full dataset availability (~200–24,000 images per group) found that model performance was positively correlated with annotation effort, with F1 averaging 0.84 and mAP50 averaging 0.91 at the maximum training dataset size. Our results suggest that the model is most effective for abundant and visually distinctive taxa, while performance declines for groups with coarse taxonomic resolution. We recommend optimizing annotation effort by targeting genus- or species-level taxa, having at least one broad-level group to capture order-level abundances, and supplementing annotations of rare groups with common but morphologically similar groups, which may further improve model performance.

SensorsVol. 26(18)
University of Victoria (CA), ASL Environmental Sciences (Canada) (CA)
Life below water
Openalex Percentile: Top 21%
Water Quality Monitoring Technologies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.