CMAE-Unet: A Study on a U-Net-Based Model for Semantic Segmentation of Unripe Tomato Images

In natural scenes, the high resemblance between unripe tomatoes and foliage, combined with severe occlusion, challenges current semantic segmentation models, causing poor accuracy and indistinct boundaries. To overcome this, we propose CMAE-UNet, a camouflage suppression and edge enhancement model based on the UNet architecture. Specifically, a Global-Local Integrated Spatial Attention (GLISA) encoder merges dual-branch dilated convolutions, residual structures, and an Efficient Multi-scale Attention mechanism to expand receptive fields and highlight targets in complex backgrounds. Furthermore, a Frequency-Domain Feature Enhancement (FFE) module leverages the Fast Fourier Transform to separate and adaptively enhance distinct frequency components, effectively mitigating camouflage interference. Additionally, a Directional Edge Enhancement (DEE) module uses three-directional learnable convolutions and spatial attention to sharpen indistinct target contours. Evaluated on a custom Tomato dataset encompassing five complex scenarios, CMAE-UNet outperforms 12 prominent methods in mIoU, Dice, and Sen metrics, yielding smoother and more precise segmentation boundaries. The model robustly withstands field interference, providing strong technological support for automated tomato detection, intelligent harvesting, and growth monitoring.

Authors

Institutions

Publication Details

Journal
AgriEngineering
Published
2026-08-25
DOI
https://doi.org/10.3390/agriengineering8090353
Primary Topic
Smart Agriculture and AI
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

CMAE-Unet: A Study on a U-Net-Based Model for Semantic Segmentation of Unripe Tomato Images

Jinfang Liu, Yongshen Liang, Y. Ye, Jianhua Zheng et al.
AgriEngineering
Smart Agriculture and AI
article

CMAE-Unet: A Study on a U-Net-Based Model for Semantic Segmentation of Unripe Tomato Images

Jinfang Liu, Yongshen Liang, Y. Ye, Jianhua Zheng, Zhaoxi Luo, Huanghui Zhao, Xiaoshan Ma, Jianru Chen, Guiming Huang
article en

Abstract

In natural scenes, the high resemblance between unripe tomatoes and foliage, combined with severe occlusion, challenges current semantic segmentation models, causing poor accuracy and indistinct boundaries. To overcome this, we propose CMAE-UNet, a camouflage suppression and edge enhancement model based on the UNet architecture. Specifically, a Global-Local Integrated Spatial Attention (GLISA) encoder merges dual-branch dilated convolutions, residual structures, and an Efficient Multi-scale Attention mechanism to expand receptive fields and highlight targets in complex backgrounds. Furthermore, a Frequency-Domain Feature Enhancement (FFE) module leverages the Fast Fourier Transform to separate and adaptively enhance distinct frequency components, effectively mitigating camouflage interference. Additionally, a Directional Edge Enhancement (DEE) module uses three-directional learnable convolutions and spatial attention to sharpen indistinct target contours. Evaluated on a custom Tomato dataset encompassing five complex scenarios, CMAE-UNet outperforms 12 prominent methods in mIoU, Dice, and Sen metrics, yielding smoother and more precise segmentation boundaries. The model robustly withstands field interference, providing strong technological support for automated tomato detection, intelligent harvesting, and growth monitoring.

AgriEngineeringVol. 8(9)
South China Agricultural University (CN), Zhongkai University of Agriculture and Engineering (CN)
Natural Science Foundation of Guangdong Province
No poverty
Openalex Percentile: Top 13%
Smart Agriculture and AI
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.