SMHA: A Style-Aware Multi-Scale and Hierarchical Alignment Framework for AIGC Image Quality Assessment

Rapid advances in text-to-image generation have made automatic quality assessment an essential component of model selection, content refinement, and system benchmarking. Unlike conventional distortions in natural images, the quality of artificial intelligence-generated images is jointly affected by visual factors such as blur, abnormal textures, and structural defects, as well as by artistic style, content authenticity, and text–image correspondence. Existing methods often struggle to distinguish legitimate artistic expression from genuine generation artifacts, capture local anomalies occurring at different visual scales, and model fine-grained text–image inconsistencies involving objects, attributes, numerosity, and spatial relations. To address these challenges, we propose SMHA, a Style-aware Multi-scale and Hierarchical Alignment model for multidimensional quality assessment of AI-generated images. SMHA comprises three core modules. First, a style-aware multi-scale quality representation module establishes a continuous style reference using external style prototypes and integrates hierarchical visual features to distinguish plausible stylistic variation from generation defects at different scales. Second, a prompt-guided hierarchical text–image alignment module follows a text-first, visual-verification interaction scheme to model fine-grained correspondence across multiple visual levels. Third, a task-specific multi-expert adaptive calibration module combines ordinal-semantic, local-structural, and Haar-frequency evidence and performs sample-adaptive residual calibration for the Quality and Authenticity tasks. Quality-assessment experiments on AGIQA-3K and AIGCIQA2023, with ArtBench-10 used for auxiliary style pretraining, demonstrate leading performance across multiple assessment dimensions and verify the effectiveness of the proposed multidimensional assessment design.

Authors

Institutions

Publication Details

Journal
Electronics
Published
2026-10-09
DOI
https://doi.org/10.3390/electronics15204594
Primary Topic
Image and Video Quality Assessment
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

SMHA: A Style-Aware Multi-Scale and Hierarchical Alignment Framework for AIGC Image Quality Assessment

Wenhao Li, Chengcheng Li, Huiying Xu, Hongdan Gu et al.
Electronics
Image and Video Quality Assessment
article

SMHA: A Style-Aware Multi-Scale and Hierarchical Alignment Framework for AIGC Image Quality Assessment

Wenhao Li, Chengcheng Li, Huiying Xu, Hongdan Gu, Xiuyuan Deng, Xinzhong Zhu
article en

Abstract

Rapid advances in text-to-image generation have made automatic quality assessment an essential component of model selection, content refinement, and system benchmarking. Unlike conventional distortions in natural images, the quality of artificial intelligence-generated images is jointly affected by visual factors such as blur, abnormal textures, and structural defects, as well as by artistic style, content authenticity, and text–image correspondence. Existing methods often struggle to distinguish legitimate artistic expression from genuine generation artifacts, capture local anomalies occurring at different visual scales, and model fine-grained text–image inconsistencies involving objects, attributes, numerosity, and spatial relations. To address these challenges, we propose SMHA, a Style-aware Multi-scale and Hierarchical Alignment model for multidimensional quality assessment of AI-generated images. SMHA comprises three core modules. First, a style-aware multi-scale quality representation module establishes a continuous style reference using external style prototypes and integrates hierarchical visual features to distinguish plausible stylistic variation from generation defects at different scales. Second, a prompt-guided hierarchical text–image alignment module follows a text-first, visual-verification interaction scheme to model fine-grained correspondence across multiple visual levels. Third, a task-specific multi-expert adaptive calibration module combines ordinal-semantic, local-structural, and Haar-frequency evidence and performs sample-adaptive residual calibration for the Quality and Authenticity tasks. Quality-assessment experiments on AGIQA-3K and AIGCIQA2023, with ArtBench-10 used for auxiliary style pretraining, demonstrate leading performance across multiple assessment dimensions and verify the effectiveness of the proposed multidimensional assessment design.

ElectronicsVol. 15(20)
Zhejiang Normal University (CN), Royal College of Art (GB)
Openalex Percentile: Top 15%
Image and Video Quality Assessment
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.