A multimodal deep learning-driven model for assessing drawing composition quality

In recent years, digital art education, computerized evaluation systems, and aesthetic computing have prioritized drawing composition quality assessment. Traditional assessments have relied on subjective judgment or single-modal analysis, limiting scalability and uniformity. Existing methods use visual characteristics or rule-based models, which cannot incorporate multimodal inputs, including visual layout and verbal descriptions or rubrics, restricting score validity. This research proposes that multimodal Deep Learning-Driven Modal(MM-DLN) be used to create a reliable and intelligent automatic drawing composition assessment framework. Visual and textual data are integrated into a single model to increase scoring precision and interpretability. The technique uses MM-DLN, a Multimodal Transformer-Based Scoring Network. It uses a Vision Transformer (ViT) to extract spatial and compositional aspects from the drawing and a Large Language Model Meta AI (Version 3) (LLaMA3) to analyze textual input. A cross-modal attention mechanism fuses these modalities, and a regression-based scoring system generates the score. The model highlights attention regions, revealing its score logic. The proposed MM-DLN model reduced MAE by 23.9% and improved PCC by 10.2% compared to the best-performing baseline, demonstrating enhanced prediction accuracy and more substantial alignment with expert-driven aesthetic evaluation. Performance was stable with noisy or stylistically diversified inputs. In conclusion, MM-DLN uses multimodal deep learning to assess drawing composition quality automatically in an accurate, scalable, and interpretable method.

Authors

Institutions

Publication Details

Journal
Journal Of Big Data
Published
2026-10-09
DOI
https://doi.org/10.1186/s40537-026-01551-0
Primary Topic
Multimodal Machine Learning Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

A multimodal deep learning-driven model for assessing drawing composition quality

Yunhe Su, Xiaowen Feng
Journal Of Big Data
Multimodal Machine Learning Applications
article

A multimodal deep learning-driven model for assessing drawing composition quality

Yunhe Su, Xiaowen Feng
article en

Abstract

In recent years, digital art education, computerized evaluation systems, and aesthetic computing have prioritized drawing composition quality assessment. Traditional assessments have relied on subjective judgment or single-modal analysis, limiting scalability and uniformity. Existing methods use visual characteristics or rule-based models, which cannot incorporate multimodal inputs, including visual layout and verbal descriptions or rubrics, restricting score validity. This research proposes that multimodal Deep Learning-Driven Modal(MM-DLN) be used to create a reliable and intelligent automatic drawing composition assessment framework. Visual and textual data are integrated into a single model to increase scoring precision and interpretability. The technique uses MM-DLN, a Multimodal Transformer-Based Scoring Network. It uses a Vision Transformer (ViT) to extract spatial and compositional aspects from the drawing and a Large Language Model Meta AI (Version 3) (LLaMA3) to analyze textual input. A cross-modal attention mechanism fuses these modalities, and a regression-based scoring system generates the score. The model highlights attention regions, revealing its score logic. The proposed MM-DLN model reduced MAE by 23.9% and improved PCC by 10.2% compared to the best-performing baseline, demonstrating enhanced prediction accuracy and more substantial alignment with expert-driven aesthetic evaluation. Performance was stable with noisy or stylistically diversified inputs. In conclusion, MM-DLN uses multimodal deep learning to assess drawing composition quality automatically in an accurate, scalable, and interpretable method.

Journal Of Big Data
Yonsei University (KR), Hongik University (KR)
Openalex Percentile: Top 15%
Multimodal Machine Learning Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

A multimodal deep learning-driven model for assessing drawing composition quality — Yunhe Su, Xiaowen Feng · Journal Of Big Data (2026) | TGRS Research Map | TGRS