CineScope-Fuse: Cross-Scale Semantic Fusion for Cinematic Aesthetic Assessment of Multimodal-LLM-Generated Videos

Multimodal large language models and diffusion transformers have made text-to-video generation accessible to film production, advertising and social-media editing, but the assessment of generated video aesthetics remains poorly aligned with how viewers diagnose cinematic failure. A video may be sharp at the frame level while failing because of jitter, implausible object motion or mismatch between the prompt and the evolving event. We introduce CineScope-Fuse, a cross-scale semantic fusion network for cinematic aesthetic assessment of multimodal-LLM-generated videos. The method decomposes video quality into spatial fidelity, temporal stability and prompt–event alignment and then recomposes them through a semantic-aware module, a unified cross-attention network and a cross-scale spatio-temporal fusion block. The design turns the three common failure families of generated videos into explicit learning targets rather than treating them as unstructured regression noise. On the proposed Cine-AIGV benchmark and five public auxiliary benchmarks, CineScope-Fuse achieved the strongest overall correlation with human ratings, reaching 0.889 SRCC and 0.884 PLCC on Cine-AIGV. Ablations further showed that semantic alignment learning, cross-dimensional attention and multi-scale temporal fusion contributed complementary gains. These results indicate that cinematic video assessment benefits from jointly modeling low-level fidelity, temporal continuity and high-level narrative alignment, especially for synthetic clips whose visual realism and semantic plausibility fail in different ways.

Authors

Institutions

Publication Details

Journal
Applied Sciences
Published
2026-09-30
DOI
https://doi.org/10.3390/app16199710
Primary Topic
Visual Attention and Saliency Detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

CineScope-Fuse: Cross-Scale Semantic Fusion for Cinematic Aesthetic Assessment of Multimodal-LLM-Generated Videos

Yuheng Li, Yanxi Dan
Applied Sciences
Visual Attention and Saliency Detection
article

CineScope-Fuse: Cross-Scale Semantic Fusion for Cinematic Aesthetic Assessment of Multimodal-LLM-Generated Videos

Yuheng Li, Yanxi Dan
article en

Abstract

Multimodal large language models and diffusion transformers have made text-to-video generation accessible to film production, advertising and social-media editing, but the assessment of generated video aesthetics remains poorly aligned with how viewers diagnose cinematic failure. A video may be sharp at the frame level while failing because of jitter, implausible object motion or mismatch between the prompt and the evolving event. We introduce CineScope-Fuse, a cross-scale semantic fusion network for cinematic aesthetic assessment of multimodal-LLM-generated videos. The method decomposes video quality into spatial fidelity, temporal stability and prompt–event alignment and then recomposes them through a semantic-aware module, a unified cross-attention network and a cross-scale spatio-temporal fusion block. The design turns the three common failure families of generated videos into explicit learning targets rather than treating them as unstructured regression noise. On the proposed Cine-AIGV benchmark and five public auxiliary benchmarks, CineScope-Fuse achieved the strongest overall correlation with human ratings, reaching 0.889 SRCC and 0.884 PLCC on Cine-AIGV. Ablations further showed that semantic alignment learning, cross-dimensional attention and multi-scale temporal fusion contributed complementary gains. These results indicate that cinematic video assessment benefits from jointly modeling low-level fidelity, temporal continuity and high-level narrative alignment, especially for synthetic clips whose visual realism and semantic plausibility fail in different ways.

Applied SciencesVol. 16(19)
Southwest Petroleum University (CN), Hainan University (CN)
Openalex Percentile: Top 14%
Visual Attention and Saliency Detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.