A Comparative Analysis of Generative AI for Architectural Visualization: A Quantitative Comparison Using Objective Evaluation Metrics and the Architectural AI Score (AIS)

This study presents a quantitative metric-based comparison of four state-of-the-art text-to-image generative AI systems for architectural visualization: OpenAI Image Generation, Imagen 4, Midjourney v8.1, and Stable Image Ultra. The models were evaluated using the Architectural Prompt Dataset (APD-50), a controlled, expert-informed prompt set consisting of 50 architectural prompts across five typologies: Modern Residential Buildings, Public Buildings, Religious and Cultural Architecture, Interior Design, and Facade and Sustainable Architecture. For each prompt, four images were generated per model–prompt combination, resulting in a corpus of 800 images. The generated images were analyzed using reference-free and image-level computer vision metrics, including CLIP Score for prompt alignment, Raw BRISQUE and Quality_BRISQUE for reference-free perceptual quality, Shannon Entropy as a diagnostic descriptor of visual information density, LPIPS as intra-prompt perceptual variation, and SSIM as intra-prompt structural similarity. In response to the limitations of treating entropy as an inherently positive indicator, Shannon Entropy was excluded from the revised composite score and retained only as a diagnostic metric. Alongside the individual metric results, this study proposes the Revised Architectural AI Score (AIS-R) as a first-stage, exploratory composite index for summarizing selected visual-output characteristics under controlled prompt conditions. AIS-R combines normalized CLIP, Quality_BRISQUE, and LPIPS components using an explicit and reproducible weighting structure. The purpose of AIS-R is not to replace professional architectural judgment or to establish a fully validated architectural-quality framework at this stage. Rather, it provides an initial operational formulation for comparing prompt alignment, reference-free perceptual quality, and intra-prompt perceptual variation within the limited scope of the APD-50 experimental setup. Linear mixed-effects models with prompt-level random intercepts were used as the primary inferential framework. Under the revised metric configuration, Midjourney v8.1 achieved the highest AIS-R score (0.6504 ± 0.0586), followed by Stable Image Ultra (0.6267 ± 0.0582), Imagen 4 (0.6069 ± 0.0707), and OpenAI Image Generation (0.5829 ± 0.0555). However, these differences should be interpreted as metric-based differences in selected visual-output characteristics rather than as evidence of architectural quality, constructability, spatial validity, scale accuracy, or professional workflow suitability. The study therefore represents an initial metric-formulation and controlled benchmarking stage, while future research should examine the correspondence between AIS-R and professional architectural judgments through expert evaluation, user studies, or workflow-based validation.

Authors

Institutions

Publication Details

Journal
Applied Sciences
Published
2026-10-06
DOI
https://doi.org/10.3390/app16199877
Primary Topic
Generative Adversarial Networks and Image Synthesis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

A Comparative Analysis of Generative AI for Architectural Visualization: A Quantitative Comparison Using Objective Evaluation Metrics and the Architectural AI Score (AIS)

Volkan Cantemir, Minel Kurtuluş, Ece Cantemir
Applied Sciences
Generative Adversarial Networks and Image Synthesis
article

A Comparative Analysis of Generative AI for Architectural Visualization: A Quantitative Comparison Using Objective Evaluation Metrics and the Architectural AI Score (AIS)

Volkan Cantemir, Minel Kurtuluş, Ece Cantemir
article en

Abstract

This study presents a quantitative metric-based comparison of four state-of-the-art text-to-image generative AI systems for architectural visualization: OpenAI Image Generation, Imagen 4, Midjourney v8.1, and Stable Image Ultra. The models were evaluated using the Architectural Prompt Dataset (APD-50), a controlled, expert-informed prompt set consisting of 50 architectural prompts across five typologies: Modern Residential Buildings, Public Buildings, Religious and Cultural Architecture, Interior Design, and Facade and Sustainable Architecture. For each prompt, four images were generated per model–prompt combination, resulting in a corpus of 800 images. The generated images were analyzed using reference-free and image-level computer vision metrics, including CLIP Score for prompt alignment, Raw BRISQUE and Quality_BRISQUE for reference-free perceptual quality, Shannon Entropy as a diagnostic descriptor of visual information density, LPIPS as intra-prompt perceptual variation, and SSIM as intra-prompt structural similarity. In response to the limitations of treating entropy as an inherently positive indicator, Shannon Entropy was excluded from the revised composite score and retained only as a diagnostic metric. Alongside the individual metric results, this study proposes the Revised Architectural AI Score (AIS-R) as a first-stage, exploratory composite index for summarizing selected visual-output characteristics under controlled prompt conditions. AIS-R combines normalized CLIP, Quality_BRISQUE, and LPIPS components using an explicit and reproducible weighting structure. The purpose of AIS-R is not to replace professional architectural judgment or to establish a fully validated architectural-quality framework at this stage. Rather, it provides an initial operational formulation for comparing prompt alignment, reference-free perceptual quality, and intra-prompt perceptual variation within the limited scope of the APD-50 experimental setup. Linear mixed-effects models with prompt-level random intercepts were used as the primary inferential framework. Under the revised metric configuration, Midjourney v8.1 achieved the highest AIS-R score (0.6504 ± 0.0586), followed by Stable Image Ultra (0.6267 ± 0.0582), Imagen 4 (0.6069 ± 0.0707), and OpenAI Image Generation (0.5829 ± 0.0555). However, these differences should be interpreted as metric-based differences in selected visual-output characteristics rather than as evidence of architectural quality, constructability, spatial validity, scale accuracy, or professional workflow suitability. The study therefore represents an initial metric-formulation and controlled benchmarking stage, while future research should examine the correspondence between AIS-R and professional architectural judgments through expert evaluation, user studies, or workflow-based validation.

Applied SciencesVol. 16(19)
İstanbul Gelişim Üniversitesi (TR)
Openalex Percentile: Top 15%
Generative Adversarial Networks and Image Synthesis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.