A Comparative Analysis of Generative AI for Architectural Visualization: A Quantitative Comparison Using Objective Evaluation Metrics and the Architectural AI Score (AIS)
This study presents a quantitative metric-based comparison of four state-of-the-art text-to-image generative AI systems for architectural visualization: OpenAI Image Generation, Imagen 4, Midjourney v8.1, and Stable Image Ultra. The models were evaluated using the Architectural Prompt Dataset (APD-50), a controlled, expert-informed prompt set consisting of 50 architectural prompts across five typologies: Modern Residential Buildings, Public Buildings, Religious and Cultural Architecture, Interior Design, and Facade and Sustainable Architecture. For each prompt, four images were generated per model–prompt combination, resulting in a corpus of 800 images. The generated images were analyzed using reference-free and image-level computer vision metrics, including CLIP Score for prompt alignment, Raw BRISQUE and Quality_BRISQUE for reference-free perceptual quality, Shannon Entropy as a diagnostic descriptor of visual information density, LPIPS as intra-prompt perceptual variation, and SSIM as intra-prompt structural similarity. In response to the limitations of treating entropy as an inherently positive indicator, Shannon Entropy was excluded from the revised composite score and retained only as a diagnostic metric. Alongside the individual metric results, this study proposes the Revised Architectural AI Score (AIS-R) as a first-stage, exploratory composite index for summarizing selected visual-output characteristics under controlled prompt conditions. AIS-R combines normalized CLIP, Quality_BRISQUE, and LPIPS components using an explicit and reproducible weighting structure. The purpose of AIS-R is not to replace professional architectural judgment or to establish a fully validated architectural-quality framework at this stage. Rather, it provides an initial operational formulation for comparing prompt alignment, reference-free perceptual quality, and intra-prompt perceptual variation within the limited scope of the APD-50 experimental setup. Linear mixed-effects models with prompt-level random intercepts were used as the primary inferential framework. Under the revised metric configuration, Midjourney v8.1 achieved the highest AIS-R score (0.6504 ± 0.0586), followed by Stable Image Ultra (0.6267 ± 0.0582), Imagen 4 (0.6069 ± 0.0707), and OpenAI Image Generation (0.5829 ± 0.0555). However, these differences should be interpreted as metric-based differences in selected visual-output characteristics rather than as evidence of architectural quality, constructability, spatial validity, scale accuracy, or professional workflow suitability. The study therefore represents an initial metric-formulation and controlled benchmarking stage, while future research should examine the correspondence between AIS-R and professional architectural judgments through expert evaluation, user studies, or workflow-based validation.
Authors
- Volkan Cantemir (ORCID: https://orcid.org/0000-0002-6632-8151)
- Minel Kurtuluş (ORCID: https://orcid.org/0000-0003-4623-0613)
- Ece Cantemir
Institutions
- İstanbul Gelişim Üniversitesi (TR)
Publication Details
- Journal
- Applied Sciences
- Published
- 2026-10-06
- DOI
- https://doi.org/10.3390/app16199877
- Primary Topic
- Generative Adversarial Networks and Image Synthesis
- Type
- article
- Field-Weighted Citation Impact
- 0.00