Human-like content analysis for generative AI with language-grounded sparse encoders

The rapid development of generative AI has transformed content creation, communication, and human development. However, this technology raises profound concerns in high-stakes domains, demanding rigorous methods to analyze and evaluate AI-generated content. While existing analytic methods often treat images as indivisible wholes, real-world AI failures generally manifest as specific visual patterns that can evade holistic detection and suit more granular and decomposed analysis. Here we introduce a content analysis tool, Language-Grounded Sparse Encoders (LanSE), which decomposes images into interpretable visual patterns with natural language descriptions. Utilizing interpretability modules and large multimodal models, LanSE can automatically identify visual patterns within data modalities. Our method discovers >5000 visual patterns with 93% human agreement, provides decomposed evaluation that outperforms existing methods, establishes the first systematic evaluation of physical plausibility, and extends to medical imaging settings. Our method’s capability to extract language-grounded patterns can be naturally adapted to numerous fields, including biology and geography, as well as other data modalities such as protein structures and time series, thereby advancing content analysis for generative AI.

Authors

Institutions

Publication Details

Journal
npj Artificial Intelligence
Published
2026-10-07
DOI
https://doi.org/10.1038/s44387-026-00152-9
Primary Topic
Digital Media Forensic Detection
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Human-like content analysis for generative AI with language-grounded sparse encoders

Qiran Zou, Trang Nguyen, Dianbo Liu, Ehsan Adeli et al.
npj Artificial Intelligence
Digital Media Forensic Detection
article

Human-like content analysis for generative AI with language-grounded sparse encoders

Qiran Zou, Trang Nguyen, Dianbo Liu, Ehsan Adeli, Srinivas Anumasa, Ching‐Yu Cheng, Yiming Tang, Yingtao Zhu, Yih Chung Tham, Arash Lagzian, Yilun Du, Zhang Ye
article en

Abstract

The rapid development of generative AI has transformed content creation, communication, and human development. However, this technology raises profound concerns in high-stakes domains, demanding rigorous methods to analyze and evaluate AI-generated content. While existing analytic methods often treat images as indivisible wholes, real-world AI failures generally manifest as specific visual patterns that can evade holistic detection and suit more granular and decomposed analysis. Here we introduce a content analysis tool, Language-Grounded Sparse Encoders (LanSE), which decomposes images into interpretable visual patterns with natural language descriptions. Utilizing interpretability modules and large multimodal models, LanSE can automatically identify visual patterns within data modalities. Our method discovers >5000 visual patterns with 93% human agreement, provides decomposed evaluation that outperforms existing methods, establishes the first systematic evaluation of physical plausibility, and extends to medical imaging settings. Our method’s capability to extract language-grounded patterns can be naturally adapted to numerous fields, including biology and geography, as well as other data modalities such as protein structures and time series, thereby advancing content analysis for generative AI.

npj Artificial Intelligence
Harvard University (US), National University of Singapore (SG), Harvard University Press (US), Stanford University (US), Tsinghua University (CN)
Openalex Percentile: Top 99%
Digital Media Forensic Detection
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.