Human-like content analysis for generative AI with language-grounded sparse encoders
The rapid development of generative AI has transformed content creation, communication, and human development. However, this technology raises profound concerns in high-stakes domains, demanding rigorous methods to analyze and evaluate AI-generated content. While existing analytic methods often treat images as indivisible wholes, real-world AI failures generally manifest as specific visual patterns that can evade holistic detection and suit more granular and decomposed analysis. Here we introduce a content analysis tool, Language-Grounded Sparse Encoders (LanSE), which decomposes images into interpretable visual patterns with natural language descriptions. Utilizing interpretability modules and large multimodal models, LanSE can automatically identify visual patterns within data modalities. Our method discovers >5000 visual patterns with 93% human agreement, provides decomposed evaluation that outperforms existing methods, establishes the first systematic evaluation of physical plausibility, and extends to medical imaging settings. Our method’s capability to extract language-grounded patterns can be naturally adapted to numerous fields, including biology and geography, as well as other data modalities such as protein structures and time series, thereby advancing content analysis for generative AI.
Authors
- Qiran Zou
- Trang Nguyen (ORCID: https://orcid.org/0009-0000-6025-8474)
- Dianbo Liu (ORCID: https://orcid.org/0000-0002-3042-9161)
- Ehsan Adeli (ORCID: https://orcid.org/0000-0002-0579-7763)
- Srinivas Anumasa
- Ching‐Yu Cheng (ORCID: https://orcid.org/0000-0003-0655-885X)
- Yiming Tang (ORCID: https://orcid.org/0000-0003-2378-8972)
- Yingtao Zhu (ORCID: https://orcid.org/0000-0002-7210-9385)
- Yih Chung Tham (ORCID: https://orcid.org/0000-0002-6752-797X)
- Arash Lagzian
- Yilun Du
- Zhang Ye (ORCID: https://orcid.org/0000-0003-3495-5596)
Institutions
- Harvard University (US)
- National University of Singapore (SG)
- Harvard University Press (US)
- Stanford University (US)
- Tsinghua University (CN)
Publication Details
- Journal
- npj Artificial Intelligence
- Published
- 2026-10-07
- DOI
- https://doi.org/10.1038/s44387-026-00152-9
- Primary Topic
- Digital Media Forensic Detection
- Type
- article
- Field-Weighted Citation Impact
- 0.00