HumaniBench: A Human-Centric Framework for Large Multimodal Models Evaluation
Although recent large multimodal models (LMMs) show impressive progress on vision–language tasks, their alignment with human-centered (HC) principles such as fairness, ethics, inclusivity, empathy, and robustness is often overlooked. Existing LMM benchmarks are largely accuracy-agnostic. We present HumaniBench, a unified framework for characterizing HC alignment across realistic, socially grounded visual contexts. It contains 32,000 expert-verified image–question pairs from real-world news imagery, each mapped to one or more HC principles through explicit metrics. Comparing 15 state-of-the-art LMMs reveals consistent trade-offs: proprietary systems lead on ethics, reasoning, and empathy, while open-source models show superior visual grounding and resilience. All models show persistent gaps in fairness and multilingual inclusivity. Chain-of-thought prompting and test-time scaling yield 8–12% gains on several HC dimensions. HumaniBench enables fine-grained analysis of alignment trade-offs not captured by conventional multimodal benchmarks. Project: https://vectorinstitute.github.io/humanibench/ Data: https://huggingface.co/vector-institute/HumaniBench Code: https://github.com/VectorInstitute/HumaniBench
Authors
- Amandeep Singh (ORCID: https://orcid.org/0000-0002-6371-0435)
- Vahid Reza Khazaie
- Mubarak Shah (ORCID: https://orcid.org/0000-0001-6172-5572)
- A. G. Hari Narayanan
- Ahmed Radwan
- Ashmal Vayani
- Mukund S. Chettiar
Institutions
- University of Central Florida (US)
- Vector Institute (CA)
Publication Details
- Journal
- ACM Transactions on Intelligent Systems and Technology
- Published
- 2026-09-07
- DOI
- https://doi.org/10.1145/3845999
- Primary Topic
- Multimodal Machine Learning Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00