LLM Hallucination of Citations in Economics Persists With Web‐Enabled Models
ABSTRACT Despite significant advances since 2023 in large language model (LLM) capabilities and the introduction of web‐enabled search functionality, citation hallucination remains a persistent problem in AI‐generated academic content. We test GPT‐4o, DeepSeek V3, web‐enabled GPT‐4o, and web‐enabled GPT‐5.1 across prompts derived from Journal of Economic Literature categories. Using a hierarchical citation validation framework, we find that even an advanced web‐enabled model produces non‐existent citations. While web‐enabled models show improvement over their predecessors, the error rates remain high for scholarly applications. Our findings suggest that the fundamental challenge of citation hallucination persists despite technological advances, with implications for academic integrity and the reliable use of AI in research contexts. We provide an online tool for researchers to run LLMs through the benchmark test using our prompts.
Authors
- Stephen E. Hill (ORCID: https://orcid.org/0000-0002-9547-9098)
- Olga Shapoval (ORCID: https://orcid.org/0000-0003-3960-9641)
- Joy A. Buchanan (ORCID: https://orcid.org/0000-0002-5759-7368)
Institutions
- Samford University (US)
Publication Details
- Journal
- Southern Economic Journal
- Published
- 2026-10-08
- DOI
- https://doi.org/10.1002/soej.70069
- Primary Topic
- Topic Modeling
- Type
- article
- Field-Weighted Citation Impact
- 0.00