Consistent Explainers or Unreliable Narrators: Systematic Differences in Consistency and Sensitivity Across Large Language Models for Group Recommendations

Group Recommender Systems (GRS) face the challenge of combining conflicting individual preferences. While Large Language Models (LLMs) are increasingly integrated into GRS, they may introduce risks regarding output inconsistency or sensitivity to group characteristics. In this paper, we analyze LLMs acting as both decision-makers and explanation generators for group recommendations. We contribute a novel dataset featuring fictitious group preferences, top-10 recommendations, and natural language explanations. Explanations are generated across multiple LLMs, testing differences between model families ( GPT-OSS vs. Mistral ) and model sizes. Our methodology evaluates recommendations and explanations against traditional social choice-based aggregation strategies across domains (abstract, low-stakes, high-stakes) and group configurations (uniform, divergent, coalitional, or minority). Findings reveal that the choice of LLM backbone is the primary driver of variance in recommendation quality and recommendation procedures detailed in the explanations. We found a clear divergence across model families. GPT-OSS provided consistent rankings, while Mistral exhibited more sensitivity to domain and group configurations. While main effects for domain and group configuration were not found, significant higher-order interactions suggest performance effects persist in specific combinations of configuration, domain, and aggregation strategy. We further discuss LLM roles in mitigating cold-start issues and group moderation, alongside drawbacks like transparency and privacy.

Authors

Institutions

Publication Details

Journal
ACM Transactions on Recommender Systems
Published
2026-09-22
DOI
https://doi.org/10.1145/3848640
Primary Topic
Recommender Systems and Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Consistent Explainers or Unreliable Narrators: Systematic Differences in Consistency and Sensitivity Across Large Language Models for Group Recommendations

Nava Tintarev, Cedric Waterschoot, Francesco Barile
ACM Transactions on Recommender Systems
Recommender Systems and Techniques
article

Consistent Explainers or Unreliable Narrators: Systematic Differences in Consistency and Sensitivity Across Large Language Models for Group Recommendations

Nava Tintarev, Cedric Waterschoot, Francesco Barile
article en

Abstract

Group Recommender Systems (GRS) face the challenge of combining conflicting individual preferences. While Large Language Models (LLMs) are increasingly integrated into GRS, they may introduce risks regarding output inconsistency or sensitivity to group characteristics. In this paper, we analyze LLMs acting as both decision-makers and explanation generators for group recommendations. We contribute a novel dataset featuring fictitious group preferences, top-10 recommendations, and natural language explanations. Explanations are generated across multiple LLMs, testing differences between model families ( GPT-OSS vs. Mistral ) and model sizes. Our methodology evaluates recommendations and explanations against traditional social choice-based aggregation strategies across domains (abstract, low-stakes, high-stakes) and group configurations (uniform, divergent, coalitional, or minority). Findings reveal that the choice of LLM backbone is the primary driver of variance in recommendation quality and recommendation procedures detailed in the explanations. We found a clear divergence across model families. GPT-OSS provided consistent rankings, while Mistral exhibited more sensitivity to domain and group configurations. While main effects for domain and group configuration were not found, significant higher-order interactions suggest performance effects persist in specific combinations of configuration, domain, and aggregation strategy. We further discuss LLM roles in mitigating cold-start issues and group moderation, alongside drawbacks like transparency and privacy.

ACM Transactions on Recommender Systems
Maastricht University (NL)
Peace, Justice and strong institutions
Openalex Percentile: Top 4%
Recommender Systems and Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Consistent Explainers or Unreliable Narrators: Systematic Differences in Consistency and Sensitivity Across Large Language Models for Group Recommendations — Nava Tintarev, Cedric Waterschoot, et al. · ACM Transactions on Recommender Systems (2026) | TGRS Research Map | TGRS