Automatic and User-Centered Evaluation of an AI-Driven Korean Recipe RAG Chatbot: Alignment, Mismatch, and Actionability
Users with limited cooking experience must simultaneously consider multiple factors, including available ingredients, cooking time, cooking equipment, health conditions, personal preferences, and practical feasibility. Therefore, simple keyword-based recipe searches make it difficult for such users to identify suitable options. This study analyzed the alignment and mismatch between automated evaluation using Retrieval-Augmented Generation Assessment (RAGAS) and user evaluation for a Korean-language recipe chatbot based on Retrieval-Augmented Generation (RAG) and examined the evaluation dimensions that explain user-perceived quality in recommendation-oriented recipe responses. To this end, three system versions were configured: Basic RAG (E0), Retrieval-enhanced RAG (E1), and Prompt-constrained RAG (E2). A total of 180 responses generated for 60 queries categorized as recommendation, procedural, storage, substitution, and quantitative queries were compared. Each response was evaluated using the RAGAS metrics Answer Relevancy and Faithfulness, five user-evaluation dimensions, and three metrics proposed in this study: Decision Support Score (DSS), Constraint Satisfaction Score (CSS), and Information Sufficiency Score (IS). The results showed that evaluation patterns differed across system configurations, query types, and evaluation dimensions, without consistent improvements across all queries. Furthermore, neither RAGAS metric was significantly correlated with user-evaluation scores in the 171-response complete-case analysis. Cases in which automated evaluation scores were high but user-evaluation scores were low occurred primarily among recommendation queries. In particular, cases with high Answer Relevancy but low Actionability showed that responses relevant to the question are not necessarily easy for users to act upon. In contrast, DSS and IS showed significant positive correlations with all five user-evaluation dimensions, whereas CSS showed significant positive correlations with Correctness, Usefulness, and Satisfaction, suggesting that decision support, constraint satisfaction, and information sufficiency may serve as important complementary dimensions for explaining user-perceived quality in recommendation-oriented recipe responses. These findings suggest that the two RAGAS metrics examined in this study—Answer Relevancy and Faithfulness—do not, on their own, fully explain user-perceived quality when evaluating recipe RAG systems, and that query-type-specific analysis and user-centric evaluation should be considered together.
Authors
- Youngjung Suh (ORCID: https://orcid.org/0009-0001-3762-9609)
- SunJae Jeong (ORCID: https://orcid.org/0009-0003-7148-8779)
Institutions
- Kongju National University (KR)
Publication Details
- Journal
- Electronics
- Published
- 2026-09-17
- DOI
- https://doi.org/10.3390/electronics15184231
- Primary Topic
- AI in Service Interactions
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- Ministry of Science and ICT, South Korea