Automatic and User-Centered Evaluation of an AI-Driven Korean Recipe RAG Chatbot: Alignment, Mismatch, and Actionability

Users with limited cooking experience must simultaneously consider multiple factors, including available ingredients, cooking time, cooking equipment, health conditions, personal preferences, and practical feasibility. Therefore, simple keyword-based recipe searches make it difficult for such users to identify suitable options. This study analyzed the alignment and mismatch between automated evaluation using Retrieval-Augmented Generation Assessment (RAGAS) and user evaluation for a Korean-language recipe chatbot based on Retrieval-Augmented Generation (RAG) and examined the evaluation dimensions that explain user-perceived quality in recommendation-oriented recipe responses. To this end, three system versions were configured: Basic RAG (E0), Retrieval-enhanced RAG (E1), and Prompt-constrained RAG (E2). A total of 180 responses generated for 60 queries categorized as recommendation, procedural, storage, substitution, and quantitative queries were compared. Each response was evaluated using the RAGAS metrics Answer Relevancy and Faithfulness, five user-evaluation dimensions, and three metrics proposed in this study: Decision Support Score (DSS), Constraint Satisfaction Score (CSS), and Information Sufficiency Score (IS). The results showed that evaluation patterns differed across system configurations, query types, and evaluation dimensions, without consistent improvements across all queries. Furthermore, neither RAGAS metric was significantly correlated with user-evaluation scores in the 171-response complete-case analysis. Cases in which automated evaluation scores were high but user-evaluation scores were low occurred primarily among recommendation queries. In particular, cases with high Answer Relevancy but low Actionability showed that responses relevant to the question are not necessarily easy for users to act upon. In contrast, DSS and IS showed significant positive correlations with all five user-evaluation dimensions, whereas CSS showed significant positive correlations with Correctness, Usefulness, and Satisfaction, suggesting that decision support, constraint satisfaction, and information sufficiency may serve as important complementary dimensions for explaining user-perceived quality in recommendation-oriented recipe responses. These findings suggest that the two RAGAS metrics examined in this study—Answer Relevancy and Faithfulness—do not, on their own, fully explain user-perceived quality when evaluating recipe RAG systems, and that query-type-specific analysis and user-centric evaluation should be considered together.

Authors

Institutions

Publication Details

Journal
Electronics
Published
2026-09-17
DOI
https://doi.org/10.3390/electronics15184231
Primary Topic
AI in Service Interactions
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Automatic and User-Centered Evaluation of an AI-Driven Korean Recipe RAG Chatbot: Alignment, Mismatch, and Actionability

Youngjung Suh, SunJae Jeong
Electronics
AI in Service Interactions
article

Automatic and User-Centered Evaluation of an AI-Driven Korean Recipe RAG Chatbot: Alignment, Mismatch, and Actionability

Youngjung Suh, SunJae Jeong
article en

Abstract

Users with limited cooking experience must simultaneously consider multiple factors, including available ingredients, cooking time, cooking equipment, health conditions, personal preferences, and practical feasibility. Therefore, simple keyword-based recipe searches make it difficult for such users to identify suitable options. This study analyzed the alignment and mismatch between automated evaluation using Retrieval-Augmented Generation Assessment (RAGAS) and user evaluation for a Korean-language recipe chatbot based on Retrieval-Augmented Generation (RAG) and examined the evaluation dimensions that explain user-perceived quality in recommendation-oriented recipe responses. To this end, three system versions were configured: Basic RAG (E0), Retrieval-enhanced RAG (E1), and Prompt-constrained RAG (E2). A total of 180 responses generated for 60 queries categorized as recommendation, procedural, storage, substitution, and quantitative queries were compared. Each response was evaluated using the RAGAS metrics Answer Relevancy and Faithfulness, five user-evaluation dimensions, and three metrics proposed in this study: Decision Support Score (DSS), Constraint Satisfaction Score (CSS), and Information Sufficiency Score (IS). The results showed that evaluation patterns differed across system configurations, query types, and evaluation dimensions, without consistent improvements across all queries. Furthermore, neither RAGAS metric was significantly correlated with user-evaluation scores in the 171-response complete-case analysis. Cases in which automated evaluation scores were high but user-evaluation scores were low occurred primarily among recommendation queries. In particular, cases with high Answer Relevancy but low Actionability showed that responses relevant to the question are not necessarily easy for users to act upon. In contrast, DSS and IS showed significant positive correlations with all five user-evaluation dimensions, whereas CSS showed significant positive correlations with Correctness, Usefulness, and Satisfaction, suggesting that decision support, constraint satisfaction, and information sufficiency may serve as important complementary dimensions for explaining user-perceived quality in recommendation-oriented recipe responses. These findings suggest that the two RAGAS metrics examined in this study—Answer Relevancy and Faithfulness—do not, on their own, fully explain user-perceived quality when evaluating recipe RAG systems, and that query-type-specific analysis and user-centric evaluation should be considered together.

ElectronicsVol. 15(18)
Kongju National University (KR)
Ministry of Science and ICT, South Korea
Openalex Percentile: Top 9%
AI in Service Interactions
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.