Reinforcement learning in recommender systems: a comprehensive review

Reinforcement learning (RL) has emerged as a promising approach to upgrade recommender systems, particularly in dynamic and evolving user environments. Conventional approaches such as collaborative filtering and matrix factorization are hampered by static modeling, data sparsity, and short-term thinking. RL, however, provides an approach that allows systems to learn and evolve in response to user interactions and optimize for immediate and long-term user engagement. This review examines the use of RL in recommender systems through the examination of over 60 works of literature spanning three main approaches: model-free, model-based, and hybrid approaches. This work also highlights several gaps in previous reviews, namely their lack of coverage of deep RL techniques and lack of systematic taxonomies for each of the three approaches. An important discovery in this area is that traditional metrics for offline evaluation such as Precision@K and NDCG cannot account for the long-term performance of policies, showing an inherent disconnect between offline evaluation and actual implementation. Multi-agent reinforcement learning (MARL) and offline RL are two of the most significant emergent areas of study, with the former allowing for complex multi-stage recommendation processes and the latter ensuring data efficiency without needing live interaction.

Authors

Institutions

Publication Details

Journal
Artificial Intelligence Review
Published
2026-09-19
DOI
https://doi.org/10.1007/s10462-026-11706-3
Primary Topic
Recommender Systems and Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Reinforcement learning in recommender systems: a comprehensive review

Nikhil Parkar, Aditi Agale, M. Marimuthu
Artificial Intelligence Review
Recommender Systems and Techniques
article

Reinforcement learning in recommender systems: a comprehensive review

Nikhil Parkar, Aditi Agale, M. Marimuthu
article en

Abstract

Reinforcement learning (RL) has emerged as a promising approach to upgrade recommender systems, particularly in dynamic and evolving user environments. Conventional approaches such as collaborative filtering and matrix factorization are hampered by static modeling, data sparsity, and short-term thinking. RL, however, provides an approach that allows systems to learn and evolve in response to user interactions and optimize for immediate and long-term user engagement. This review examines the use of RL in recommender systems through the examination of over 60 works of literature spanning three main approaches: model-free, model-based, and hybrid approaches. This work also highlights several gaps in previous reviews, namely their lack of coverage of deep RL techniques and lack of systematic taxonomies for each of the three approaches. An important discovery in this area is that traditional metrics for offline evaluation such as Precision@K and NDCG cannot account for the long-term performance of policies, showing an inherent disconnect between offline evaluation and actual implementation. Multi-agent reinforcement learning (MARL) and offline RL are two of the most significant emergent areas of study, with the former allowing for complex multi-stage recommendation processes and the latter ensuring data efficiency without needing live interaction.

Artificial Intelligence Review
Vellore Institute of Technology University (IN)
Openalex Percentile: Top 4%
Recommender Systems and Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Reinforcement learning in recommender systems: a comprehensive review — Nikhil Parkar, Aditi Agale, et al. · Artificial Intelligence Review (2026) | TGRS Research Map | TGRS