Model-free inverse reinforcement learning for linear discrete-time deterministic systems
Designing appropriate objective functions for autonomous agents is a challenging task that often involves cumbersome reward engineering. Inverse reinforcement learning (IRL) provides an alternative by inferring the underlying cost structure from observed behaviours. This paper introduces two novel model-free IRL techniques tailored to linear discrete-time deterministic systems. The first is an online method that leverages a single trajectory, enhanced with exploratory inputs, to quickly approximate the expert's objective. The second is an offline iterative approach that begins with a single dataset for parameter estimation and, if necessary, incorporates additional datasets to satisfy certain conditions, all while eliminating the need to repeatedly solve a forward control problem. Notable innovations of our work include a novel formulation of the discrete-time Hamilton-Jacobi-Bellman equation for online estimation and a new controller representation that facilitates model-free solution development in offline estimation. Numerical simulations validate the effectiveness of the proposed methods, demonstrating that the estimated cost functions produce learner policies that closely match those of the expert.
Authors
- Mohan Rajesh Elara (ORCID: https://orcid.org/0000-0001-6504-1530)
- Bùi Vũ Minh (ORCID: https://orcid.org/0000-0003-0222-166X)
- Anh Vu Le (ORCID: https://orcid.org/0000-0002-4804-7540)
- Hamed Jabbari Asl (ORCID: https://orcid.org/0000-0002-3040-8539)
Institutions
- Izmir Institute of Technology (TR)
- Ton Duc Thang University (VN)
- Urmia University of Technology (IR)
- Singapore University of Technology and Design (SG)
- Trường ĐH Nguyễn Tất Thành (VN)
Publication Details
- Journal
- International Journal of Systems Science
- Published
- 2026-09-29
- DOI
- https://doi.org/10.1080/00207721.2026.2737394
- Primary Topic
- Adaptive Dynamic Programming Control
- Type
- article
- Field-Weighted Citation Impact
- 0.00