An efficient on-policy actor–critic architecture for adaptive bitrate streaming using deep reinforcement learning

Abstract Adaptive video streaming is a widely used way of distributing video material over the Internet. It employs adaptive bitrate (ABR) algorithms to ensure a high quality of experience (QoE) for end users under a variety of network situations. However, the dynamic and varied characteristics of emerging 5th generation (5G) and beyond 5G (B5G) networks frequently causes instability in ABR algorithms. Despite advances in learning-based ABR algorithms, particularly those that use cutting-edge deep reinforcement learning (DRL) technology, they typically achieve mediocre performance and struggle to offer optimal results consistently across a variety of network environments. Recent observations reveal that Comyco, a quality-aware adaptive video streaming technology, provides great QoE over 3G networks but suffers limits across 4G and 5G networks. To solve this issue, we offer GRAdient Sharing actor-critic approach with Policy optimization (GRASP), a revolutionary on-policy DRL-based ABR technique. We compare GRASP to various on-policy and off-policy DRL techniques such as Pensieve, GRAAB, Comyco, PPO-ABR and SAC-ABR and other cutting-edge DRL-based ABR algorithms and non DRL-based ABR algorithms including BB, RB and BOLA. For comparison, we utilize mm Wave 4G traces and Lumos 5G dataset. Our results show that the proposed GRASP algorithm outperforms Comyco and other cutting-edge DRL-based ABR methods under a variety of network situations. GRASP consistently achieves higher QoE over the Comyco DRL-based framework up to 11.44%, and 16.56% for 5G and 4G networks respectively, and even higher QoE for other DRL-based and non DRL-based ABR schemes. Meanwhile, GRASP reduces the stall ratio compared to Comyco up to 67% and 70% for 5G and 4G networks respectively.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-04
DOI
https://doi.org/10.1038/s41598-026-69078-1
Primary Topic
Image and Video Quality Assessment
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

An efficient on-policy actor–critic architecture for adaptive bitrate streaming using deep reinforcement learning

Paresh Saxena, Manik Gupta, Mandan Naresh
Scientific Reports
Image and Video Quality Assessment
article

An efficient on-policy actor–critic architecture for adaptive bitrate streaming using deep reinforcement learning

Paresh Saxena, Manik Gupta, Mandan Naresh
article en

Abstract

Abstract Adaptive video streaming is a widely used way of distributing video material over the Internet. It employs adaptive bitrate (ABR) algorithms to ensure a high quality of experience (QoE) for end users under a variety of network situations. However, the dynamic and varied characteristics of emerging 5th generation (5G) and beyond 5G (B5G) networks frequently causes instability in ABR algorithms. Despite advances in learning-based ABR algorithms, particularly those that use cutting-edge deep reinforcement learning (DRL) technology, they typically achieve mediocre performance and struggle to offer optimal results consistently across a variety of network environments. Recent observations reveal that Comyco, a quality-aware adaptive video streaming technology, provides great QoE over 3G networks but suffers limits across 4G and 5G networks. To solve this issue, we offer GRAdient Sharing actor-critic approach with Policy optimization (GRASP), a revolutionary on-policy DRL-based ABR technique. We compare GRASP to various on-policy and off-policy DRL techniques such as Pensieve, GRAAB, Comyco, PPO-ABR and SAC-ABR and other cutting-edge DRL-based ABR algorithms and non DRL-based ABR algorithms including BB, RB and BOLA. For comparison, we utilize mm Wave 4G traces and Lumos 5G dataset. Our results show that the proposed GRASP algorithm outperforms Comyco and other cutting-edge DRL-based ABR methods under a variety of network situations. GRASP consistently achieves higher QoE over the Comyco DRL-based framework up to 11.44%, and 16.56% for 5G and 4G networks respectively, and even higher QoE for other DRL-based and non DRL-based ABR schemes. Meanwhile, GRASP reduces the stall ratio compared to Comyco up to 67% and 70% for 5G and 4G networks respectively.

Scientific Reports
Manipal Academy of Higher Education (IN), Birla Institute of Technology and Science - Hyderabad Campus (IN), Birla Institute of Technology and Science, Pilani (IN)
Openalex Percentile: Top 12%
Image and Video Quality Assessment
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.