Покрокове завдання-орієнтоване навчання планування руху колаборативного робота з 6 ступенями свободи за допомогою пакета Unity ML Agents

The object of research is the process of training 6 DOF Collaborative Robot motion planning. This is important since, for a robot with a few DOF, the training process by RL is relatively easier compare to robots with high degrees of freedom. One of the most problematic areas is the requirement of additional parameters for each DOF inside the RL makes the agents’ task harder. This can be best described as agent attempting to change multiple numerical float values sliders until the correct solution is found. In addition to this, multiple configurations can also lead to a potentially solution and confuse the agent. During the research, a potential solution of using multiple simpler agents was explored. As simulation environment, a repository from ROS and Unity collaboration with UR3E was chosen and modified to test the stepwise control strategy. This strategy itself consisted of two agents performing separate tasks. For the first agent’s task, rotating of the base joint is performed to align with the target object. For the second agent’s task, movement of shoulder and elbow to lower the gripper closer to the target is performed. Obtained results showed that the first agent learned rapidly to align the base with the object and correct position was reached each time and second agent also had success although with lower precision. The first agent achieved a mean reward of around 79 by step 500000 with time elapsed of around 420 seconds. The second agent archived a mean reward of around 33 by step 500000 with time elapsed of around 1055 seconds. Therefore, proving the potential of joint motion separation when training with RL. This method can be used in other robots to aid in faster and more precise learning. It can also be used on other Universal Robots models and robots made by other manufacturers.

Authors

Institutions

Publication Details

Journal
The Scientific Issues of Ternopil Volodymyr Hnatiuk National Pedagogical University Series pedagogy
Published
2026-08-31
Primary Topic
Teleoperation and Haptic Systems
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Покрокове завдання-орієнтоване навчання планування руху колаборативного робота з 6 ступенями свободи за допомогою пакета Unity ML Agents

Rufat Mammadzada, Latafat Gardasova
The Scientific Issues of Ternopil Volodymyr Hnatiuk National Pedagogical University Series pedagogy
Teleoperation and Haptic Systems
article

Покрокове завдання-орієнтоване навчання планування руху колаборативного робота з 6 ступенями свободи за допомогою пакета Unity ML Agents

Rufat Mammadzada, Latafat Gardasova
article en

Abstract

The object of research is the process of training 6 DOF Collaborative Robot motion planning. This is important since, for a robot with a few DOF, the training process by RL is relatively easier compare to robots with high degrees of freedom. One of the most problematic areas is the requirement of additional parameters for each DOF inside the RL makes the agents’ task harder. This can be best described as agent attempting to change multiple numerical float values sliders until the correct solution is found. In addition to this, multiple configurations can also lead to a potentially solution and confuse the agent. During the research, a potential solution of using multiple simpler agents was explored. As simulation environment, a repository from ROS and Unity collaboration with UR3E was chosen and modified to test the stepwise control strategy. This strategy itself consisted of two agents performing separate tasks. For the first agent’s task, rotating of the base joint is performed to align with the target object. For the second agent’s task, movement of shoulder and elbow to lower the gripper closer to the target is performed. Obtained results showed that the first agent learned rapidly to align the base with the object and correct position was reached each time and second agent also had success although with lower precision. The first agent achieved a mean reward of around 79 by step 500000 with time elapsed of around 420 seconds. The second agent archived a mean reward of around 33 by step 500000 with time elapsed of around 1055 seconds. Therefore, proving the potential of joint motion separation when training with RL. This method can be used in other robots to aid in faster and more precise learning. It can also be used on other Universal Robots models and robots made by other manufacturers.

The Scientific Issues of Ternopil Volodymyr Hnatiuk National Pedagogical University Series pedagogy
Azerbaijan State Oil and Industry University (AZ)
Peace, Justice and strong institutions
Openalex Percentile: Top 19%
Teleoperation and Haptic Systems
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.