Please wait ...
0% Complete
Home
/
11th International Symposium on Telecommunication (IST'2024)
Modified Double-DQN: addressing stability
Authors :
Shervin Halat
1
Mohammad Mehdi Ebadzadeh
2
Kiana Amani
3
1- Department of Computer Engineering, Amirkabir University of Technology (Polytechnique)
2- Department of Computer Engineering, Amirkabir University of Technology (Polytechnique)
3- sharif university of technology
Keywords :
stability،overestimation،moving targets،deep reinforcement learning،Q-learning،DQN،Double-DQN
Abstract :
Inspired by Double Q-learning algorithm, the Double-DQN (DDQN) algorithm was originally proposed in order to address the overestimation issue in the original DQN algorithm. The DDQN has successfully shown both theoretically and empirically the importance of decoupling in terms of action evaluation and selection in computation of target values; although, all the benefits were acquired with only a simple adaption to DQN algorithm, minimal possible change as it was mentioned by the authors. Nevertheless, there seems a roll-back in the proposed algorithm of DDQN since the parameters of policy network are emerged again in the target value function which were initially withdrawn by DQN with the hope of tackling the serious issue of moving targets and the instability caused by it (i.e., by moving targets) in the process of learning. Therefore, in this paper three modifications to the DDQN algorithm are proposed with the hope of maintaining the performance in the terms of both stability and overestimation. These modifications are focused on the logic of decoupling the best action selection and evaluation in the target value function and the logic of tackling the moving targets issue. Each of these modifications have their own pros and cons compared to the others. The mentioned pros and cons mainly refer to the execution time required for the corresponding algorithm and the stability provided by the corresponding algorithm. Also, in terms of overestimation, none of the modifications seem to underperform compared to the original DDQN if not outperform it. With the intention of evaluating the efficacy of the proposed modifications, multiple empirical experiments along with theoretical experiments were conducted. The results obtained are represented and discussed in this article.
Papers List
List of archived papers
Analyzing the Challenges of Applying Artificial Intelligence in the Field of Employment and Offering Solutions to Improve Skills and Workforce Training
Azam sadat Mortazavi Kahangi - Hassan Yeganeh - Anita Hadizadeh
Improve Quality of Systems via Dynamic Cost-Aware Selection of Web-Services
Ali Jahani - Fattaneh Taghiyareh
LSTM-based Framework for 5G Resource Allocation Prediction
Amin Pourmahbobi - Hamed Tabrizchi
Shot Noise Optimal Filters for Phase-Diversity Homodyne Receivers
Mohammad Sheikhbahaei - Amir R. Forouzan
The cost of renting a fiber network from one operator to another in Iran
Reza Mehri - Davoud Ranjbarrafie - Elham Ziaeipour
A Novel Bidirectional Distributed Cross Domain Solution Architecture
Sajjad Ansariyan - Mohammadali Doostari
An Efficient Motor Imagery BCI Classification Architecture Using Non-linear Network Tuning suitable for Cloud-Edge Environments
Abolfazl Danayi - Hamid Soltanian-Zadeh
NOMA-SWIPT optimization with PSO and machine learning algorithm
Arshia Saeidifar - Khashayar Saremi - Bahareh Akhbari
Key Performance Indicators in Information Sharing and Analysis Centers (ISAC)
Fatemeh Valinezhad - Afsaneh Madani
Cryptocurrency volatility prediction based on price, return and volatility cross-correlation using LSTM
Masoud Omidvari Abarghouie - Sasan H. Alizadeh - Ahmad Khademzadeh
more
Samin Hamayesh - Version 44.9.0