فصلنامه علمی کارافن

فصلنامه علمی کارافن

طراحی و پیاده سازی یک الگوریتم تخصیص وظایف برای کنترل آرایش‌بندی و تخصیص نقش‌ها در سامانه‌های چند رباتی با استفاده از یادگیری تقویتی

نوع مقاله : مقاله پژوهشی (کاربردی)

نویسندگان
گروه مهندسی برق، دانشکده برق و کامپیوتر، دانشگاه ملی مهارت، تهران، ایران.
چکیده
این پژوهش به توسعه یک روش پویا برای تخصیص نقش و کنترل آرایش‌بندی در سیستم‌های چند رباتی، با هدف بهبود چرخش و انتقال در آرایش‌های ثابت انتقالی می‌پردازد. در روش پیشنهادی، به جای جداسازی مراحل تخصیص وظایف و کنترل آرایش‌بندی، آرایش‌بندی به عنوان بخشی از الگوریتم تخصیص وظایف در نظر گرفته شده است. این رویکرد با استفاده از یادگیری تقویتی Q و نظریه تصمیم‌گیری چندمعیاره (MCDM) پیاده‌سازی شده است. رفتارهای ربات‌ها به صورت پویا نقش‌های خود را تعیین می‌کنند و همزمان موقعیت آرایش‌بندی را نیز مشخص می‌نمایند. برای ارزیابی عملکرد الگوریتم، از شبیه‌سازی در محیط‌های مختلف با ابعاد ۵×۵، ۸×۸ و ۲۰×۲۰ استفاده شده است. نتایج نشان می‌دهد که روش پیشنهادی به‌خوبی همگرا شده و در مقایسه با روش‌های مشابه، عملکرد بهتری دارد. همچنین، با بهره‌گیری از الگوریتم ژنتیک، پارامترهای کلیدی روش MCDM بهینه‌سازی شده‌اند تا کارایی سیستم افزایش یابد. این پژوهش نشان می‌دهد که ادغام یادگیری تقویتی با نظریه تصمیم‌گیری چندمعیاره، به‌ویژه در محیط‌های پویا و نامشخص می‌تواند راه‌حلی مؤثر برای کنترل آرایش‌بندی در سیستم‌های چند رباتی ارائه دهد.
کلیدواژه‌ها
موضوعات

عنوان مقاله English

Design and Implementation of a Task Allocation Algorithm for Formation Control and Role Assignment in Multi-Robot Systems Using Reinforcement Learning

نویسندگان English

Mahdi Ardestani
Mahdi siavash
Department of Electrical Engineering, Faculty of Electrical and Computer Engineering, Technical and Vocational University (TVU), Tehran, Iran
چکیده English

This research focuses on developing a dynamic method for role assignment and formation control in multi-robot systems, aiming to enhance rotation and transition in fixed translational formations. In the proposed method, instead of separating the stages of task allocation and formation control, formation is considered as an integral part of the task allocation algorithm. This approach is implemented using Q-learning reinforcement learning, and multi-criteria decision-making (MCDM) theory. Robot behaviors dynamically determine their roles while simultaneously selecting their positions within the formation. To evaluate the algorithm's performance, simulations were conducted in various environments with grid sizes of 5×5, 8×8, and 20×20. The results demonstrate that the proposed method converges effectively and outperforms comparable methods. Furthermore, by employing a genetic algorithm, the key parameters of the MCDM method were optimized to enhance system efficiency. This research indicates that integrating reinforcement learning with multi-criteria decision-making, particularly in dynamic and uncertain environments, can offer an effective solution for formation control in multi-robot systems.

کلیدواژه‌ها English

Role assignment
formation control
multi-robot systems
Q-learning
multi-criteria decision-making
genetic algorithm
[1]    Schwab, K. (2023, May 31). The Fourth Industrial Revolution. Encyclopedia Britannica. https://www.britannica.com/topic/The-Fourth-Industrial-Revolution-2119734.
[2]    M. Dorigo, G. Theraulaz and V. Trianni, "Swarm Robotics: Past, Present, and Future [Point of View]," in Proceedings of the IEEE, vol. 109, no. 7, pp. 1152-1165, July 2021, https://doi.org/10.1109/JPROC.2021.3072740.
[3]    Wolf, D.F., Sukhatme, G.S. Mobile Robot Simultaneous Localization and Mapping in Dynamic Environments. Auton Robot 19, 53–65 (2005). https://doi.org/10.1007/s10514-005-0606-4.
[4]    Choi, D., & Kim, D. (2021). Intelligent Multi-Robot System for Collaborative Object Transportation Tasks in Rough Terrains. Electronics, 10(12), 1499. https://doi.org/10.3390/electronics10121499.
[1][5] Pedersen, S.M., Fountas, S., Sørensen, C.G., Van Evert, F.K., Blackmore, B.S. (2017). Robotic Seeding: Economic Perspectives. In: Pedersen, S., Lind, K. (eds) Precision Agriculture: Technology and Economic Perspectives. Progress in Precision Agriculture. Springer, Cham.. https://doi.org/10.1007/978-3-319-68715-5_8.
[6]    Murphy, R.R. et al. (2008). Search and Rescue Robotics. In: Siciliano, B., Khatib, O. (eds) Springer Handbook of Robotics. Springer, Berlin, Heidelberg. https://doi.org/10.1007/978-3-540-30301-5_51.
[7]    Ren, Wei. (2007). Multi-Vehicle consensus with a time-varying reference state. Systems & Control Letters. 56. 474-483. https://doi.org/10.1016/j.sysconle.2007.01.002.
[8]    Olfati-Saber, Reza. (2006). Flocking for Multi-Agent Dynamic Systems: Algorithms and Theory. Automatic Control, IEEE Transactions on. 51. 401 - 420. https://doi.org/10.1109/TAC.2005.864190.
[1][9] Şahin, E. (2022). Swarm robotics: From sources of inspiration to domains of application. In Swarm robotics (3342,10–20). Springer. https://doi.org/10.1007/978-3-540-30552-1_2.
[10]  J. Cortes, S. Martinez, T. Karatas and F. Bullo, "Coverage control for mobile sensing networks," in IEEE Transactions on Robotics and Automation, vol. 20, no. 2, pp. 243-255, April 2004, https://doi.org/10.1109/TRA.2004.824698.
[11]  McCreery, H. F., Bilek, J., Nagpal, R., & Breed, M. D. (2019). Effects of load mass and size on cooperative transport in ants over multiple transport challenges. The Journal of experimental biology, 222(Pt 17), jeb206821. https://doi.org/10.1242/jeb.206821.
[12]  De La Cruz, C., & Carelli, R. (2008). Dynamic model based formation control and obstacle avoidance of multi-robot systems. Robotica, 26(3), 345–356. https://doi.org/10.1017/S0263574707004092.
[13]  Mesbahi, M., & Egerstedt, M. (2010). Graph theoretic methods in multiagent networks, https://press.princeton.edu/books/hardcover/9780691140612/graph-theoretic-methods-in-multiagent-networks
[14]  Thrun, S., Burgard, W., & Fox, D. (2021). Probabilistic robotics (2nd ed.). MIT Press. https://mitpress.mit.edu/9780262201629/probabilistic-robotics.
[15]  Karim, N. A., & Ardestani, M. A. (2016, January). Takagi-Sugeno fuzzy formation control of nonholonomic robots. In 2016 4th International Conference on Control, Instrumentation, and Automation (ICCIA) (pp. 178-183). https://doi.org/10.1109/ICCIAutom.2016.7483157.
[16]  LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning for robot perception. Nature, 521(7553), 436–444. https://www.nature.com/articles/nature14539
[17]  Wang, Y., & de Silva, C. W. (2008). A machine-learning approach to multi-robot coordination. Engineering Applications of Artificial Intelligence, 21(3), 470-484.https://doi.org/10.1016/j.engappai.2007.05.006.
[18]  Mnih, V., Kavukcuoglu, K., & Silver, D. (2015). Human-level control through deep reinforcement learning. Nature, 518(7540), 529–533. https://www.nature.com/articles/nature14236.
[19]  Aykin, Can & Knopp, Martin & Dieopold, Klaus. (2018). Deep Reinforcement Learning for Formation Control. 1-5. https://doi.org/ 10.1109/ROMAN.2018.8525765.
[20]  Kamel, M. A., Yu, X., & Zhang, Y. (2020). Real-time fault-tolerant formation control of multiple WMRs based on hybrid GA–PSO algorithm. IEEE Transactions on Automation Science and Engineering, 18(3), 1263-1276, https://doi.org/10.1109/TASE.2020.3000507.
[21]  Azevedo, C., & Lima, P. U. (2025). Formal and scalable multi-robot coordination methods for long horizon tasks with time uncertainty. Robotics and Autonomous Systems, 105103. https://doi.org/10.1016/j.robot.2025.105103.
[22]  Shome, R., Solovey, K., Dobson, A. et al.(2020). dRRT*: Scalable and informed asymptotically-optimal multi-robot motion planning. Auton Robot 44, 443–467. https://doi.org/10.1007/s10514-019-09832-9.
[23]  Gilmartin, M. J. (2005). INTRODUCTION TO AUTONOMOUS MOBILE ROBOTS, by Roland Siegwart and Illah R. Nourbakhsh, MIT Press, 2004, xiii+321 pp., ISBN 0-262-19502-X. (Hardback, £27.95). Robotica, 23(2), 271–272. https://doi.org/10.1017/S0263574705221628.
[24]  Zhou, L., & Tokekar, P. (2021). Multi-robot coordination and planning in uncertain and adversarial environments. Current Robotics Reports, 2(2), 147-157. https://doi.org/10.1007/s43154-021-00046-5.
[25]  Kim, In-Cheol. (2006). Dynamic Role Assignment for Multi-agent Cooperation. 4263. 221-229. http://dx.doi.org/10.1007/11902140_25.
[26]  Xia, Y., Zhu, J. & Zhu, L. Dynamic role discovery and assignment in multi-agent task decomposition. Complex Intell. Syst. 9, 6211–6222 (2023). https://doi.org/10.1007/s40747-023-01071-x.
[27]  Paolo, G., Benechehab, A., Cherkaoui, H., Thomas, A., & K'egl, B. (2025). TAG: A Decentralized Framework for Multi-Agent Hierarchical Reinforcement Learning. ArXiv, abs/2502.15425.https://doi.org/10.48550/arXiv.2502.15425.
[28]  Du, W., Ding, S. (2021). A survey on multi-agent deep reinforcement learning: from the perspective of challenges and applications. Artif Intell Rev 54, 3215–3238. https://doi.org/10.1007/s10462-020-09938-y.
[1][29] Ghani Nori Alsaedi, A., Jalali Varnamkhasti, M., Jasim Mohammed, H., & Aghajani, M. (2025). Integrating Multi-Criteria Decision Analysis with Deep Reinforcement Learning: A Novel Framework for Intelligent Decision-Making in Iraqi Industries. International Journal of Mathematical Modelling & Computations, 14(2). https://doi.org/10.71932/ijm.2024.1197765.
[30]  Zhenglei He, Kim-Phuc Tran, Sebastien Thomassey, Xianyi Zeng, Jie Xu, Changhai Yi, (2021), A deep reinforcement learning based multi-criteria decision support system for optimizing textile chemical process, Computers in Industry, Volume 125, https://doi.org/10.1016/j.compind.2020.103373.
[31]  Deb, K., Pratap, A., Agarwal, S., & Meyarivan, T. (2002). A fast and elitist multiobjective genetic algorithm: NSGA‑II. IEEE Transactions on Evolutionary Computation, 6(2), 182–197. https://doi.org/10.1109/4235.996017.
[32]  Esmaeelinia ketabi A A, Pourkhaghan Shahrezaei Z, Danan Keshideh M. (2025). comparison of single-objective (GA) and multi-objective (NSGA-II) optimization methods in optimizing Iran's electricity generation portfolio.. QEER ; 21 (84) :235-266, URL: http://iiesj.ir/article-1-1605-en.html.
[33]  Chen, Q., Wang, R., Lyu, M., & Zhang, J. (2024). Transformer-Based Reinforcement Learning for Multi-Robot Autonomous Exploration. Sensors, 24(16), 5083. https://doi.org/10.3390/s24165083.
[34]  Y. Zhou, J. Xiao, Y. Zhou and G. Loianno, (2022), "Multi-Robot Collaborative Perception With Graph Neural Networks," in IEEE Robotics and Automation Letters, vol. 7, no. 2, pp. 2289-2296, https://doi.org/ 10.1109/LRA.2022.3141661.
[35]  Hafez, A.T., Givigi, S.N., & Yousefi, S. (2018). Unmanned Aerial Vehicles Formation Using Learning Based Model Predictive Control. Asian Journal of Control, 20, 1014 - 1026. https://doi.org/10.1002/asjc.1774
[36]  AlinaghizadehArdestani, M., & Vakili, A. (2020). Output feedback Controller design for HVAC system with delayed based Robust control approach. Karafan Quarterly Scientific Journal, 17(1), 85-95. https://doi.org/10.48301/kssa.2020.112758,
 (in persian) 
[37]  Noori, A. (2022). A New Method for Detecting Influential Nodes in Social Network Graphs Using Deep Learning Techniques. Karafan Journal, 19(1), 607-628. https://doi.org/ 10.48301/kssa.2022.310565.1786. (in persian) 
دوره 23، شماره 1
فنی و مهندسی
بهار 1405
صفحه 335-361

  • تاریخ دریافت 07 مرداد 1404
  • تاریخ بازنگری 25 آذر 1404
  • تاریخ پذیرش 05 اسفند 1404