Karafan Journal

Karafan Journal

Design and Implementation of a Task Allocation Algorithm for Formation Control and Role Assignment in Multi-Robot Systems Using Reinforcement Learning

Document Type : Original Article

Authors
Department of Electrical Engineering, Faculty of Electrical and Computer Engineering, Technical and Vocational University (TVU), Tehran, Iran
Abstract
This research focuses on developing a dynamic method for role assignment and formation control in multi-robot systems, aiming to enhance rotation and transition in fixed translational formations. In the proposed method, instead of separating the stages of task allocation and formation control, formation is considered as an integral part of the task allocation algorithm. This approach is implemented using Q-learning reinforcement learning, and multi-criteria decision-making (MCDM) theory. Robot behaviors dynamically determine their roles while simultaneously selecting their positions within the formation. To evaluate the algorithm's performance, simulations were conducted in various environments with grid sizes of 5×5, 8×8, and 20×20. The results demonstrate that the proposed method converges effectively and outperforms comparable methods. Furthermore, by employing a genetic algorithm, the key parameters of the MCDM method were optimized to enhance system efficiency. This research indicates that integrating reinforcement learning with multi-criteria decision-making, particularly in dynamic and uncertain environments, can offer an effective solution for formation control in multi-robot systems.
Keywords
Subjects

[1]    Schwab, K. (2023, May 31). The Fourth Industrial Revolution. Encyclopedia Britannica. https://www.britannica.com/topic/The-Fourth-Industrial-Revolution-2119734.
[2]    M. Dorigo, G. Theraulaz and V. Trianni, "Swarm Robotics: Past, Present, and Future [Point of View]," in Proceedings of the IEEE, vol. 109, no. 7, pp. 1152-1165, July 2021, https://doi.org/10.1109/JPROC.2021.3072740.
[3]    Wolf, D.F., Sukhatme, G.S. Mobile Robot Simultaneous Localization and Mapping in Dynamic Environments. Auton Robot 19, 53–65 (2005). https://doi.org/10.1007/s10514-005-0606-4.
[4]    Choi, D., & Kim, D. (2021). Intelligent Multi-Robot System for Collaborative Object Transportation Tasks in Rough Terrains. Electronics, 10(12), 1499. https://doi.org/10.3390/electronics10121499.
[1][5] Pedersen, S.M., Fountas, S., Sørensen, C.G., Van Evert, F.K., Blackmore, B.S. (2017). Robotic Seeding: Economic Perspectives. In: Pedersen, S., Lind, K. (eds) Precision Agriculture: Technology and Economic Perspectives. Progress in Precision Agriculture. Springer, Cham.. https://doi.org/10.1007/978-3-319-68715-5_8.
[6]    Murphy, R.R. et al. (2008). Search and Rescue Robotics. In: Siciliano, B., Khatib, O. (eds) Springer Handbook of Robotics. Springer, Berlin, Heidelberg. https://doi.org/10.1007/978-3-540-30301-5_51.
[7]    Ren, Wei. (2007). Multi-Vehicle consensus with a time-varying reference state. Systems & Control Letters. 56. 474-483. https://doi.org/10.1016/j.sysconle.2007.01.002.
[8]    Olfati-Saber, Reza. (2006). Flocking for Multi-Agent Dynamic Systems: Algorithms and Theory. Automatic Control, IEEE Transactions on. 51. 401 - 420. https://doi.org/10.1109/TAC.2005.864190.
[1][9] Şahin, E. (2022). Swarm robotics: From sources of inspiration to domains of application. In Swarm robotics (3342,10–20). Springer. https://doi.org/10.1007/978-3-540-30552-1_2.
[10]  J. Cortes, S. Martinez, T. Karatas and F. Bullo, "Coverage control for mobile sensing networks," in IEEE Transactions on Robotics and Automation, vol. 20, no. 2, pp. 243-255, April 2004, https://doi.org/10.1109/TRA.2004.824698.
[11]  McCreery, H. F., Bilek, J., Nagpal, R., & Breed, M. D. (2019). Effects of load mass and size on cooperative transport in ants over multiple transport challenges. The Journal of experimental biology, 222(Pt 17), jeb206821. https://doi.org/10.1242/jeb.206821.
[12]  De La Cruz, C., & Carelli, R. (2008). Dynamic model based formation control and obstacle avoidance of multi-robot systems. Robotica, 26(3), 345–356. https://doi.org/10.1017/S0263574707004092.
[13]  Mesbahi, M., & Egerstedt, M. (2010). Graph theoretic methods in multiagent networks, https://press.princeton.edu/books/hardcover/9780691140612/graph-theoretic-methods-in-multiagent-networks
[14]  Thrun, S., Burgard, W., & Fox, D. (2021). Probabilistic robotics (2nd ed.). MIT Press. https://mitpress.mit.edu/9780262201629/probabilistic-robotics.
[15]  Karim, N. A., & Ardestani, M. A. (2016, January). Takagi-Sugeno fuzzy formation control of nonholonomic robots. In 2016 4th International Conference on Control, Instrumentation, and Automation (ICCIA) (pp. 178-183). https://doi.org/10.1109/ICCIAutom.2016.7483157.
[16]  LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning for robot perception. Nature, 521(7553), 436–444. https://www.nature.com/articles/nature14539
[17]  Wang, Y., & de Silva, C. W. (2008). A machine-learning approach to multi-robot coordination. Engineering Applications of Artificial Intelligence, 21(3), 470-484.https://doi.org/10.1016/j.engappai.2007.05.006.
[18]  Mnih, V., Kavukcuoglu, K., & Silver, D. (2015). Human-level control through deep reinforcement learning. Nature, 518(7540), 529–533. https://www.nature.com/articles/nature14236.
[19]  Aykin, Can & Knopp, Martin & Dieopold, Klaus. (2018). Deep Reinforcement Learning for Formation Control. 1-5. https://doi.org/ 10.1109/ROMAN.2018.8525765.
[20]  Kamel, M. A., Yu, X., & Zhang, Y. (2020). Real-time fault-tolerant formation control of multiple WMRs based on hybrid GA–PSO algorithm. IEEE Transactions on Automation Science and Engineering, 18(3), 1263-1276, https://doi.org/10.1109/TASE.2020.3000507.
[21]  Azevedo, C., & Lima, P. U. (2025). Formal and scalable multi-robot coordination methods for long horizon tasks with time uncertainty. Robotics and Autonomous Systems, 105103. https://doi.org/10.1016/j.robot.2025.105103.
[22]  Shome, R., Solovey, K., Dobson, A. et al.(2020). dRRT*: Scalable and informed asymptotically-optimal multi-robot motion planning. Auton Robot 44, 443–467. https://doi.org/10.1007/s10514-019-09832-9.
[23]  Gilmartin, M. J. (2005). INTRODUCTION TO AUTONOMOUS MOBILE ROBOTS, by Roland Siegwart and Illah R. Nourbakhsh, MIT Press, 2004, xiii+321 pp., ISBN 0-262-19502-X. (Hardback, £27.95). Robotica, 23(2), 271–272. https://doi.org/10.1017/S0263574705221628.
[24]  Zhou, L., & Tokekar, P. (2021). Multi-robot coordination and planning in uncertain and adversarial environments. Current Robotics Reports, 2(2), 147-157. https://doi.org/10.1007/s43154-021-00046-5.
[25]  Kim, In-Cheol. (2006). Dynamic Role Assignment for Multi-agent Cooperation. 4263. 221-229. http://dx.doi.org/10.1007/11902140_25.
[26]  Xia, Y., Zhu, J. & Zhu, L. Dynamic role discovery and assignment in multi-agent task decomposition. Complex Intell. Syst. 9, 6211–6222 (2023). https://doi.org/10.1007/s40747-023-01071-x.
[27]  Paolo, G., Benechehab, A., Cherkaoui, H., Thomas, A., & K'egl, B. (2025). TAG: A Decentralized Framework for Multi-Agent Hierarchical Reinforcement Learning. ArXiv, abs/2502.15425.https://doi.org/10.48550/arXiv.2502.15425.
[28]  Du, W., Ding, S. (2021). A survey on multi-agent deep reinforcement learning: from the perspective of challenges and applications. Artif Intell Rev 54, 3215–3238. https://doi.org/10.1007/s10462-020-09938-y.
[1][29] Ghani Nori Alsaedi, A., Jalali Varnamkhasti, M., Jasim Mohammed, H., & Aghajani, M. (2025). Integrating Multi-Criteria Decision Analysis with Deep Reinforcement Learning: A Novel Framework for Intelligent Decision-Making in Iraqi Industries. International Journal of Mathematical Modelling & Computations, 14(2). https://doi.org/10.71932/ijm.2024.1197765.
[30]  Zhenglei He, Kim-Phuc Tran, Sebastien Thomassey, Xianyi Zeng, Jie Xu, Changhai Yi, (2021), A deep reinforcement learning based multi-criteria decision support system for optimizing textile chemical process, Computers in Industry, Volume 125, https://doi.org/10.1016/j.compind.2020.103373.
[31]  Deb, K., Pratap, A., Agarwal, S., & Meyarivan, T. (2002). A fast and elitist multiobjective genetic algorithm: NSGA‑II. IEEE Transactions on Evolutionary Computation, 6(2), 182–197. https://doi.org/10.1109/4235.996017.
[32]  Esmaeelinia ketabi A A, Pourkhaghan Shahrezaei Z, Danan Keshideh M. (2025). comparison of single-objective (GA) and multi-objective (NSGA-II) optimization methods in optimizing Iran's electricity generation portfolio.. QEER ; 21 (84) :235-266, URL: http://iiesj.ir/article-1-1605-en.html.
[33]  Chen, Q., Wang, R., Lyu, M., & Zhang, J. (2024). Transformer-Based Reinforcement Learning for Multi-Robot Autonomous Exploration. Sensors, 24(16), 5083. https://doi.org/10.3390/s24165083.
[34]  Y. Zhou, J. Xiao, Y. Zhou and G. Loianno, (2022), "Multi-Robot Collaborative Perception With Graph Neural Networks," in IEEE Robotics and Automation Letters, vol. 7, no. 2, pp. 2289-2296, https://doi.org/ 10.1109/LRA.2022.3141661.
[35]  Hafez, A.T., Givigi, S.N., & Yousefi, S. (2018). Unmanned Aerial Vehicles Formation Using Learning Based Model Predictive Control. Asian Journal of Control, 20, 1014 - 1026. https://doi.org/10.1002/asjc.1774
[36]  AlinaghizadehArdestani, M., & Vakili, A. (2020). Output feedback Controller design for HVAC system with delayed based Robust control approach. Karafan Quarterly Scientific Journal, 17(1), 85-95. https://doi.org/10.48301/kssa.2020.112758,
 (in persian) 
[37]  Noori, A. (2022). A New Method for Detecting Influential Nodes in Social Network Graphs Using Deep Learning Techniques. Karafan Journal, 19(1), 607-628. https://doi.org/ 10.48301/kssa.2022.310565.1786. (in persian) 
Volume 23, Issue 1
Technical and Engineering
Spring 2026
Pages 335-361

  • Receive Date 29 July 2025
  • Revise Date 16 December 2025
  • Accept Date 24 February 2026