Meta-path guided policy distillation for resilient coordination in autonomous unmanned swarm
Abstract
Enhancing the resilience of Autonomous Unmanned Swarms (AUS) requires policies that remain effective under severe, structured disruptions while respecting the heterogeneous semantics of inter–subsystem interactions. Existing reinforcement learning (RL) approaches typically aggregate first–order neighborhoods in a path–agnostic manner, thereby blurring typed, ordered, and directed multi–hop dependencies encoded by domain meta–paths. We propose MPGPD-RC , a M eta- P ath G uided P olicy D istillation framework for R esilient C oordination that couples: (i) meta-path–guided embeddings learned by path-specific graph attention with contrastive reconstruction and attention fusion, and (ii) a teacher–student scheme in which a PPO teacher trained with a relaxed meta-path mask provides trajectories, and a student aligns both action distributions (KL) and trajectory-level structural codes via path-aware contrastive learning. Empirical evaluations validate that MPGPD-RC consistently surpasses state-of-the-art baselines across diverse perturbation scenarios by modeling complex, high-order dependencies that underpin resilient coordination.
Article Details
Authors (12)
Xingye Han
Huifang Wang
Qiang Jia
YingDong Gou
Bo Li
Jiancheng Liu
Sichuan Provincial Institute of Cultural Relics and Archaeology
Zaikun Han
Gang Hou
Key Laboratory for Green Chemical Technology of Ministry of Education, School of Chemical Engineering and Technology
Ke Li
Junxiong Ye
Yuqing Lin
Department of Pharmacy at the Second Affiliated Hospital, Harbin Medical University
Siwen Wei