Execution uncertainties,such as motion delay or confrontation in pursuit-evasion problems with rapidly changing states,affect the task performance of multi-Unmanned Aerial Vehicle(UAV)systems.This may lead to the fail...Execution uncertainties,such as motion delay or confrontation in pursuit-evasion problems with rapidly changing states,affect the task performance of multi-Unmanned Aerial Vehicle(UAV)systems.This may lead to the failure of the initial task assignment scheme.To address this problem,this paper takes the interception scenario as a typical case.It proposes a distributed dynamic task assignment algorithm based on an evolving task performance model to reassign UAVs to tasks in an event-triggered manner.This paper combines the underlying execution model with the interception effectiveness model to design the evolving task performance model.This model describes the UAV task performance in a finite interval by predicting and integrating the states and actions of the intercepted UAVs and the targets under execution uncertainty.The discrete task monitor triggers reassignment based on the severity of the task performance deviation.The Consensus-based Auction Algorithm(CBAA)is extended to optimize the task performance function and efficiently give the reassignment scheme.Simulation results demonstrate the feasibility and effectiveness of the proposed algorithm.展开更多
In multiple Unmanned Aerial Vehicles(UAV)systems,achieving efficient navigation is essential for executing complex tasks and enhancing autonomy.Traditional navigation methods depend on predefined control strategies an...In multiple Unmanned Aerial Vehicles(UAV)systems,achieving efficient navigation is essential for executing complex tasks and enhancing autonomy.Traditional navigation methods depend on predefined control strategies and trajectory planning and often perform poorly in complex environments.To improve the UAV-environment interaction efficiency,this study proposes a multi-UAV integrated navigation algorithm based on Deep Reinforcement Learning(DRL).This algorithm integrates the Inertial Navigation System(INS),Global Navigation Satellite System(GNSS),and Visual Navigation System(VNS)for comprehensive information fusion.Specifically,an improved multi-UAV integrated navigation algorithm called Information Fusion with MultiAgent Deep Deterministic Policy Gradient(IF-MADDPG)was developed.This algorithm enables UAVs to learn collaboratively and optimize their flight trajectories in real time.Through simulations and experiments,test scenarios in GNSS-denied environments were constructed to evaluate the effectiveness of the algorithm.The experimental results demonstrate that the IF-MADDPG algorithm significantly enhances the collaborative navigation capabilities of multiple UAVs in formation maintenance and GNSS-denied environments.Additionally,it has advantages in terms of mission completion time.This study provides a novel approach for efficient collaboration in multi-UAV systems,which significantly improves the robustness and adaptability of navigation systems.展开更多
This paper proposes a Multi-Agent Attention Proximal Policy Optimization(MA2PPO)algorithm aiming at the problems such as credit assignment,low collaboration efficiency and weak strategy generalization ability existing...This paper proposes a Multi-Agent Attention Proximal Policy Optimization(MA2PPO)algorithm aiming at the problems such as credit assignment,low collaboration efficiency and weak strategy generalization ability existing in the cooperative pursuit tasks of multiple unmanned aerial vehicles(UAVs).Traditional algorithms often fail to effectively identify critical cooperative relationships in such tasks,leading to low capture efficiency and a significant decline in performance when the scale expands.To tackle these issues,based on the proximal policy optimization(PPO)algorithm,MA2PPO adopts the centralized training with decentralized execution(CTDE)framework and introduces a dynamic decoupling mechanism,that is,sharing the multi-head attention(MHA)mechanism for critics during centralized training to solve the credit assignment problem.This method enables the pursuers to identify highly correlated interactions with their teammates,effectively eliminate irrelevant and weakly relevant interactions,and decompose large-scale cooperation problems into decoupled sub-problems,thereby enhancing the collaborative efficiency and policy stability among multiple agents.Furthermore,a reward function has been devised to facilitate the pursuers to encircle the escapee by combining a formation reward with a distance reward,which incentivizes UAVs to develop sophisticated cooperative pursuit strategies.Experimental results demonstrate the effectiveness of the proposed algorithm in achieving multi-UAV cooperative pursuit and inducing diverse cooperative pursuit behaviors among UAVs.Moreover,experiments on scalability have demonstrated that the algorithm is suitable for large-scale multi-UAV systems.展开更多
This paper studies the countermeasure design problems of distributed resilient time-varying formation-tracking control for multi-UAV systems with single-way communications against composite attacks,including denial-of...This paper studies the countermeasure design problems of distributed resilient time-varying formation-tracking control for multi-UAV systems with single-way communications against composite attacks,including denial-of-services(DoS)attacks,false-data injection attacks,camouflage attacks,and actuation attacks(AAs).Inspired by the concept of digital twin,a new two-layered protocol equipped with a safe and private twin layer(TL)is proposed,which decouples the above problems into the defense scheme against DoS attacks on the TL and the defense scheme against AAs on the cyber-physical layer.First,a topologyrepairing strategy against frequency-constrained DoS attacks is implemented via a Zeno-free event-triggered estimation scheme,which saves communication resources considerably.The upper bound of the reaction time needed to launch the repaired topology after the occurrence of DoS attacks is calculated.Second,a decentralized adaptive and chattering-relief controller against potentially unbounded AAs is designed.Moreover,this novel adaptive controller can achieve uniformly ultimately bounded convergence,whose error bound can be given explicitly.The practicability and validity of this new two-layered protocol are shown via a simulation example and a UAV swarm experiment equipped with both Ultra-WideBand and WiFi communication channels.展开更多
There are many interesting flocking phenomena in nature,such as joint predation and group migration,and the intrinsic communication patterns of flocking are essential for studying group behavior.Traditional models of ...There are many interesting flocking phenomena in nature,such as joint predation and group migration,and the intrinsic communication patterns of flocking are essential for studying group behavior.Traditional models of communication such as the pigeon flock model and the wolf pack model define all agents within a perceptual distance as the neighborhoods,and some models have fixed communicating numbers.There is a significant impact on the quality of the flocking formation when encountering poor initial state of the flocking,multiple obstacles,or loss of certain agents.To solve this problem,this paper proposes a local communication model with nearest agents in four directions.Based on this model and behavioral method,two distributed flocking formation algorithms are designed in this paper for different scenarios,namely the flocking algorithm and the circular formation algorithm.Numerical simulation results show that the flocking can pass through the obstacle area and re-formation smoothly,and also the formation quality of the flocking is better compared with the traditional communication model.展开更多
With the rapid development of Unmanned Aerial Vehicle(UAV)technology,one of the emerging fields is to utilize multi-UAV as a team under autonomous control in a complex environment.Among the challenges in fully achievi...With the rapid development of Unmanned Aerial Vehicle(UAV)technology,one of the emerging fields is to utilize multi-UAV as a team under autonomous control in a complex environment.Among the challenges in fully achieving autonomous control,Cooperative task assignment stands out as the key function.In this paper,we analyze the importance and difficulties of multiUAV cooperative task assignment in characterizing scenarios and obtaining high-quality solutions.Furthermore,we present three promising directions for future research:Cooperative task assignment in a dynamic complex environment,in an unmanned-manned aircraft system and in a UAV swarm.Our goal is to provide a brief review of multi-UAV cooperative task assignment for readers to further explore.展开更多
This paper investigates subcarrier and power allocation in a multi-UAV OFDM system.The study considers a practical scenario,where certain subcarriers are unavailable for dynamic subcarrier allocation,on account of pre...This paper investigates subcarrier and power allocation in a multi-UAV OFDM system.The study considers a practical scenario,where certain subcarriers are unavailable for dynamic subcarrier allocation,on account of pre-allocation for burst transmissions.We first propose a novel iterative algorithm to jointly optimize subcarrier and power allocation,so as to maximize the sum rate of the uplink transmission in the multiUAV OFDM system.The key idea behind our solution is converting the nontrivial allocation problem into a weighted mean square error(MSE) problem.By this means,the allocation problem can be solved by the alternating optimization method.Besides,aiming at a lower-complexity solution,we propose a heuristic allocation scheme,where subcarrier allocation and transmit power allocation are separately optimized.In the heuristic scheme,closedform solution can be obtained for power allocation.Simulation results demonstrate that in the presence of stretched subcarrier resource,the proposed iterative joint optimization algorithm can significantly outperform the heuristic scheme,offering a higher sum rate.展开更多
With the increasing maturity of multi-UAV technology and its broad applications in scenarios such as UAV roundup tasks,this paper proposes a novel approach to enhance interception efficiency and system robustness by a...With the increasing maturity of multi-UAV technology and its broad applications in scenarios such as UAV roundup tasks,this paper proposes a novel approach to enhance interception efficiency and system robustness by addressing insufficient historical data utilization and inadequate environmental explo-ration.The multi-UAV roundup problem is formulated as a Markov Decision Process(MDP),and an Improved Cross-Entropy Method with Intrinsic Curiosity-enhanced Multi-Agent Twin Delayed Deep Deterministic Policy Gradient(I2C-MATD3)is designed.Specifically,an Improved Cross-Entropy Method(ICEM)based on global elite samples rapidly optimizes training strategies while generating extensive experience for a Multi-Agent Twin Delayed Deep Deterministic Policy Gradient algorithm augmented with intrinsic curiosity rewards(IC-MATD3).In turn,IC-MATD3 guides the optimization direction of ICEM,enabling a synergistic interaction that facilitates effective historical data exploitation and pro-active environmental exploration for UAV agents to accomplish roundup tasks.Experiments in complex scenarios demonstrate that the proposed algorithm achieves superior training efficiency and conver-gence performance compared to state-of-the-art multi-agent reinforcement learning(MARL)methods.Robustness tests and ablation experiments further validate its enhanced generalizability and robustness.展开更多
To address the challenge of achieving decentralized,scalable,and adaptive control for large-scale multiple unmanned aerial vehicle(multi-UAV)swarms in dynamic urban environments with obstacles and wind perturbations,w...To address the challenge of achieving decentralized,scalable,and adaptive control for large-scale multiple unmanned aerial vehicle(multi-UAV)swarms in dynamic urban environments with obstacles and wind perturbations,we proposed a hybrid framework integrating adaptive reinforcement learning(RL),multi-modal perception fusion,and enhanced pigeon flock optimization(PFO)with curiosity-driven exploration to enable robust autonomous and formation control.The framework leverages meta-learning to optimize RL policies for real-time adaptation,fuses sensor data for precise state estimation,and enhances PFO with learned leader-follower dynamics and exploration rewards to maintain cohesive formations and explore uncertain areas.For swarms of 10–30 UAVs,it achieves 34%faster convergence,61%reduced stability root mean square error(RMSE),88%fewer collisions and 85.6%–92.3%success rates in target detection and encirclement,outperforming standard multi-agent RL,pure PFO,and single-modality RL.Three-dimensional trajectory visualizations confirm cohesive formations,collision-free maneuvers,and efficient exploration in urban search-and-rescue scenarios.Innovations include meta-RL for rapid adaptation,multi-modal fusion for robust perception,and curiosity-driven PFO for scalable,decentralized control,advancing real-world multi-UAV swarm autonomy and coordination.展开更多
Cooperative multi-UAV search requires jointly optimizing wide-area coverage,rapid target discovery,and endurance under sensing and motion constraints.Resolving this coupling enables scalable coordination with high dat...Cooperative multi-UAV search requires jointly optimizing wide-area coverage,rapid target discovery,and endurance under sensing and motion constraints.Resolving this coupling enables scalable coordination with high data efficiency and mission reliability.We formulate this problem as a discounted Markov decision process on an occupancy grid with a cellwise Bayesian belief update,yielding a Markov state that couples agent poses with a probabilistic target field.On this belief–MDP we introduce a segment-conditioned latent-intent framework,in which a discrete intent head selects a latent skill every K steps and an intra-segment GRU policy generates per-step control conditioned on the fixed intent;both components are trained end-to-end with proximal updates under a centralized critic.On the 50×50 grid,coverage and discovery convergence times are reduced by up to 48%and 40%relative to a flat actor-critic benchmark,and the aggregated convergence metric improves by about 12%compared with a stateof-the-art hierarchical method.Qualitative analyses further reveal stable spatial sectorization,low path overlap,and fuel-aware patrolling,indicating that segment-conditioned latent intents provide an effective and scalable mechanism for coordinated multi-UAV search.展开更多
To address real-time path planning requirements for multi-unmanned aerial vehicle(multi-UAV)collaboration in environments,this study proposes an improved multi-agent deep deterministic policy gradient algorithm with p...To address real-time path planning requirements for multi-unmanned aerial vehicle(multi-UAV)collaboration in environments,this study proposes an improved multi-agent deep deterministic policy gradient algorithm with prioritized experience replay(PER-MADDPG).By designing a multi-dimensional state representation incorporating relative positions,velocity vectors,and obstacle distance fields,we construct a composite reward function integrating safe obstacle avoidance,formation maintenance,and energy efficiency for environment perception and multiobjective collaborative optimization.The prioritized experience replay mechanism dynamically adjusts sampling weights based on temporal difference(TD)errors,enhancing learning efficiency for high-value samples.Simulation experiments demonstrate that our method generates real-time collaborative paths in 3D complex obstacle environments,reducing training time by 25.3%and 16.8%compared to traditional MADDPG and multi-agent twin delayed deep deterministic policy gradient(MATD3)algorithms respectively,while achieving smaller path length variances among UAVs.Results validate the effectiveness of prioritized experience replay in multi-agent collaborative decision-making.展开更多
Unmanned aerial vehicles(UAVs)are widely used in situations with uncertain and risky areas lacking network coverage.In natural disasters,timely delivery of first aid supplies is crucial.Current UAVs face risks such as...Unmanned aerial vehicles(UAVs)are widely used in situations with uncertain and risky areas lacking network coverage.In natural disasters,timely delivery of first aid supplies is crucial.Current UAVs face risks such as crashing into birds or unexpected structures.Airdrop systems with parachutes risk dispersing payloads away from target locations.The objective here is to use multiple UAVs to distribute payloads cooperatively to assigned locations.The civil defense department must balance coverage,accurate landing,and flight safety while considering battery power and capability.Deep Q-network(DQN)models are commonly used in multi-UAV path planning to effectively represent the surroundings and action spaces.Earlier strategies focused on advanced DQNs for UAV path planning in different configurations,but rarely addressed non-cooperative scenarios and disaster environments.This paper introduces a new DQN framework to tackle challenges in disaster environments.It considers unforeseen structures and birds that could cause UAV crashes and assumes urgent landing zones and winch-based airdrop systems for precise delivery and return.A new DQN model is developed,which incorporates the battery life,safe flying distance between UAVs,and remaining delivery points to encode surrounding hazards into the state space and Q-networks.Additionally,a unique reward system is created to improve UAV action sequences for better delivery coverage and safe landings.The experimental results demonstrate that multi-UAV first aid delivery in disaster environments can achieve advanced performance.展开更多
Multiple UAVs cooperative target search has been widely used in various environments,such as emergency rescue and traffic monitoring.However,uncertain communication network among UAVs exhibits unstable links and rapid...Multiple UAVs cooperative target search has been widely used in various environments,such as emergency rescue and traffic monitoring.However,uncertain communication network among UAVs exhibits unstable links and rapid topological fluctuations due to mission complexity and unpredictable environmental states.This limitation hinders timely information sharing and insightful path decisions for UAVs,resulting in inefficient or even failed collaborative search.Aiming at this issue,this paper proposes a multi-UAV cooperative search strategy by developing a real-time trajectory decision that incorporates autonomous connectivity to reinforce multi-UAV collaboration and achieve search acceleration in uncertain search environments.Specifically,an autonomous connectivity strategy based on node cognitive information and network states is introduced to enable effective message transmission and adapt to the dynamic network environment.Based on the fused information,we formalize the trajectory planning as a multiobjective optimization problem by jointly considering search performance and UAV energy harnessing.A multi-agent deep reinforcement learning based algorithm is proposed to solve it,where the reward-guided real-time path is determined to achieve an energyefficient search.Finally,extensive experimental results show that the proposed algorithm outperforms existing works in terms of average search rate and coverage rate with reduced energy consumption under uncertain search environments.展开更多
This study introduces a novel algorithm known as the dung beetle optimization algorithm based on bounded reflection optimization andmulti-strategy fusion(BFDBO),which is designed to tackle the complexities associated ...This study introduces a novel algorithm known as the dung beetle optimization algorithm based on bounded reflection optimization andmulti-strategy fusion(BFDBO),which is designed to tackle the complexities associated with multi-UAV collaborative trajectory planning in intricate battlefield environments.Initially,a collaborative planning cost function for the multi-UAV system is formulated,thereby converting the trajectory planning challenge into an optimization problem.Building on the foundational dung beetle optimization(DBO)algorithm,BFDBO incorporates three significant innovations:a boundary reflection mechanism,an adaptive mixed exploration strategy,and a dynamic multi-scale mutation strategy.These enhancements are intended to optimize the equilibrium between local exploration and global exploitation,facilitating the discovery of globally optimal trajectories thatminimize the cost function.Numerical simulations utilizing the CEC2022 benchmark function indicate that all three enhancements of BFDBOpositively influence its performance,resulting in accelerated convergence and improved optimization accuracy relative to leading optimization algorithms.In two battlefield scenarios of varying complexities,BFDBO achieved a minimum of a 39% reduction in total trajectory planning costs when compared to DBO and three other highperformance variants,while also demonstrating superior average runtime.This evidence underscores the effectiveness and applicability of BFDBO in practical,real-world contexts.展开更多
Aiming at the problem of low convergence efficiency of traditional multi-UAV path planning algorithms in unknown complex environments,this paper proposes a deep reinforcement learning algorithm incorporating the atten...Aiming at the problem of low convergence efficiency of traditional multi-UAV path planning algorithms in unknown complex environments,this paper proposes a deep reinforcement learning algorithm incorporating the attention mechanism.The method is based on the Soft Actor-Critic(SAC)framework,which introduces a multi-attention mechanism in the Critic network,dynamically learns the dependency relationship between intelligences,and realizes key information screening and conflict avoidance.An environment with multiple random obstacles is designed to simulate complex emergent situations.The results show that the proposed algorithm significantly improves the mission success rate and average reward,significantly extends the survival time and exploration range of the UAVs,and verifies the effectiveness of the attention mechanism in enhancing the efficiency,robustness,and long-term planning capability of multi-UAV collaboration,as compared to the baseline method that does not use attention.展开更多
In emergency communication scenarios,exploiting Unmanned Aerial Vehicles(UAVs)as relays to provide wireless communication services for ground users has emerged as a promising application.A key challenge in this resour...In emergency communication scenarios,exploiting Unmanned Aerial Vehicles(UAVs)as relays to provide wireless communication services for ground users has emerged as a promising application.A key challenge in this resource-constrained application is deploying the minimum number of UAVs to form an aerial backhaul network to ensure coverage,which composes the Number and Placement Optimization for the Backhaul-Aware Network Deployment(NPO-BAND)problem.In this paper,we first formulate the NPO-BAND problem based on the geometric disk coverage model.Then,we propose a low-complexity heuristic method to solve this NP-hard problem.The proposed method contains a Very Important Point-Choosing(VIPC)strategy and a Backhaul-Aware Local Coverage(BALC)algorithm.Specifically,the VIPC strategy weighs up the backhaul connectivity constraint and the ground user coverage to choose the VIP,while the BALC algorithm solves the extended 1-center problem to determine the deployment location of each UAV.Simulation results show that the proposed method can effectively reduce the number of deployed UAVs,saving up to 25%-50%of that compared to existing methods across varying numbers and area sizes in clustered distribution patterns of ground users.展开更多
In this paper,we present a distributed framework for the lidar-based relative state estimator which achieves highly accurate,real-time trajectory estimation of multiple Unmanned Aerial Vehicles(UAVs)in GPS-denied envi...In this paper,we present a distributed framework for the lidar-based relative state estimator which achieves highly accurate,real-time trajectory estimation of multiple Unmanned Aerial Vehicles(UAVs)in GPS-denied environments.The system builds atop a factor graph,and only on-board sensors and computing power are utilized.Benefiting from the keyframe strategy,each UAV performs relative state estimation individually and broadcasts very partial information without exchanging raw data.The complete system runs in real-time and is evaluated with three experiments in different environments.Experimental results show that the proposed distributed approach offers comparable performance with a centralized method in terms of accuracy and real-time performance.The flight test demonstrates that the proposed relative state estimation framework is able to be used for aggressive flights over 5 m/s.展开更多
Unmanned Aerial Vehicle(UAV)-assisted Vehicular Edge Computing Networks(VECNs)have emerged as a promising solution to enhance service quality for ground vehicle users.However,the growing demands from users and the lim...Unmanned Aerial Vehicle(UAV)-assisted Vehicular Edge Computing Networks(VECNs)have emerged as a promising solution to enhance service quality for ground vehicle users.However,the growing demands from users and the limited computing and storage resources of UAVs present significant challenges in designing an efficient edge service caching scheme to minimize latency.Moreover,the integration of service caching and task offloading complicates the support of complex tasks by a single UAV.To address these challenges,this paper proposes a novel two-tier UAV-assisted VECNs framework.In this framework,multi-rotor UAVs function as hovering nodes for computational offloading,while a fixed-wing UAV serves as a mobile auxiliary cloud platform,forming a cohesive UAV group.User tasks are structured into a task chain based on the available UAVs.We integrate a joint service chain caching and task offloading scheme that considers UAV computing and storage capacities,duplicate caching,and dynamic transmission latency.To optimize task chain completion latency,we propose an Attention-based Multi-Agent Deep Q-Network(A-MADQN)algorithm.This algorithm incorporates an attention mechanism to narrow the UAV selection space,enabling the selected UAVs to collaboratively make caching and task offloading decisions.Numerical results demonstrate that the proposed algorithm significantly enhances system processing efficiency and reduces task completion latency compared to the benchmark approaches.展开更多
Unmanned aerial vehicles(UAVs)are becoming a common solution to urban mobility,and traffic monitoring as well,owing to their ability to be deployed flexibly,ability to see a broader area and real-time sensing.However,...Unmanned aerial vehicles(UAVs)are becoming a common solution to urban mobility,and traffic monitoring as well,owing to their ability to be deployed flexibly,ability to see a broader area and real-time sensing.However,the reliability of UAV-assisted traffic systems can be compromised through identity spoofing,Sybil attacks,false data injection,and trajectory manipulation.Current authentication techniques primarily verify cryptographic identities but often cannot detect when a claimed identity is inconsistent with physical movement patterns and settings.To overcome this drawback,this paper presents a context-aware identity validation system,CIV-UAV,for UAV-based urban traffic surveillance.The paradigm combines a model of cryptographic validation,model mobility,on-the-fly visual,road-network,temporal continuity,anomaly scoring,and multi-UAV consensus into a cohesive trust-based validation model.The risk-adaptive policy also adjusts the validation strictness based on the seriousness of the situation and the level of uncertainty.The outcomes of simulations indicate that CIV-UAV enhances identity validation,lowers the false detection and false acceptance rates,and reinforces the detection of spoofing,Sybil behaviour,path forgery,injection of fake events,and vision-communication mismatch attacks.The suggested architecture provides an identity validation system that is easy to implement and can be upgraded to a next-generation UAV-intelligent transportation network.展开更多
基金co-supported by the Natural Science Foundation of Hunan Province,China(No.2025JJ20055)the Science and Technology Innovation Program of Hunan Province,China(No.2022RC1095)the Joint Funds of the National Natural Science Foundation of China(No.U23B2032)。
摘要Execution uncertainties,such as motion delay or confrontation in pursuit-evasion problems with rapidly changing states,affect the task performance of multi-Unmanned Aerial Vehicle(UAV)systems.This may lead to the failure of the initial task assignment scheme.To address this problem,this paper takes the interception scenario as a typical case.It proposes a distributed dynamic task assignment algorithm based on an evolving task performance model to reassign UAVs to tasks in an event-triggered manner.This paper combines the underlying execution model with the interception effectiveness model to design the evolving task performance model.This model describes the UAV task performance in a finite interval by predicting and integrating the states and actions of the intercepted UAVs and the targets under execution uncertainty.The discrete task monitor triggers reassignment based on the severity of the task performance deviation.The Consensus-based Auction Algorithm(CBAA)is extended to optimize the task performance function and efficiently give the reassignment scheme.Simulation results demonstrate the feasibility and effectiveness of the proposed algorithm.
基金co-supported by the National Natural Science Foundation of China(Nos.92371201 and 52192633)the Natural Science Foundation of Shaanxi Province of China(No.2022JC-03)the Aeronautical Science Foundation of China(No.ASFC-20220019070002)。
摘要In multiple Unmanned Aerial Vehicles(UAV)systems,achieving efficient navigation is essential for executing complex tasks and enhancing autonomy.Traditional navigation methods depend on predefined control strategies and trajectory planning and often perform poorly in complex environments.To improve the UAV-environment interaction efficiency,this study proposes a multi-UAV integrated navigation algorithm based on Deep Reinforcement Learning(DRL).This algorithm integrates the Inertial Navigation System(INS),Global Navigation Satellite System(GNSS),and Visual Navigation System(VNS)for comprehensive information fusion.Specifically,an improved multi-UAV integrated navigation algorithm called Information Fusion with MultiAgent Deep Deterministic Policy Gradient(IF-MADDPG)was developed.This algorithm enables UAVs to learn collaboratively and optimize their flight trajectories in real time.Through simulations and experiments,test scenarios in GNSS-denied environments were constructed to evaluate the effectiveness of the algorithm.The experimental results demonstrate that the IF-MADDPG algorithm significantly enhances the collaborative navigation capabilities of multiple UAVs in formation maintenance and GNSS-denied environments.Additionally,it has advantages in terms of mission completion time.This study provides a novel approach for efficient collaboration in multi-UAV systems,which significantly improves the robustness and adaptability of navigation systems.
基金supported by the National Research and Development Program of China under Grant JCKY2018607C019in part by the Key Laboratory Fund of UAV of Northwestern Polytechnical University under Grant 2021JCJQLB0710L.
摘要This paper proposes a Multi-Agent Attention Proximal Policy Optimization(MA2PPO)algorithm aiming at the problems such as credit assignment,low collaboration efficiency and weak strategy generalization ability existing in the cooperative pursuit tasks of multiple unmanned aerial vehicles(UAVs).Traditional algorithms often fail to effectively identify critical cooperative relationships in such tasks,leading to low capture efficiency and a significant decline in performance when the scale expands.To tackle these issues,based on the proximal policy optimization(PPO)algorithm,MA2PPO adopts the centralized training with decentralized execution(CTDE)framework and introduces a dynamic decoupling mechanism,that is,sharing the multi-head attention(MHA)mechanism for critics during centralized training to solve the credit assignment problem.This method enables the pursuers to identify highly correlated interactions with their teammates,effectively eliminate irrelevant and weakly relevant interactions,and decompose large-scale cooperation problems into decoupled sub-problems,thereby enhancing the collaborative efficiency and policy stability among multiple agents.Furthermore,a reward function has been devised to facilitate the pursuers to encircle the escapee by combining a formation reward with a distance reward,which incentivizes UAVs to develop sophisticated cooperative pursuit strategies.Experimental results demonstrate the effectiveness of the proposed algorithm in achieving multi-UAV cooperative pursuit and inducing diverse cooperative pursuit behaviors among UAVs.Moreover,experiments on scalability have demonstrated that the algorithm is suitable for large-scale multi-UAV systems.
基金This work was supported in part by the National Natural Science Foundation of China(61903258)Guangdong Basic and Applied Basic Research Foundation(2022A1515010234)+1 种基金the Project of Department of Education of Guangdong Province(2022KTSCX105)Qatar National Research Fund(NPRP12C-0814-190012).
摘要This paper studies the countermeasure design problems of distributed resilient time-varying formation-tracking control for multi-UAV systems with single-way communications against composite attacks,including denial-of-services(DoS)attacks,false-data injection attacks,camouflage attacks,and actuation attacks(AAs).Inspired by the concept of digital twin,a new two-layered protocol equipped with a safe and private twin layer(TL)is proposed,which decouples the above problems into the defense scheme against DoS attacks on the TL and the defense scheme against AAs on the cyber-physical layer.First,a topologyrepairing strategy against frequency-constrained DoS attacks is implemented via a Zeno-free event-triggered estimation scheme,which saves communication resources considerably.The upper bound of the reaction time needed to launch the repaired topology after the occurrence of DoS attacks is calculated.Second,a decentralized adaptive and chattering-relief controller against potentially unbounded AAs is designed.Moreover,this novel adaptive controller can achieve uniformly ultimately bounded convergence,whose error bound can be given explicitly.The practicability and validity of this new two-layered protocol are shown via a simulation example and a UAV swarm experiment equipped with both Ultra-WideBand and WiFi communication channels.
基金Jilin Province Development and Reform Commission under Grant[2020C018-2]Jilin Province Key R&D Plan Project under Grant[20200401113GX].
摘要There are many interesting flocking phenomena in nature,such as joint predation and group migration,and the intrinsic communication patterns of flocking are essential for studying group behavior.Traditional models of communication such as the pigeon flock model and the wolf pack model define all agents within a perceptual distance as the neighborhoods,and some models have fixed communicating numbers.There is a significant impact on the quality of the flocking formation when encountering poor initial state of the flocking,multiple obstacles,or loss of certain agents.To solve this problem,this paper proposes a local communication model with nearest agents in four directions.Based on this model and behavioral method,two distributed flocking formation algorithms are designed in this paper for different scenarios,namely the flocking algorithm and the circular formation algorithm.Numerical simulation results show that the flocking can pass through the obstacle area and re-formation smoothly,and also the formation quality of the flocking is better compared with the traditional communication model.
基金supported in part by the National Natural Science Foundation of China(Nos.61671031,61722102,91738301)。
摘要With the rapid development of Unmanned Aerial Vehicle(UAV)technology,one of the emerging fields is to utilize multi-UAV as a team under autonomous control in a complex environment.Among the challenges in fully achieving autonomous control,Cooperative task assignment stands out as the key function.In this paper,we analyze the importance and difficulties of multiUAV cooperative task assignment in characterizing scenarios and obtaining high-quality solutions.Furthermore,we present three promising directions for future research:Cooperative task assignment in a dynamic complex environment,in an unmanned-manned aircraft system and in a UAV swarm.Our goal is to provide a brief review of multi-UAV cooperative task assignment for readers to further explore.
基金supported by China NSF Grants(61631020)the Fundamental Research Funds for the Central Universities(NP2018103,NE2017103,NC2017003)
摘要This paper investigates subcarrier and power allocation in a multi-UAV OFDM system.The study considers a practical scenario,where certain subcarriers are unavailable for dynamic subcarrier allocation,on account of pre-allocation for burst transmissions.We first propose a novel iterative algorithm to jointly optimize subcarrier and power allocation,so as to maximize the sum rate of the uplink transmission in the multiUAV OFDM system.The key idea behind our solution is converting the nontrivial allocation problem into a weighted mean square error(MSE) problem.By this means,the allocation problem can be solved by the alternating optimization method.Besides,aiming at a lower-complexity solution,we propose a heuristic allocation scheme,where subcarrier allocation and transmit power allocation are separately optimized.In the heuristic scheme,closedform solution can be obtained for power allocation.Simulation results demonstrate that in the presence of stretched subcarrier resource,the proposed iterative joint optimization algorithm can significantly outperform the heuristic scheme,offering a higher sum rate.
基金the National Key Lab-oratory of Air-based Information Perception and Fusion(Grant No.202510)the Key Research and Development Program of Shaanxi Province(Grant No.2023-GHZD-33)+2 种基金the Fundamental Research Funds for the Central Universities(Grant No.H20250607)the Open Project of the State Key Laboratory of Intelligent Game(Grant No.ZBKF-23-05)the National Nature Science Foundation of China(Grant No.62003267)to provide fund for conducting experiments.
摘要With the increasing maturity of multi-UAV technology and its broad applications in scenarios such as UAV roundup tasks,this paper proposes a novel approach to enhance interception efficiency and system robustness by addressing insufficient historical data utilization and inadequate environmental explo-ration.The multi-UAV roundup problem is formulated as a Markov Decision Process(MDP),and an Improved Cross-Entropy Method with Intrinsic Curiosity-enhanced Multi-Agent Twin Delayed Deep Deterministic Policy Gradient(I2C-MATD3)is designed.Specifically,an Improved Cross-Entropy Method(ICEM)based on global elite samples rapidly optimizes training strategies while generating extensive experience for a Multi-Agent Twin Delayed Deep Deterministic Policy Gradient algorithm augmented with intrinsic curiosity rewards(IC-MATD3).In turn,IC-MATD3 guides the optimization direction of ICEM,enabling a synergistic interaction that facilitates effective historical data exploitation and pro-active environmental exploration for UAV agents to accomplish roundup tasks.Experiments in complex scenarios demonstrate that the proposed algorithm achieves superior training efficiency and conver-gence performance compared to state-of-the-art multi-agent reinforcement learning(MARL)methods.Robustness tests and ablation experiments further validate its enhanced generalizability and robustness.
基金supported by the National Natural Science Foundation of China(No.62350048)。
摘要To address the challenge of achieving decentralized,scalable,and adaptive control for large-scale multiple unmanned aerial vehicle(multi-UAV)swarms in dynamic urban environments with obstacles and wind perturbations,we proposed a hybrid framework integrating adaptive reinforcement learning(RL),multi-modal perception fusion,and enhanced pigeon flock optimization(PFO)with curiosity-driven exploration to enable robust autonomous and formation control.The framework leverages meta-learning to optimize RL policies for real-time adaptation,fuses sensor data for precise state estimation,and enhances PFO with learned leader-follower dynamics and exploration rewards to maintain cohesive formations and explore uncertain areas.For swarms of 10–30 UAVs,it achieves 34%faster convergence,61%reduced stability root mean square error(RMSE),88%fewer collisions and 85.6%–92.3%success rates in target detection and encirclement,outperforming standard multi-agent RL,pure PFO,and single-modality RL.Three-dimensional trajectory visualizations confirm cohesive formations,collision-free maneuvers,and efficient exploration in urban search-and-rescue scenarios.Innovations include meta-RL for rapid adaptation,multi-modal fusion for robust perception,and curiosity-driven PFO for scalable,decentralized control,advancing real-world multi-UAV swarm autonomy and coordination.
摘要Cooperative multi-UAV search requires jointly optimizing wide-area coverage,rapid target discovery,and endurance under sensing and motion constraints.Resolving this coupling enables scalable coordination with high data efficiency and mission reliability.We formulate this problem as a discounted Markov decision process on an occupancy grid with a cellwise Bayesian belief update,yielding a Markov state that couples agent poses with a probabilistic target field.On this belief–MDP we introduce a segment-conditioned latent-intent framework,in which a discrete intent head selects a latent skill every K steps and an intra-segment GRU policy generates per-step control conditioned on the fixed intent;both components are trained end-to-end with proximal updates under a centralized critic.On the 50×50 grid,coverage and discovery convergence times are reduced by up to 48%and 40%relative to a flat actor-critic benchmark,and the aggregated convergence metric improves by about 12%compared with a stateof-the-art hierarchical method.Qualitative analyses further reveal stable spatial sectorization,low path overlap,and fuel-aware patrolling,indicating that segment-conditioned latent intents provide an effective and scalable mechanism for coordinated multi-UAV search.
基金supported by the open project of National Key Laboratory of Air-Based Information Perception and Fusion(No.202462)。
摘要To address real-time path planning requirements for multi-unmanned aerial vehicle(multi-UAV)collaboration in environments,this study proposes an improved multi-agent deep deterministic policy gradient algorithm with prioritized experience replay(PER-MADDPG).By designing a multi-dimensional state representation incorporating relative positions,velocity vectors,and obstacle distance fields,we construct a composite reward function integrating safe obstacle avoidance,formation maintenance,and energy efficiency for environment perception and multiobjective collaborative optimization.The prioritized experience replay mechanism dynamically adjusts sampling weights based on temporal difference(TD)errors,enhancing learning efficiency for high-value samples.Simulation experiments demonstrate that our method generates real-time collaborative paths in 3D complex obstacle environments,reducing training time by 25.3%and 16.8%compared to traditional MADDPG and multi-agent twin delayed deep deterministic policy gradient(MATD3)algorithms respectively,while achieving smaller path length variances among UAVs.Results validate the effectiveness of prioritized experience replay in multi-agent collaborative decision-making.
基金supported by the Committee of Science of the Ministry of Education and Science of the Republic of Kazakhstan under Grant No.249015/0224.
摘要Unmanned aerial vehicles(UAVs)are widely used in situations with uncertain and risky areas lacking network coverage.In natural disasters,timely delivery of first aid supplies is crucial.Current UAVs face risks such as crashing into birds or unexpected structures.Airdrop systems with parachutes risk dispersing payloads away from target locations.The objective here is to use multiple UAVs to distribute payloads cooperatively to assigned locations.The civil defense department must balance coverage,accurate landing,and flight safety while considering battery power and capability.Deep Q-network(DQN)models are commonly used in multi-UAV path planning to effectively represent the surroundings and action spaces.Earlier strategies focused on advanced DQNs for UAV path planning in different configurations,but rarely addressed non-cooperative scenarios and disaster environments.This paper introduces a new DQN framework to tackle challenges in disaster environments.It considers unforeseen structures and birds that could cause UAV crashes and assumes urgent landing zones and winch-based airdrop systems for precise delivery and return.A new DQN model is developed,which incorporates the battery life,safe flying distance between UAVs,and remaining delivery points to encode surrounding hazards into the state space and Q-networks.Additionally,a unique reward system is created to improve UAV action sequences for better delivery coverage and safe landings.The experimental results demonstrate that multi-UAV first aid delivery in disaster environments can achieve advanced performance.
基金supported by National Natural Science Foundation of China(No.62202449 and No.62472410)National Key Research and Development Program of China(2021YFB2900102)。
摘要Multiple UAVs cooperative target search has been widely used in various environments,such as emergency rescue and traffic monitoring.However,uncertain communication network among UAVs exhibits unstable links and rapid topological fluctuations due to mission complexity and unpredictable environmental states.This limitation hinders timely information sharing and insightful path decisions for UAVs,resulting in inefficient or even failed collaborative search.Aiming at this issue,this paper proposes a multi-UAV cooperative search strategy by developing a real-time trajectory decision that incorporates autonomous connectivity to reinforce multi-UAV collaboration and achieve search acceleration in uncertain search environments.Specifically,an autonomous connectivity strategy based on node cognitive information and network states is introduced to enable effective message transmission and adapt to the dynamic network environment.Based on the fused information,we formalize the trajectory planning as a multiobjective optimization problem by jointly considering search performance and UAV energy harnessing.A multi-agent deep reinforcement learning based algorithm is proposed to solve it,where the reward-guided real-time path is determined to achieve an energyefficient search.Finally,extensive experimental results show that the proposed algorithm outperforms existing works in terms of average search rate and coverage rate with reduced energy consumption under uncertain search environments.
基金funded by the National Defense Science and Technology Innovation project,grant number ZZKY20223103the Basic Frontier InnovationProject at the Engineering University of PAP,grant number WJY202429+2 种基金the Basic Frontier lnnovation Project at the Engineering University of PAP,grant number WJY202408the Graduate Student Funding Priority Project,grant number JYWJ2024B006Key project of National Social Science Foundation,grant number 2023-SKJJ-A-116.
摘要This study introduces a novel algorithm known as the dung beetle optimization algorithm based on bounded reflection optimization andmulti-strategy fusion(BFDBO),which is designed to tackle the complexities associated with multi-UAV collaborative trajectory planning in intricate battlefield environments.Initially,a collaborative planning cost function for the multi-UAV system is formulated,thereby converting the trajectory planning challenge into an optimization problem.Building on the foundational dung beetle optimization(DBO)algorithm,BFDBO incorporates three significant innovations:a boundary reflection mechanism,an adaptive mixed exploration strategy,and a dynamic multi-scale mutation strategy.These enhancements are intended to optimize the equilibrium between local exploration and global exploitation,facilitating the discovery of globally optimal trajectories thatminimize the cost function.Numerical simulations utilizing the CEC2022 benchmark function indicate that all three enhancements of BFDBOpositively influence its performance,resulting in accelerated convergence and improved optimization accuracy relative to leading optimization algorithms.In two battlefield scenarios of varying complexities,BFDBO achieved a minimum of a 39% reduction in total trajectory planning costs when compared to DBO and three other highperformance variants,while also demonstrating superior average runtime.This evidence underscores the effectiveness and applicability of BFDBO in practical,real-world contexts.
摘要Aiming at the problem of low convergence efficiency of traditional multi-UAV path planning algorithms in unknown complex environments,this paper proposes a deep reinforcement learning algorithm incorporating the attention mechanism.The method is based on the Soft Actor-Critic(SAC)framework,which introduces a multi-attention mechanism in the Critic network,dynamically learns the dependency relationship between intelligences,and realizes key information screening and conflict avoidance.An environment with multiple random obstacles is designed to simulate complex emergent situations.The results show that the proposed algorithm significantly improves the mission success rate and average reward,significantly extends the survival time and exploration range of the UAVs,and verifies the effectiveness of the attention mechanism in enhancing the efficiency,robustness,and long-term planning capability of multi-UAV collaboration,as compared to the baseline method that does not use attention.
基金supported by the Natural Science Foundation of China under Grant No.91948303。
摘要In emergency communication scenarios,exploiting Unmanned Aerial Vehicles(UAVs)as relays to provide wireless communication services for ground users has emerged as a promising application.A key challenge in this resource-constrained application is deploying the minimum number of UAVs to form an aerial backhaul network to ensure coverage,which composes the Number and Placement Optimization for the Backhaul-Aware Network Deployment(NPO-BAND)problem.In this paper,we first formulate the NPO-BAND problem based on the geometric disk coverage model.Then,we propose a low-complexity heuristic method to solve this NP-hard problem.The proposed method contains a Very Important Point-Choosing(VIPC)strategy and a Backhaul-Aware Local Coverage(BALC)algorithm.Specifically,the VIPC strategy weighs up the backhaul connectivity constraint and the ground user coverage to choose the VIP,while the BALC algorithm solves the extended 1-center problem to determine the deployment location of each UAV.Simulation results show that the proposed method can effectively reduce the number of deployed UAVs,saving up to 25%-50%of that compared to existing methods across varying numbers and area sizes in clustered distribution patterns of ground users.
基金supported by the National Key Research and Development Program of China(No.2018AAA0102401)the National Natural Science Foundation of China(Nos.62022060,61773278,61873340).
摘要In this paper,we present a distributed framework for the lidar-based relative state estimator which achieves highly accurate,real-time trajectory estimation of multiple Unmanned Aerial Vehicles(UAVs)in GPS-denied environments.The system builds atop a factor graph,and only on-board sensors and computing power are utilized.Benefiting from the keyframe strategy,each UAV performs relative state estimation individually and broadcasts very partial information without exchanging raw data.The complete system runs in real-time and is evaluated with three experiments in different environments.Experimental results show that the proposed distributed approach offers comparable performance with a centralized method in terms of accuracy and real-time performance.The flight test demonstrates that the proposed relative state estimation framework is able to be used for aggressive flights over 5 m/s.
基金supported by Guangdong Basic and Applied Basic Research Foundation(Grant No.2024A1515012745)。
摘要Unmanned Aerial Vehicle(UAV)-assisted Vehicular Edge Computing Networks(VECNs)have emerged as a promising solution to enhance service quality for ground vehicle users.However,the growing demands from users and the limited computing and storage resources of UAVs present significant challenges in designing an efficient edge service caching scheme to minimize latency.Moreover,the integration of service caching and task offloading complicates the support of complex tasks by a single UAV.To address these challenges,this paper proposes a novel two-tier UAV-assisted VECNs framework.In this framework,multi-rotor UAVs function as hovering nodes for computational offloading,while a fixed-wing UAV serves as a mobile auxiliary cloud platform,forming a cohesive UAV group.User tasks are structured into a task chain based on the available UAVs.We integrate a joint service chain caching and task offloading scheme that considers UAV computing and storage capacities,duplicate caching,and dynamic transmission latency.To optimize task chain completion latency,we propose an Attention-based Multi-Agent Deep Q-Network(A-MADQN)algorithm.This algorithm incorporates an attention mechanism to narrow the UAV selection space,enabling the selected UAVs to collaboratively make caching and task offloading decisions.Numerical results demonstrate that the proposed algorithm significantly enhances system processing efficiency and reduces task completion latency compared to the benchmark approaches.
基金supported by the Korea Institute of Marine Science and Technology Promotion(KIMST),in 2022 through the Project is Development and Demonstration of Data Platform for AI Based Safe Fishing Vessel Design(Grant Number:RS-2022-KS221571).
摘要Unmanned aerial vehicles(UAVs)are becoming a common solution to urban mobility,and traffic monitoring as well,owing to their ability to be deployed flexibly,ability to see a broader area and real-time sensing.However,the reliability of UAV-assisted traffic systems can be compromised through identity spoofing,Sybil attacks,false data injection,and trajectory manipulation.Current authentication techniques primarily verify cryptographic identities but often cannot detect when a claimed identity is inconsistent with physical movement patterns and settings.To overcome this drawback,this paper presents a context-aware identity validation system,CIV-UAV,for UAV-based urban traffic surveillance.The paradigm combines a model of cryptographic validation,model mobility,on-the-fly visual,road-network,temporal continuity,anomaly scoring,and multi-UAV consensus into a cohesive trust-based validation model.The risk-adaptive policy also adjusts the validation strictness based on the seriousness of the situation and the level of uncertainty.The outcomes of simulations indicate that CIV-UAV enhances identity validation,lowers the false detection and false acceptance rates,and reinforces the detection of spoofing,Sybil behaviour,path forgery,injection of fake events,and vision-communication mismatch attacks.The suggested architecture provides an identity validation system that is easy to implement and can be upgraded to a next-generation UAV-intelligent transportation network.