In this paper,a self-developed master-slave follow-up disc cutter is used to conduct rock-breaking tests on hard sandstone samples.Different working parameters were employed in the tests(e.g.cutting depth,cutting spee...In this paper,a self-developed master-slave follow-up disc cutter is used to conduct rock-breaking tests on hard sandstone samples.Different working parameters were employed in the tests(e.g.cutting depth,cutting speed,cutting angle,and rotational speed)in order to explore their influences on cutting performance.The results indicate that the thrust,torque,vibration velocity,and roughness all increased continuously with increase of the propulsion speed and cutting depth.At the same time,the specific energy consumption was found to decrease continuously.As the rotational speed was increased,the thrust increased at first and then decreased.In contrast,the torque and roughness continuously decreased,and the specific energy consumption and vibration speed continuously increased.When the cutting angle was increased,the thrust remained unchanged.However,the torque,specific energy consumption,and vibration speed all decreased continuously,and the roughness increased continuously.The temperature of the surface of the cutting tool was found to be relatively uniformly distributed during the rock-breaking process;the highest temperatures generated were in the range of 200-300℃.As the propulsion speed,cutting depth,and cutting angle were increased,the proportion of tensile fractures produced appeared to increase and the proportion of shear fractures decreased.As the rotational speed was increased,the proportion of tensile fractures decreased and the proportion of shear fractures increased.The results could provide useful information on the rock-breaking behavior involved and can be used to offer technical support for engineers using master-slave follow-up disc cutters in the field.展开更多
The problem of maneuvering for a servicing spacecraft(inspector)to inspect a noncooperative spacecraft(evader)in cislunar space is investigated in this paper.The evader,which may be a malfunctioning or uncontrolled sa...The problem of maneuvering for a servicing spacecraft(inspector)to inspect a noncooperative spacecraft(evader)in cislunar space is investigated in this paper.The evader,which may be a malfunctioning or uncontrolled satellite,introduces uncertainties due to its potential maneuvering capabilities.To address this challenge,the scenario is modeled as a special orbital game,incorporating the unique complexities of the cislunar environment.A variable-duration,turn-based inspection and anti-inspection game model is designed.The model defines both players'rules,constraints,and victory conditions,providing a framework for non-cooperative inspection.Strategies for both players are developed and validated based on their dynamical properties.The inspector's strategy integrates two-body Lambert transfers with shooting methods,while the evader's strategy aims to maximize the inspector's fuel consumption.Simulation results show that the evader's optimal strategy involves deliberate fluctuations in its lunar periapsis altitude,with the inspector's requiredΔV up to eight times greater than the evader's.The impact of game constraints is evaluated,and the effectiveness of deploying the inspector in low lunar orbit is compared with the inspector at the Earth-Moon Lagrange point L1.The strengths and weaknesses of both are shown.These findings provide valuable insights for future orbital servicing and orbital games.展开更多
To uncover the decision-making mechanisms and evolutionary dynamics of multiple stakeholders in highway noise pollution control,a three-party evolutionary game model involving the government,operators,and the public i...To uncover the decision-making mechanisms and evolutionary dynamics of multiple stakeholders in highway noise pollution control,a three-party evolutionary game model involving the government,operators,and the public is constructed.The operation period is divided into different stages for differentiated analysis.A simulation analysis was performed on the Lituo sinking section of the Beijing-Hong Kong-Macao Highway to assess the impact of variations in critical elements on the system.The results indicate that the Lituo sinking section of the Beijing-Hong Kong-Macao Highway is currently in its early stage of development,with the corresponding strategies being active regulation,excessive emissions,and supervision.When the cost of the government’s active regulation decreases from 1×105 to 5×104 yuan,the system converges more rapidly toward the active regulation strategy.When the cost of the operator’s excessive emissions increases from 14.08×106 to 20.00×106 yuan,the system drives the operator toward the standardized emission strategy.In addition,when the cost of public supervision decreases from 15×104 to 5×104 yuan and the compensation paid by operators to the public increases from 1.288×106 to 2.576×106 yuan,the system converges more quickly toward the supervision strategy.The cost of the operator’s excessive emissions serves as the core decision variable for achieving the ideal equilibrium in the three-party game involving government active regulation,operator standardized emissions,and public supervision.展开更多
This paper proposes a two-phase game guidance strategy for the three-body confrontation scenario according to the linear quadratic differential game method,which includes an Attacker,an Interceptor,and a Target.The in...This paper proposes a two-phase game guidance strategy for the three-body confrontation scenario according to the linear quadratic differential game method,which includes an Attacker,an Interceptor,and a Target.The interception probabilities between the Attacker and the Interceptor-Target team are estimated using the probability density function.The desired zero-effort miss distances associated with the interception probabilities for the three-body conflict are acquired by virtue of the gradient descent method.The game combat is divided into two phases by introducing the switching time.In Phase 1,a differential game strategy is developed to guide the zero-control miss distance between the Attacker and the Interceptor-Target into the desired position,which guarantees that the Attacker has the maximizing probability of intercepting the Target and the minimizing probability of being captured by the Interceptor.In Phase 2,a differential game guidance strategy is proposed to ensure that the Attacker evades the pursuit of the Interceptor and intercepts the Target at the preset impact angle.Finally,numerical simulation verifies the effectiveness of the two-stage game guidance strategy.展开更多
To alleviate the conflicts resulting from lane-changing in human-machine co-driving,this study investigates four game scenarios:human-driven lanechanging vehicles(HLVs)and human-driven target vehicles(HTVs),HLVs and A...To alleviate the conflicts resulting from lane-changing in human-machine co-driving,this study investigates four game scenarios:human-driven lanechanging vehicles(HLVs)and human-driven target vehicles(HTVs),HLVs and Autonomous target vehicles(ATVs),autonomous lane-changing vehicles(ALVs)and HTVs,and ALVs and ATVs.An evolutionary game model is formulated by integrating safety,efficiency,comfort,fatigue,and energy-saving utilities.The strategy selection process of decision-makers is simulated by replicator dynamics equations.Based on the Friedman method,the stability of equilibrium points in different situations is analyzed,and the sensitivity of evolutionary results to utility parameters is clarified.Results show that the strategy profile of(lane-changing,not-giving-way)that causes conflicts only exists temporarily under certain conditions and will not become a long-term stable point.The game characteristics under different scenarios are obviously different;that is,the game evolutionary paths for autonomous vehicles are the most stable,while the adjustment of game strategies for human-driven vehicles is more frequent.Additionally,energy-saving promotes the formation of the strategy profile of(lane-changing,giving-way),while fatigue limits the frequent lane-changing of human-driven vehicles.The evolutionary game model proposed can provide theoretical support for improving the performance of the human-machine co-driving system.展开更多
Cooperation,fairness,trust,and resource coordination are cornerstones of modern civilization,yet their emergence remains inadequately explained,largely due to persistent discrepancies between theoretical predictions a...Cooperation,fairness,trust,and resource coordination are cornerstones of modern civilization,yet their emergence remains inadequately explained,largely due to persistent discrepancies between theoretical predictions and behavioral experiments.Part of this gap may arise from the imitation learning paradigm commonly used in prior theoretical models,which assumes individuals merely copy successful neighbors according to predetermined,fixed rules.This review examines recent advances in evolutionary game dynamics that employ reinforcement learning(RL)as an alternative paradigm.In RL,individuals learn through trial and error and intro spec tively refine their strategies based on environmental feedback.We begin by introducing key concepts in evolutionary game theory and the two learning paradigms,then synthesize progress in applying RL to elucidate cooperation,trust,fairness,optimal resource coordination,and ecological dynamics.Collectively,these studies indicate that RL offers a promising unified framework for understanding the diverse social and ecological phenomena observed in human and natural systems.展开更多
As a special type of dynamic game,Pursuit-Evasion Games(PEGs)have expanded their application range from initial military confrontations to areas such as navigation control and aerospace,demonstrating broad applicabili...As a special type of dynamic game,Pursuit-Evasion Games(PEGs)have expanded their application range from initial military confrontations to areas such as navigation control and aerospace,demonstrating broad applicability and significant value in addressing a wide array of modern complex decision-making problems.Traditional optimal control methods based on differential game theory are classic approaches to solve PEG problems.However,these methods often struggle to perform well in complex environments,nonlinear systems,and situations involving highly uncertain participant behaviors.In recent years,rapidly developing Reinforcement Learning(RL)techniques has provided new avenues for PEG research.RL is capable of adapting to environmental changes through efficient online computation and feedback-driven learning,exhibiting strong generalization capabilities.Therefore,this survey presents a detailed and systematic review of PEG research based on RL methods.First,it classifies and discusses key RL algorithms and theoretical foundations in PEGs according to different forms of strategy learning.Then,it summarizes typical application scenarios,including tactical combat,unmanned systems control,and spacecraft interception,demonstrating the potential and effectiveness of RL in addressing real-world challenges.Finally,the survey explores current challenges and future opportunities in applying RL to PEGs,with the aim of promoting further research on more effective and practical solutions.展开更多
An attack-resilient distributed Nash equilibrium(NE) seeking problem is addressed for noncooperative games of networked systems under malicious cyber-attacks,i.e.,false data injection(FDI) attacks.Different from many ...An attack-resilient distributed Nash equilibrium(NE) seeking problem is addressed for noncooperative games of networked systems under malicious cyber-attacks,i.e.,false data injection(FDI) attacks.Different from many existing distributed NE seeking works,it is practical and challenging to get resilient adaptively distributed NE seeking under unknown and unbounded FDI attacks.An attack-resilient NE seeking algorithm that is distributed(i.e.,independent of global information on the graph's algebraic connectivity,Lipschitz and monotone constants of pseudo-gradients,or number of players),is presented by means of incorporating the consensus-based gradient play with a distributed attack identifier so as to achieve simultaneous NE seeking and attack identification asymptotically.Another key characteristic is that FDI attacks are allowed to be unknown and unbounded.By exploiting nonsmooth analysis and stability theory,the global asymptotic convergence of the developed algorithm to the NE is ensured.Moreover,we extend this design to further consider the attack-resilient NE seeking of double-integrator players.Lastly,numerical simulation and practical experiment results are presented to validate the developed algorithms' effectiveness.展开更多
Large language models(LLMs)have made remarkable advances in natural language processing,demonstrating great potential in modelling structured sequences.However,adapting these capabilities to machine gaming tasks such ...Large language models(LLMs)have made remarkable advances in natural language processing,demonstrating great potential in modelling structured sequences.However,adapting these capabilities to machine gaming tasks such as Go remains challenging due to limitations in strategy generalisation and optimisation efficiency.This paper presents multitype game optimisation(MyGO),a two-stage fine-tuning framework tailored for two-player perfect information board games,exploring the applicability of LLMs to nonlinguistic decision-making domains.In the supervised fine-tuning stage,we propose a unified structural encoding method,action semantic unit(ASU),which efficiently converts heterogeneous game records into discrete token sequences compatible with LLMs.In the reinforcement learning stage,we design TA-PPO(token-level adaptive proximal policy optimisation),an enhanced PPO-based algorithm to address the issue of sparse feedback commonly encountered in game reinforcement learning.Experimental results demonstrate that the fine-tuned models achieve superior or comparable performance to traditional game-playing algorithms in terms of strategy quality,rule generalisation and inference efficiency.This work provides a scalable paradigm for fine-tuning LLMs in complex decision-making tasks and lays a foundation for future research in game AI and generalisable strategy optimisation.展开更多
Multi-domain competition is developing for disintegrating the component of the opponent’s operational system and winning advantage in decision space.Island air defense is a typical multi-domain security problem,which...Multi-domain competition is developing for disintegrating the component of the opponent’s operational system and winning advantage in decision space.Island air defense is a typical multi-domain security problem,which dramatically increases the complexity of decision-making by considering different factors such as multi-stages decisions,multi-domain settings,imperfection information,and uncertain events.However,current research on island air defense security problems is sparse and lacks consideration of key factors.To provide support for assisting human commanders to take wise decisions in a complex environment,we build a multi-domain multi-state island air defense model and propose responding solving algorithms.We study the whole progress of island air defense and propose a multi-domain,multi-stage imperfection information security game that formulates critical characters in the adversarial scenario of island air defense.In addition,considering a bounded rational opponent’s possible strategies,we propose an opponent-aware Monte Carlo counterfactual regret minimization algorithm for learning a robust defensive strategy in the security game.We evaluate our methods in various adversarial scenarios.The results show that our equilibrium learning method can effectively play against an opponent with bounded rationality and significantly outperform some advanced algorithms.展开更多
Weapon target assignment(WTA)problem is a critical problem in multiplatform confrontation.This paper studies a static WTA problem with heterogeneous weapons in multi-platform air combat scenarios,called heterogeneous ...Weapon target assignment(WTA)problem is a critical problem in multiplatform confrontation.This paper studies a static WTA problem with heterogeneous weapons in multi-platform air combat scenarios,called heterogeneous WTA(HWTA)problem.Heterogeneous indicates that the engagement platforms carry multiple kinds of weapons for different tactical purposes.The targets assigned and the weapons used by one side’s platforms will affect the survival probability and capability of the other side’s platforms.The goal of each side in HWTA is to find a solution to determine the kind of weapon used and the target assigned for each platform,so as to maximize their combat effectiveness.The problem is formulated as a two-player noncooperative game model with considering the conflicts between the engaged sides.The Nash equilibrium is an effective solution to the game in which no player has an incentive to deviate.However,the number of pure strategies in HWTA increases exponentially with the engagement platforms.To improve computing efficiency,a double oracle algorithm with constructive heuristic(DOCH)is developed,within which the constructive heuristic is embedded to solve the oracle subproblems efficiently.Numerical experiments are conducted to verify the effectiveness of the DOCH.The results show that the DOCH can find effective strategies for platforms to improve combat effectiveness.Moreover,the DOCH can find high-quality solutions in seconds,significantly outperforming the state-of-the-art algorithms in terms of computational efficiency,especially for large-scale problems.展开更多
This paper suggests a way to improve teamwork and reduce uncertainties in operations by using a game theory approach involving multiple virtual power plants(VPP).A generalized credibility-based fuzzy chance constraint...This paper suggests a way to improve teamwork and reduce uncertainties in operations by using a game theory approach involving multiple virtual power plants(VPP).A generalized credibility-based fuzzy chance constraint programming approach is adopted to address uncertainties stemming from renewable generation and load demand within individual VPPs,while robust optimization techniques manage electricity and thermal price volatilities.Building upon this foundation,a hierarchical Nash-Stackelberg game model is established across multiple VPPs.Within each VPP,a Stackelberg game resolves the strategic interaction between the operator and photovoltaic prosumers(PVP).Among VPPs,a cooperative Nash bargaining model coordinates alliance formation.The problem is decomposed into two subproblems:maximizing coalitional benefits,and allocating cooperative surpluses via payment bargaining,solved distributively using the alternating direction method of multipliers(ADMM).Case studies demonstrate that the proposed strategy significantly enhances the economic efficiency and uncertainty resilience of multi-VPP alliances.展开更多
Discrete memristive neuron systems have attracted considerable attention due to their nonlinear dynamical properties,low computational overhead,and ease of hardware implementation.For the practical engineering applica...Discrete memristive neuron systems have attracted considerable attention due to their nonlinear dynamical properties,low computational overhead,and ease of hardware implementation.For the practical engineering applications of discrete memristive neuron systems,effective control remains a key issue.Parameter identification using intelligent optimization algorithms is an important approach for controlling complex nonlinear systems.However,classical algorithms are prone to falling into local optima and often exhibit high computational complexity,resulting in slow convergence.Therefore,a new algorithm named adaptive chaos game optimization(ACGO)is proposed to address these issues.By introducing a differential evolution mutation strategy and a Cauchy adaptive parameter mechanism,the ACGO algorithm can effectively balance global exploration and local exploitation capabilities.To verify the effectiveness of the proposed algorithm,it is applied to parameter identification in five discrete memristive neuron maps(DMNMs)and compared with seven intelligent optimization algorithms.Simulation results demonstrate that the ACGO algorithm achieves higher accuracy and faster convergence.In addition,an in-depth investigation is conducted into the effects of sample size and objective function on identification performance.The results indicate that setting the sample size to 4 and selecting the mean squared error(MSE)as the objective function can achieve better identification performance and a high level of robustness.展开更多
In the cloud-edge collaborative network,advanced persistent threats(APTs)pose a serious security risk to critical network assets.Although network deception defense can mislead attackers’cognition,its effectiveness de...In the cloud-edge collaborative network,advanced persistent threats(APTs)pose a serious security risk to critical network assets.Although network deception defense can mislead attackers’cognition,its effectiveness depends on dynamically selecting appropriate rotation timings of the deception defense.However,the deployment of deception resources and state updates is not completed instantaneously,and existing methods ignore the state transition delay and the dynamic interaction between the attackers and defenders during the real attack and defense process.To address this,we propose a deception defense timing selection method based on the time-delayed FlipIt game.Firstly,a network state evolution model integrating state transition delay is constructed,and the dynamic transfer process between node states is characterized by a set of delay differential equations.Secondly,a cloud-edge collaborative defense architecture is designed.On this basis,a time-delayed FlipIt game model(TD-FlipIt)is established,and the gate control mechanism is introduced to formalize the defense cooling period as a constraint for the rotation action of deception resources.Subsequently,we use the multi-agent deep deterministic policy gradient(MADDPG)algorithm to solve the rotation strategy for deception defense timing.Experimental results show that the proposed method can effectively optimize the selection of defense timing,ensuring defense effectiveness while reducing resource consumption,and providing effective support for defense in the cloud-edge collaborative environment.展开更多
In strategic decision-making tasks,determining how to assign limited costly resource towards the defender and the attacker is a central problem.However,it is hard for pre-allocated resource assignment to adapt to dyna...In strategic decision-making tasks,determining how to assign limited costly resource towards the defender and the attacker is a central problem.However,it is hard for pre-allocated resource assignment to adapt to dynamic fighting scenarios,and exists situations where the scenario and rule of the Colonel Blotto(CB)game are too restrictive in real world.To address these issues,a support stage is added as supplementary for pre-allocated results,in which a novel two-stage competitive resource assignment problem is formulated based on CB game and stochastic Lanchester equation(SLE).Further,the force attrition in these two stages is formulated as a stochastic progress to consider the complex fighting progress,including the case that the player with fewer resources defeats the player with more resources and wins the battlefield.For solving this two-stage resource assignment problem,nested solving and no-regret learning are proposed to search the optimal resource assignment strategies.Numerical experiments are taken to analyze the effectiveness of the proposed model and study the assignment strategies in various cases.展开更多
Conflicting interests among multiple stakeholders and regulatory imbalances are challenges faced by agricultural ecosystems in highland regions.Current governance models largely fail to reflect complex interactions am...Conflicting interests among multiple stakeholders and regulatory imbalances are challenges faced by agricultural ecosystems in highland regions.Current governance models largely fail to reflect complex interactions among stakeholders,resulting in suboptimal outcomes.To address this problem,this study constructs a four party asymmetric evolutionary game model that includes the government,regulatory agencies,agricultural enterprises,and local residents to analyze stakeholder strategy choices under different benefit conditions.It innovatively incorporates public supervision incentives and dynamic reward-penalty mechanisms into the regulatory system to demonstrate how the government can influence agricultural enterprises’strategic choices through mechanism design.The evolution of the system is analyzed using a dynamic equation,and numerical simulations are performed using data on the ecological baseline of the plateau region and socioeconomic parameters.The findings indicate that effective long-term governance of plateau agricultural ecological security depends on the establishment of a reasonable cost-benefit distribution mechanism.Moreover,public participation in supervision mechanisms can effectively guide agricultural enterprises toward eco-friendly behavior.Additionally,dynamic reward-punishment mechanisms are advantageous for achieving long-term ecological security goals.The scientific value of this study lies in proposing a collaborative regulatory mechanism for plateau agricultural ecological security from a game theory perspective,and its research paradigm provides a theoretical reference for complex regulatory problems in other fields.Based on these findings,we recommend a collaborative governance mechanism that combines dynamic rewards and punishments with public participation to improve governance efficiency and achieve the sustainable development of the plateau’s economy and ecology.展开更多
Building on the existing symmetric quantization model of the dynamic Cournot duopoly game(CDG)with asymmetric information,we extend it to an asymmetric quantization model and study the stability of the quantum Bayesia...Building on the existing symmetric quantization model of the dynamic Cournot duopoly game(CDG)with asymmetric information,we extend it to an asymmetric quantization model and study the stability of the quantum Bayesian Nash equilibrium(QBNE)under heterogeneous expectations.We analyze the influence of various parameters on the stability of QBNE,with a particular focus on the impact of the parameter α on system stability.The results show that when α1,the quantum strategy of the symmetric quantization model is more conducive to stabilizing the market.展开更多
This paper studies an indefinite mean-field game with Markov jump parameters,where all agents'diffusion terms depend on control variables and both state and control average terms(x.(N),u.(N))are considered.O...This paper studies an indefinite mean-field game with Markov jump parameters,where all agents'diffusion terms depend on control variables and both state and control average terms(x.(N),u.(N))are considered.One notable aspect is the relaxation of the assumption regarding the positivity or non-negativity of weight matrices within costs,allowing for zero or even negative values.By virtue of mean-field methods and decomposition techniques,we have derived decentralized strategies presented by Hamiltonian systems and a new type of consistency condition system.These systems consist of fully coupled regime-switching forward-backward stochastic differential equations that do not conform to the Monotonicity condition.The well-posedness of these strategies is established by employing a relaxed compensator method with an easily verifiable Condition(RC)and the decomposition technique.Furthermore,we demonstrate that the resulting decentralized strategies achieve anϵ-Nash equilibrium in the indefinite case without any assumptions on admissible control sets using novel estimates of the disturbed state and cost function.Finally,our theoretical results are applied to resolve a class of mean-variance portfolio selection problems.We provide corresponding numerical simulation results and economic explanations.展开更多
Dear Editor,This letter deals with incentive design problem for noncooperative dynamical systems to achieve social welfare maximization with uncertain types of mixed dynamics.The agents are allowed to liberally choose...Dear Editor,This letter deals with incentive design problem for noncooperative dynamical systems to achieve social welfare maximization with uncertain types of mixed dynamics.The agents are allowed to liberally choose either the pseudo-gradient or best-response dynamics for continuous-time decision-making.展开更多
Dear Editor,This letter proposes a reinforcement learning-based predictive learning algorithm for unknown continuous-time nonlinear systems with observation loss.Firstly,we construct a temporal nonzero-sum game over p...Dear Editor,This letter proposes a reinforcement learning-based predictive learning algorithm for unknown continuous-time nonlinear systems with observation loss.Firstly,we construct a temporal nonzero-sum game over predictive control input sequences,deriving multiple optimal predictive control input sequences from its solution.展开更多
基金study was supported by the National Key Research and Development Program of China(Grant No.2023YFC2907202)the National Natural Science Foundation of China(Grant No.52404116)the Postdoctoral Fellowship Program of CPSF(Grant No.GZB20240129).
摘要In this paper,a self-developed master-slave follow-up disc cutter is used to conduct rock-breaking tests on hard sandstone samples.Different working parameters were employed in the tests(e.g.cutting depth,cutting speed,cutting angle,and rotational speed)in order to explore their influences on cutting performance.The results indicate that the thrust,torque,vibration velocity,and roughness all increased continuously with increase of the propulsion speed and cutting depth.At the same time,the specific energy consumption was found to decrease continuously.As the rotational speed was increased,the thrust increased at first and then decreased.In contrast,the torque and roughness continuously decreased,and the specific energy consumption and vibration speed continuously increased.When the cutting angle was increased,the thrust remained unchanged.However,the torque,specific energy consumption,and vibration speed all decreased continuously,and the roughness increased continuously.The temperature of the surface of the cutting tool was found to be relatively uniformly distributed during the rock-breaking process;the highest temperatures generated were in the range of 200-300℃.As the propulsion speed,cutting depth,and cutting angle were increased,the proportion of tensile fractures produced appeared to increase and the proportion of shear fractures decreased.As the rotational speed was increased,the proportion of tensile fractures decreased and the proportion of shear fractures increased.The results could provide useful information on the rock-breaking behavior involved and can be used to offer technical support for engineers using master-slave follow-up disc cutters in the field.
基金supported by the National Key R&D Pro-gram of China:Gravitational Wave Detection Project(Nos.2021YFC2026,2021YFC2202601,2021YFC2202603)the National Natural Science Foundation of China(Nos.12172288 and 12472046)。
摘要The problem of maneuvering for a servicing spacecraft(inspector)to inspect a noncooperative spacecraft(evader)in cislunar space is investigated in this paper.The evader,which may be a malfunctioning or uncontrolled satellite,introduces uncertainties due to its potential maneuvering capabilities.To address this challenge,the scenario is modeled as a special orbital game,incorporating the unique complexities of the cislunar environment.A variable-duration,turn-based inspection and anti-inspection game model is designed.The model defines both players'rules,constraints,and victory conditions,providing a framework for non-cooperative inspection.Strategies for both players are developed and validated based on their dynamical properties.The inspector's strategy integrates two-body Lambert transfers with shooting methods,while the evader's strategy aims to maximize the inspector's fuel consumption.Simulation results show that the evader's optimal strategy involves deliberate fluctuations in its lunar periapsis altitude,with the inspector's requiredΔV up to eight times greater than the evader's.The impact of game constraints is evaluated,and the effectiveness of deploying the inspector in low lunar orbit is compared with the inspector at the Earth-Moon Lagrange point L1.The strengths and weaknesses of both are shown.These findings provide valuable insights for future orbital servicing and orbital games.
基金The Natural Science Foundation of Heilongjiang Province(No.LH2023E011)Open Fund of National Key Laboratory of Green and Long-Life Road Engineering in Extreme Environment in Changsha University of Science and Technology(No.kfj230105).
摘要To uncover the decision-making mechanisms and evolutionary dynamics of multiple stakeholders in highway noise pollution control,a three-party evolutionary game model involving the government,operators,and the public is constructed.The operation period is divided into different stages for differentiated analysis.A simulation analysis was performed on the Lituo sinking section of the Beijing-Hong Kong-Macao Highway to assess the impact of variations in critical elements on the system.The results indicate that the Lituo sinking section of the Beijing-Hong Kong-Macao Highway is currently in its early stage of development,with the corresponding strategies being active regulation,excessive emissions,and supervision.When the cost of the government’s active regulation decreases from 1×105 to 5×104 yuan,the system converges more rapidly toward the active regulation strategy.When the cost of the operator’s excessive emissions increases from 14.08×106 to 20.00×106 yuan,the system drives the operator toward the standardized emission strategy.In addition,when the cost of public supervision decreases from 15×104 to 5×104 yuan and the compensation paid by operators to the public increases from 1.288×106 to 2.576×106 yuan,the system converges more quickly toward the supervision strategy.The cost of the operator’s excessive emissions serves as the core decision variable for achieving the ideal equilibrium in the three-party game involving government active regulation,operator standardized emissions,and public supervision.
基金co-supported by the National Natural Science Foundation of China(No.62273119)。
摘要This paper proposes a two-phase game guidance strategy for the three-body confrontation scenario according to the linear quadratic differential game method,which includes an Attacker,an Interceptor,and a Target.The interception probabilities between the Attacker and the Interceptor-Target team are estimated using the probability density function.The desired zero-effort miss distances associated with the interception probabilities for the three-body conflict are acquired by virtue of the gradient descent method.The game combat is divided into two phases by introducing the switching time.In Phase 1,a differential game strategy is developed to guide the zero-control miss distance between the Attacker and the Interceptor-Target into the desired position,which guarantees that the Attacker has the maximizing probability of intercepting the Target and the minimizing probability of being captured by the Interceptor.In Phase 2,a differential game guidance strategy is proposed to ensure that the Attacker evades the pursuit of the Interceptor and intercepts the Target at the preset impact angle.Finally,numerical simulation verifies the effectiveness of the two-stage game guidance strategy.
基金supported by the National Natural Science Foundation of China(52172314)the Natural Science Foundation of Shandong Province,China(ZR2024MG021 and ZR2024QG023)the Special Funding Project of Taishan Scholar Engineering.
摘要To alleviate the conflicts resulting from lane-changing in human-machine co-driving,this study investigates four game scenarios:human-driven lanechanging vehicles(HLVs)and human-driven target vehicles(HTVs),HLVs and Autonomous target vehicles(ATVs),autonomous lane-changing vehicles(ALVs)and HTVs,and ALVs and ATVs.An evolutionary game model is formulated by integrating safety,efficiency,comfort,fatigue,and energy-saving utilities.The strategy selection process of decision-makers is simulated by replicator dynamics equations.Based on the Friedman method,the stability of equilibrium points in different situations is analyzed,and the sensitivity of evolutionary results to utility parameters is clarified.Results show that the strategy profile of(lane-changing,not-giving-way)that causes conflicts only exists temporarily under certain conditions and will not become a long-term stable point.The game characteristics under different scenarios are obviously different;that is,the game evolutionary paths for autonomous vehicles are the most stable,while the adjustment of game strategies for human-driven vehicles is more frequent.Additionally,energy-saving promotes the formation of the strategy profile of(lane-changing,giving-way),while fatigue limits the frequent lane-changing of human-driven vehicles.The evolutionary game model proposed can provide theoretical support for improving the performance of the human-machine co-driving system.
基金supported by the National Natural Science Foundation of China(Grants Nos.12075144,12165014)the Fundamental Research Funds for the Central Universities(Grant No.GK202401002)the Key Research and Development Program of Ningxia in China(Grant No.2021BEB04032)。
摘要Cooperation,fairness,trust,and resource coordination are cornerstones of modern civilization,yet their emergence remains inadequately explained,largely due to persistent discrepancies between theoretical predictions and behavioral experiments.Part of this gap may arise from the imitation learning paradigm commonly used in prior theoretical models,which assumes individuals merely copy successful neighbors according to predetermined,fixed rules.This review examines recent advances in evolutionary game dynamics that employ reinforcement learning(RL)as an alternative paradigm.In RL,individuals learn through trial and error and intro spec tively refine their strategies based on environmental feedback.We begin by introducing key concepts in evolutionary game theory and the two learning paradigms,then synthesize progress in applying RL to elucidate cooperation,trust,fairness,optimal resource coordination,and ecological dynamics.Collectively,these studies indicate that RL offers a promising unified framework for understanding the diverse social and ecological phenomena observed in human and natural systems.
基金supported by the National Science and Technology Major Project,China(No.2022ZD0119703)the National Natural Science Foundation of China(No.62273044)the National Natural Science Foundation of China National Science Fund for Distinguished Young Scholars(No.62025301)。
摘要As a special type of dynamic game,Pursuit-Evasion Games(PEGs)have expanded their application range from initial military confrontations to areas such as navigation control and aerospace,demonstrating broad applicability and significant value in addressing a wide array of modern complex decision-making problems.Traditional optimal control methods based on differential game theory are classic approaches to solve PEG problems.However,these methods often struggle to perform well in complex environments,nonlinear systems,and situations involving highly uncertain participant behaviors.In recent years,rapidly developing Reinforcement Learning(RL)techniques has provided new avenues for PEG research.RL is capable of adapting to environmental changes through efficient online computation and feedback-driven learning,exhibiting strong generalization capabilities.Therefore,this survey presents a detailed and systematic review of PEG research based on RL methods.First,it classifies and discusses key RL algorithms and theoretical foundations in PEGs according to different forms of strategy learning.Then,it summarizes typical application scenarios,including tactical combat,unmanned systems control,and spacecraft interception,demonstrating the potential and effectiveness of RL in addressing real-world challenges.Finally,the survey explores current challenges and future opportunities in applying RL to PEGs,with the aim of promoting further research on more effective and practical solutions.
基金supported in part by the National Natural Science Foundation of China(62373022,U2241217,62141604)Beijing Natural Science Foundation(4252043,JQ23019)+4 种基金the Fundamental Research Funds for the Central Universities(JKF-2025037448805,JKF-2025086098295)the Aeronautical Science Fund(2023Z034051001)the Academic Excellence Foundation of BUAA for Ph.D. Studentsthe Science and Technology Innovation2030—Key Project of New Generation Artificial Intelligence(2020AAA0108200)the National Key Research and Development Program of China(2022YFB3305600)。
摘要An attack-resilient distributed Nash equilibrium(NE) seeking problem is addressed for noncooperative games of networked systems under malicious cyber-attacks,i.e.,false data injection(FDI) attacks.Different from many existing distributed NE seeking works,it is practical and challenging to get resilient adaptively distributed NE seeking under unknown and unbounded FDI attacks.An attack-resilient NE seeking algorithm that is distributed(i.e.,independent of global information on the graph's algebraic connectivity,Lipschitz and monotone constants of pseudo-gradients,or number of players),is presented by means of incorporating the consensus-based gradient play with a distributed attack identifier so as to achieve simultaneous NE seeking and attack identification asymptotically.Another key characteristic is that FDI attacks are allowed to be unknown and unbounded.By exploiting nonsmooth analysis and stability theory,the global asymptotic convergence of the developed algorithm to the NE is ensured.Moreover,we extend this design to further consider the attack-resilient NE seeking of double-integrator players.Lastly,numerical simulation and practical experiment results are presented to validate the developed algorithms' effectiveness.
基金supported in part by the National Natural Science Foundation of China under Grants 62276285 and 62236011。
摘要Large language models(LLMs)have made remarkable advances in natural language processing,demonstrating great potential in modelling structured sequences.However,adapting these capabilities to machine gaming tasks such as Go remains challenging due to limitations in strategy generalisation and optimisation efficiency.This paper presents multitype game optimisation(MyGO),a two-stage fine-tuning framework tailored for two-player perfect information board games,exploring the applicability of LLMs to nonlinguistic decision-making domains.In the supervised fine-tuning stage,we propose a unified structural encoding method,action semantic unit(ASU),which efficiently converts heterogeneous game records into discrete token sequences compatible with LLMs.In the reinforcement learning stage,we design TA-PPO(token-level adaptive proximal policy optimisation),an enhanced PPO-based algorithm to address the issue of sparse feedback commonly encountered in game reinforcement learning.Experimental results demonstrate that the fine-tuned models achieve superior or comparable performance to traditional game-playing algorithms in terms of strategy quality,rule generalisation and inference efficiency.This work provides a scalable paradigm for fine-tuning LLMs in complex decision-making tasks and lays a foundation for future research in game AI and generalisable strategy optimisation.
基金supported by the National Natural Science Foundation of China(92271108,61702528,61806212,62173336).
摘要Multi-domain competition is developing for disintegrating the component of the opponent’s operational system and winning advantage in decision space.Island air defense is a typical multi-domain security problem,which dramatically increases the complexity of decision-making by considering different factors such as multi-stages decisions,multi-domain settings,imperfection information,and uncertain events.However,current research on island air defense security problems is sparse and lacks consideration of key factors.To provide support for assisting human commanders to take wise decisions in a complex environment,we build a multi-domain multi-state island air defense model and propose responding solving algorithms.We study the whole progress of island air defense and propose a multi-domain,multi-stage imperfection information security game that formulates critical characters in the adversarial scenario of island air defense.In addition,considering a bounded rational opponent’s possible strategies,we propose an opponent-aware Monte Carlo counterfactual regret minimization algorithm for learning a robust defensive strategy in the security game.We evaluate our methods in various adversarial scenarios.The results show that our equilibrium learning method can effectively play against an opponent with bounded rationality and significantly outperform some advanced algorithms.
基金supported by the National Natural Science Foundation of China(72571094,71871079,72271076,72001004)Anhui Provincial Natural Science Foundation(2308085QG233)+1 种基金Anhui Province Postdoctoral Research Activities Funds(2022B587)Talent Research Fund of Hefei University(24RC75).
摘要Weapon target assignment(WTA)problem is a critical problem in multiplatform confrontation.This paper studies a static WTA problem with heterogeneous weapons in multi-platform air combat scenarios,called heterogeneous WTA(HWTA)problem.Heterogeneous indicates that the engagement platforms carry multiple kinds of weapons for different tactical purposes.The targets assigned and the weapons used by one side’s platforms will affect the survival probability and capability of the other side’s platforms.The goal of each side in HWTA is to find a solution to determine the kind of weapon used and the target assigned for each platform,so as to maximize their combat effectiveness.The problem is formulated as a two-player noncooperative game model with considering the conflicts between the engaged sides.The Nash equilibrium is an effective solution to the game in which no player has an incentive to deviate.However,the number of pure strategies in HWTA increases exponentially with the engagement platforms.To improve computing efficiency,a double oracle algorithm with constructive heuristic(DOCH)is developed,within which the constructive heuristic is embedded to solve the oracle subproblems efficiently.Numerical experiments are conducted to verify the effectiveness of the DOCH.The results show that the DOCH can find effective strategies for platforms to improve combat effectiveness.Moreover,the DOCH can find high-quality solutions in seconds,significantly outperforming the state-of-the-art algorithms in terms of computational efficiency,especially for large-scale problems.
基金supported by Science and Technology Project of SGCC(Research on Distributed Cooperative Control of Virtual Power Plants Based on Hybrid Game)(5700-202418337A-2-1-ZX).
摘要This paper suggests a way to improve teamwork and reduce uncertainties in operations by using a game theory approach involving multiple virtual power plants(VPP).A generalized credibility-based fuzzy chance constraint programming approach is adopted to address uncertainties stemming from renewable generation and load demand within individual VPPs,while robust optimization techniques manage electricity and thermal price volatilities.Building upon this foundation,a hierarchical Nash-Stackelberg game model is established across multiple VPPs.Within each VPP,a Stackelberg game resolves the strategic interaction between the operator and photovoltaic prosumers(PVP).Among VPPs,a cooperative Nash bargaining model coordinates alliance formation.The problem is decomposed into two subproblems:maximizing coalitional benefits,and allocating cooperative surpluses via payment bargaining,solved distributively using the alternating direction method of multipliers(ADMM).Case studies demonstrate that the proposed strategy significantly enhances the economic efficiency and uncertainty resilience of multi-VPP alliances.
基金supported by the National Natural Science Foundation of China(Grant Nos.62501516 and 62572419)the Natural Science Foundation of Hunan Province(Grant Nos.2025JJ50391 and 2025JJ50392)the Research Foundation of the Education Department of Hunan Province(Grant Nos.23B0131 and 24A0124)。
摘要Discrete memristive neuron systems have attracted considerable attention due to their nonlinear dynamical properties,low computational overhead,and ease of hardware implementation.For the practical engineering applications of discrete memristive neuron systems,effective control remains a key issue.Parameter identification using intelligent optimization algorithms is an important approach for controlling complex nonlinear systems.However,classical algorithms are prone to falling into local optima and often exhibit high computational complexity,resulting in slow convergence.Therefore,a new algorithm named adaptive chaos game optimization(ACGO)is proposed to address these issues.By introducing a differential evolution mutation strategy and a Cauchy adaptive parameter mechanism,the ACGO algorithm can effectively balance global exploration and local exploitation capabilities.To verify the effectiveness of the proposed algorithm,it is applied to parameter identification in five discrete memristive neuron maps(DMNMs)and compared with seven intelligent optimization algorithms.Simulation results demonstrate that the ACGO algorithm achieves higher accuracy and faster convergence.In addition,an in-depth investigation is conducted into the effects of sample size and objective function on identification performance.The results indicate that setting the sample size to 4 and selecting the mean squared error(MSE)as the objective function can achieve better identification performance and a high level of robustness.
基金supported in part by the National Key Research and Development Program of China under Grants 2024YFB2906704 and 2023YFB2903902in part by the State Key Laboratory of Advanced Communication Networks underGrant FFX24641X028in part by the Science and Technology Innovation Leading Talents Subsidy Project of Central Plains under Grant 244200510038.
摘要In the cloud-edge collaborative network,advanced persistent threats(APTs)pose a serious security risk to critical network assets.Although network deception defense can mislead attackers’cognition,its effectiveness depends on dynamically selecting appropriate rotation timings of the deception defense.However,the deployment of deception resources and state updates is not completed instantaneously,and existing methods ignore the state transition delay and the dynamic interaction between the attackers and defenders during the real attack and defense process.To address this,we propose a deception defense timing selection method based on the time-delayed FlipIt game.Firstly,a network state evolution model integrating state transition delay is constructed,and the dynamic transfer process between node states is characterized by a set of delay differential equations.Secondly,a cloud-edge collaborative defense architecture is designed.On this basis,a time-delayed FlipIt game model(TD-FlipIt)is established,and the gate control mechanism is introduced to formalize the defense cooling period as a constraint for the rotation action of deception resources.Subsequently,we use the multi-agent deep deterministic policy gradient(MADDPG)algorithm to solve the rotation strategy for deception defense timing.Experimental results show that the proposed method can effectively optimize the selection of defense timing,ensuring defense effectiveness while reducing resource consumption,and providing effective support for defense in the cloud-edge collaborative environment.
基金supported by the National Natural Science Foundation of China(61702528,61806212,62173336)。
摘要In strategic decision-making tasks,determining how to assign limited costly resource towards the defender and the attacker is a central problem.However,it is hard for pre-allocated resource assignment to adapt to dynamic fighting scenarios,and exists situations where the scenario and rule of the Colonel Blotto(CB)game are too restrictive in real world.To address these issues,a support stage is added as supplementary for pre-allocated results,in which a novel two-stage competitive resource assignment problem is formulated based on CB game and stochastic Lanchester equation(SLE).Further,the force attrition in these two stages is formulated as a stochastic progress to consider the complex fighting progress,including the case that the player with fewer resources defeats the player with more resources and wins the battlefield.For solving this two-stage resource assignment problem,nested solving and no-regret learning are proposed to search the optimal resource assignment strategies.Numerical experiments are taken to analyze the effectiveness of the proposed model and study the assignment strategies in various cases.
基金supported by the Special Program of National Social Science Fund of China[Grant No.25VHQ035].
摘要Conflicting interests among multiple stakeholders and regulatory imbalances are challenges faced by agricultural ecosystems in highland regions.Current governance models largely fail to reflect complex interactions among stakeholders,resulting in suboptimal outcomes.To address this problem,this study constructs a four party asymmetric evolutionary game model that includes the government,regulatory agencies,agricultural enterprises,and local residents to analyze stakeholder strategy choices under different benefit conditions.It innovatively incorporates public supervision incentives and dynamic reward-penalty mechanisms into the regulatory system to demonstrate how the government can influence agricultural enterprises’strategic choices through mechanism design.The evolution of the system is analyzed using a dynamic equation,and numerical simulations are performed using data on the ecological baseline of the plateau region and socioeconomic parameters.The findings indicate that effective long-term governance of plateau agricultural ecological security depends on the establishment of a reasonable cost-benefit distribution mechanism.Moreover,public participation in supervision mechanisms can effectively guide agricultural enterprises toward eco-friendly behavior.Additionally,dynamic reward-punishment mechanisms are advantageous for achieving long-term ecological security goals.The scientific value of this study lies in proposing a collaborative regulatory mechanism for plateau agricultural ecological security from a game theory perspective,and its research paradigm provides a theoretical reference for complex regulatory problems in other fields.Based on these findings,we recommend a collaborative governance mechanism that combines dynamic rewards and punishments with public participation to improve governance efficiency and achieve the sustainable development of the plateau’s economy and ecology.
基金Project supported by the National Natural Science Foundation of China(Grant No.12461054)the Science and Technology Key Foundation of Guizhou Province,China(Grant No.2025089)。
摘要Building on the existing symmetric quantization model of the dynamic Cournot duopoly game(CDG)with asymmetric information,we extend it to an asymmetric quantization model and study the stability of the quantum Bayesian Nash equilibrium(QBNE)under heterogeneous expectations.We analyze the influence of various parameters on the stability of QBNE,with a particular focus on the impact of the parameter α on system stability.The results show that when α1,the quantum strategy of the symmetric quantization model is more conducive to stabilizing the market.
基金supported by the National Key Research and Development Program of China(2023YFA1009200)the National Natural Science Foundation of China(12401583,12571482,12521001)+2 种基金the Taishan Scholars Climbing Program of Shandong(TSPD20210302)the Basic Research Program of Jiangsu(BK20240416)the General Program of Philosophy and Social Science Research(PSSR)of Shandong Higher Education Institutions(2024ZSMS007)。
摘要This paper studies an indefinite mean-field game with Markov jump parameters,where all agents'diffusion terms depend on control variables and both state and control average terms(x.(N),u.(N))are considered.One notable aspect is the relaxation of the assumption regarding the positivity or non-negativity of weight matrices within costs,allowing for zero or even negative values.By virtue of mean-field methods and decomposition techniques,we have derived decentralized strategies presented by Hamiltonian systems and a new type of consistency condition system.These systems consist of fully coupled regime-switching forward-backward stochastic differential equations that do not conform to the Monotonicity condition.The well-posedness of these strategies is established by employing a relaxed compensator method with an easily verifiable Condition(RC)and the decomposition technique.Furthermore,we demonstrate that the resulting decentralized strategies achieve anϵ-Nash equilibrium in the indefinite case without any assumptions on admissible control sets using novel estimates of the disturbed state and cost function.Finally,our theoretical results are applied to resolve a class of mean-variance portfolio selection problems.We provide corresponding numerical simulation results and economic explanations.
基金supported by the National Key R&D Program of China(2022ZD0119604)the National Natural Science Foundation of China(NSFC)(62222308,62173181,62221004)the Basic Program of Jiangsu(BK20253024)。
摘要Dear Editor,This letter deals with incentive design problem for noncooperative dynamical systems to achieve social welfare maximization with uncertain types of mixed dynamics.The agents are allowed to liberally choose either the pseudo-gradient or best-response dynamics for continuous-time decision-making.
基金supported by the National Natural Science Foundation of China(62433014,62373287,62573324,62333005,62273255)in part by the International Exchange Program for Graduate Students of Tongji University(4360143306)+3 种基金in part by the Fundamental Research Funds for Central Universities(22120230311)supported by DeutscheForschungsgemeinschaft(DFG,German Research Foundation)under Germany’s Excellence Strategy(EXC 2075390740016,468094890)support by the Stuttgart Center for Simulation Science(SimTech)the International Max Planck Research School for Intelligent Systems(IMPRS-IS)for supporting Y.Xie。
摘要Dear Editor,This letter proposes a reinforcement learning-based predictive learning algorithm for unknown continuous-time nonlinear systems with observation loss.Firstly,we construct a temporal nonzero-sum game over predictive control input sequences,deriving multiple optimal predictive control input sequences from its solution.