In order to achieve comprehensive,highly efficient,and multi-objective precise optimization of fiber structural parameters and further enhance the transmission capacity of optical communication systems,a homogeneous w...In order to achieve comprehensive,highly efficient,and multi-objective precise optimization of fiber structural parameters and further enhance the transmission capacity of optical communication systems,a homogeneous weakly coupled seven-core fiber based on trench-assisted structures is designed.Particle Swarm Optimization(PSO)is introduced to replace traditional empirical designs or local scanning methods.First,a multi-objective fitness function incorporating constraints such as dispersion,cutoff wavelength,ef-fective mode field area,and coating loss is established.Then,the algorithm performs a global search to pre-cisely determine the optimal structural parameters under standard dimensional constraints.Simulation results demonstrate that with a fiber core pitch of 45μm,the optimized fiber achieves an ultra-low inter-core crosstalk of below−90 dB/km at a wavelength of 1550 nm.This design scheme not only effectively resolves the conflict between crosstalk suppression and spatial utilization in multi-core fibers but also proves the effi-ciency and reliability of the PSO algorithm in complex fiber structural design,providing an important theor-etical basis and technical support for the research and manufacturing of ultra-large-capacity optical commu-nication systems.展开更多
Although supplying extensive design space,the curse of dimensionality restricts the widespread application of largescale topology optimization in practical engineering.Various acceleration techniques have been integra...Although supplying extensive design space,the curse of dimensionality restricts the widespread application of largescale topology optimization in practical engineering.Various acceleration techniques have been integrated with topology optimization,achieving significant attention and progress in large-scale problems.This work aims to investigate how much benefit can be obtained by combining parallel computing and machine learning techniques to enhance the efficiency of large-scale topology optimization algorithms.Accordingly,a parallel problem independent machine learning(PIML)-enhanced topology optimization method is proposed.The PIML model substantially reduces the dimension of the condensed stiffness matrix and its computational cost,and parallel computing reduces the workload per process and enables the application of a parallel multigrid solver.Besides,several techniques,such as matrix-free implementation,direct condensation of uniform coarse elements,and adjusting computational resource limits,have been developed to enhance computational efficiency.The weak scaling efficiency,strong scaling speedup,and maximum achievable efficiency of the proposed method are validated across multiple numerical examples,showing significant improvement in the tractable problem size and solution efficiency compared to traditional topology optimization algorithms.展开更多
Accurate estimation of photovoltaic(PV)parameters is essential for optimizing solar module perfor-mance and enhancing resource efficiency in renewable energy systems.This study presents a process innovation by introdu...Accurate estimation of photovoltaic(PV)parameters is essential for optimizing solar module perfor-mance and enhancing resource efficiency in renewable energy systems.This study presents a process innovation by introducing,for the first time,the Triangulation Topology Aggregation Optimizer(TTAO)integrated with parallel computing to address PV parameter estimation challenges.The effectiveness and robustness of TTAO are rigorously evaluated using two standard benchmark datasets(KC200GT and R.T.C.France solar cells)and a real-world dataset(Poly70W solar module)under single-,double-,and triple-diode configurations.Results show that TTAO consistently achieves superior accuracy by producing the lowest RMSE values and faster convergence compared to state-of-the-art metaheuristic algorithms.In addition,the integration of parallel computing significantly enhances computational efficiency,reducing execution time by up to 85%without compromising accuracy.Validation using real-world data further demonstrates TTAO’s adaptability and practical relevance in renewable energy systems,effectively bridging the gap between theoretical modeling and real-world implementation for PV system monitoring and optimization,contributing to climate mitigation through improved solar energy performance.展开更多
Magnetic Resonance Imaging(MRI)has a pivotal role in medical image analysis,for its ability in supporting disease detection and diagnosis.Fuzzy C-Means(FCM)clustering is widely used for MRI segmentation due to its abi...Magnetic Resonance Imaging(MRI)has a pivotal role in medical image analysis,for its ability in supporting disease detection and diagnosis.Fuzzy C-Means(FCM)clustering is widely used for MRI segmentation due to its ability to handle image uncertainty.However,the latter still has countless limitations,including sensitivity to initialization,susceptibility to local optima,and high computational cost.To address these limitations,this study integrates Grey Wolf Optimization(GWO)with FCM to enhance cluster center selection,improving segmentation accuracy and robustness.Moreover,to further refine optimization,Fuzzy Entropy Clustering was utilized for its distinctive features from other traditional objective functions.Fuzzy entropy effectively quantifies uncertainty,leading to more well-defined clusters,improved noise robustness,and better preservation of anatomical structures in MRI images.Despite these advantages,the iterative nature of GWO and FCM introduces significant computational overhead,which restricts their applicability to high-resolution medical images.To overcome this bottleneck,we propose a Parallelized-GWO-based FCM(P-GWO-FCM)approach using GPU acceleration,where both GWO optimization and FCM updates(centroid computation and membership matrix updates)are parallelized.By concurrently executing these processes,our approach efficiently distributes the computational workload,significantly reducing execution time while maintaining high segmentation accuracy.The proposed parallel method,P-GWO-FCM,was evaluated on both simulated and clinical brain MR images,focusing on segmenting white matter,gray matter,and cerebrospinal fluid regions.The results indicate significant improvements in segmentation accuracy,achieving a Jaccard Similarity(JS)of 0.92,a Partition Coefficient Index(PCI)of 0.91,a Partition Entropy Index(PEI)of 0.25,and a Davies-Bouldin Index(DBI)of 0.30.Experimental comparisons demonstrate that P-GWO-FCM outperforms existing methods in both segmentation accuracy and computational efficiency,making it a promising solution for real-time medical image segmentation.展开更多
The Dynamical Density Functional Theory(DDFT)algorithm,derived by associating classical Density Functional Theory(DFT)with the fundamental Smoluchowski dynamical equation,describes the evolution of inhomo-geneous flui...The Dynamical Density Functional Theory(DDFT)algorithm,derived by associating classical Density Functional Theory(DFT)with the fundamental Smoluchowski dynamical equation,describes the evolution of inhomo-geneous fluid density distributions over time.It plays a significant role in studying the evolution of density distributions over time in inhomogeneous systems.The Sunway Bluelight II supercomputer,as a new generation of China’s developed supercomputer,possesses powerful computational capabilities.Porting and optimizing industrial software on this platform holds significant importance.For the optimization of the DDFT algorithm,based on the Sunway Bluelight II supercomputer and the unique hardware architecture of the SW39000 processor,this work proposes three acceleration strategies to enhance computational efficiency and performance,including direct parallel optimization,local-memory constrained optimization for CPEs,and multi-core groups collaboration and communication optimization.This method combines the characteristics of the program’s algorithm with the unique hardware architecture of the Sunway Bluelight II supercomputer,optimizing the storage and transmission structures to achieve a closer integration of software and hardware.For the first time,this paper presents Sunway-Dynamical Density Functional Theory(SW-DDFT).Experimental results show that SW-DDFT achieves a speedup of 6.67 times within a single-core group compared to the original DDFT implementation,with six core groups(a total of 384 CPEs),the maximum speedup can reach 28.64 times,and parallel efficiency can reach 71%,demonstrating excellent acceleration performance.展开更多
This paper addresses the parallel control of autonomous surface vehicles subject to external disturbances,state constraints,and input constraints in complex ocean environments with multiple obstacles.A safety-certifie...This paper addresses the parallel control of autonomous surface vehicles subject to external disturbances,state constraints,and input constraints in complex ocean environments with multiple obstacles.A safety-certified parallel model predictive control scheme with collision-avoiding capability is proposed for autonomous surface vehicles in the framework of parallel control.Specifically,an extended state observer is designed by leveraging historical and real-time data for concurrent learning to map the motion of autonomous surface vehicles from its physical system to its artificial counterpart.A parallel model predictive control law is developed on the basis of the artificial system for both physical and artificial autonomous surface vehicles to realize virtual-physical tracking control of vehicles subject to state and input constraints.To ensure safety,highorder discrete control barrier functions are encoded in the parallel model predictive control law as safety constraints such that collision avoidance with obstacles can be achieved.A recedinghorizon constrained optimization problem is constructed with the safety constraints encoded by control barrier functions for parallel model predictive control of autonomous surface vehicles and solved via neurodynamic optimization with projection neural networks.The effectiveness and characteristics of the proposed method are demonstrated via simulations for the safe trajectory tracking and automatic berthing of autonomous surface vehicles.展开更多
High-dimensional and incomplete(HDI) matrices are commonly encountered in various big data-related applications for illustrating the complex interactions among numerous entities, like the user-item interactions in a c...High-dimensional and incomplete(HDI) matrices are commonly encountered in various big data-related applications for illustrating the complex interactions among numerous entities, like the user-item interactions in a commercial recommender system or the user-user interactions in a social network services system. The factorization of such an HDI matrix can embed the involved entities into the low-dimensional feature space for acquiring their principal representation, which is a vital task in various application scenes and is often established through the Latent Factor Analysis(LFA). Nevertheless, an HDI matrix can be huge when the corresponding application explodes to involve millions of users, items, or other interactive nodes. In this case, a parallel optimization algorithm is desired for raising the scalability and time efficiency of an LFA model. This paper provides a comprehensive review of the existing parallel optimization algorithms for the LFA model. Specifically, it performs: 1) discussion and summary of these algorithms based on computing architecture and mode, 2) empirical studies of representative models, and3) summary of the current challenges and future directions in this domain. This survey aims to offer an exhaustive review of Parallel Optimization Algorithms for High-Dimensional and Incomplete Matrix Factorization, thereby fostering further research in this field.展开更多
Developing parallel applications on heterogeneous processors is facing the challenges of 'memory wall',due to limited capacity of local storage,limited bandwidth and long latency for memory access. Aiming at t...Developing parallel applications on heterogeneous processors is facing the challenges of 'memory wall',due to limited capacity of local storage,limited bandwidth and long latency for memory access. Aiming at this problem,a parallelization approach was proposed with six memory optimization schemes for CG,four schemes of them aiming at all kinds of sparse matrix-vector multiplication (SPMV) operation. Conducted on IBM QS20,the parallelization approach can reach up to 21 and 133 times speedups with size A and B,respectively,compared with single power processor element. Finally,the conclusion is drawn that the peak bandwidth of memory access on Cell BE can be obtained in SPMV,simple computation is more efficient on heterogeneous processors and loop-unrolling can hide local storage access latency while executing scalar operation on SIMD cores.展开更多
Dimensional synthesis is one of the most difficult issues in the field of parallel robots with actuation redundancy. To deal with the optimal design of a redundantly actuated parallel robot used for ankle rehabilitati...Dimensional synthesis is one of the most difficult issues in the field of parallel robots with actuation redundancy. To deal with the optimal design of a redundantly actuated parallel robot used for ankle rehabilitation, a methodology of dimensional synthesis based on multi-objective optimization is presented. First, the dimensional synthesis of the redundant parallel robot is formulated as a nonlinear constrained multi-objective optimization problem. Then four objective functions, separately reflecting occupied space, input/output transmission and torque performances, and multi-criteria constraints, such as dimension, interference and kinematics, are defined. In consideration of the passive exercise of plantar/dorsiflexion requiring large output moment, a torque index is proposed. To cope with the actuation redundancy of the parallel robot, a new output transmission index is defined as well. The multi-objective optimization problem is solved by using a modified Differential Evolution(DE) algorithm, which is characterized by new selection and mutation strategies. Meanwhile, a special penalty method is presented to tackle the multi-criteria constraints. Finally, numerical experiments for different optimization algorithms are implemented. The computation results show that the proposed indices of output transmission and torque, and constraint handling are effective for the redundant parallel robot; the modified DE algorithm is superior to the other tested algorithms, in terms of the ability of global search and the number of non-dominated solutions. The proposed methodology of multi-objective optimization can be also applied to the dimensional synthesis of other redundantly actuated parallel robots only with rotational movements.展开更多
The dynamic dexterity is an important issue for manipulator design, some indices were proposed for analyzing dynamic dexterity, but they can evaluate the dynamic performance just at one pose in the workspaee of the ma...The dynamic dexterity is an important issue for manipulator design, some indices were proposed for analyzing dynamic dexterity, but they can evaluate the dynamic performance just at one pose in the workspaee of the manipulator, and can't be applied to dynamic design expediently. Much work has been done in the kinematic optimization, but the work in the dynamic optimization is much less. A global dynamic condition number index is proposed and applied to the dynamic optimization design the parallel manipulator. This paper deals with the dynamic manipulability and dynamic optimization of a two degree-of-freedom (DOF) parallel manipulator. The particular velocity and particular angular velocity matrices of each moving part about the part's pivot point are derived fi'om the kinematic formulation of the manipulator, and the inertial force and inertial movement are obtained utilizing Newton-Euler formulation, then the inverse dynamic model of the parallel manipulator is proposed based on the virtual work principle. The general inertial ellipsoid and dynamic manipulability ellipsoid are applied to evaluate the dynamic performance of the manipulator, a global dynamic condition number index based on the condition number of general inertial matrix in the workspace is proposed, and then the link lengths of the manipulator is redesigned to optimize the dynamic manipulability by this index. The dynamic manipulability of the origin mechanism and the optimized mechanism are compared, the result shows that the optimized one is much better. The global dynamic condition number index has good effect in evaluating the dynamic dexterity of the whole workspace, and is efficient in the dynamic optimal design of the parallel manipulator.展开更多
In the development of modern DSP, more and more use of C/C++ as a development language has become a trend. Optimizationof C/C++ program has become an important link of the DSP software development. This article de...In the development of modern DSP, more and more use of C/C++ as a development language has become a trend. Optimizationof C/C++ program has become an important link of the DSP software development. This article describes the structure features ofTMS320C6678 processor, illustrates the principle of efficient optimization method for C/C++, and analyzes the results.展开更多
This paper describes parallel simulation of the memory/computation-intensive acoustic wave equation with CPU template buffer optimization. Considering the 8-core CPU shared storage platform as an example,we obtain a o...This paper describes parallel simulation of the memory/computation-intensive acoustic wave equation with CPU template buffer optimization. Considering the 8-core CPU shared storage platform as an example,we obtain a one-time speed-up ratio of 6.7× compared with the serial program by using a coarse-grained OpenMP parallel scheme. Then,data is vectorized on the template buffer using the single instruction-multiple data(SIMD) technique to further exploit the computing potential of the CPUs. We apply an 8-channel parallel vector to simulate seismic wavefields with the 256-bit advanced vector extensions(AVX) instruction set. This increases the computing bandwidth,thus eliminating a significant volume of the computing instructions and obtaining a secondary speed-up ratio of 3–7×. In addition,we use 32-byte data alignment,shortest data direction vectorization,and loop tiling optimization algorithm to achieve faster program execution. Finally,we analyze the factors affecting the secondary speed-up of AVX through three-dimensional modeling experiments with the salt model.The results indicate that the memory,cache,and register can better cooperate with each other and the speed-up is increased by optimizing the AVX algorithm.展开更多
The parallel mechanisms have the disadvantage of small workspace and complication in kinematics and dynamics. An optimizing design for the parallel mechanisms can improve the motion performance relatively, but not gua...The parallel mechanisms have the disadvantage of small workspace and complication in kinematics and dynamics. An optimizing design for the parallel mechanisms can improve the motion performance relatively, but not guarantee the design results which satisfy the various practical requirements simultaneously. In this paper, a dynamical and optimal synthesis method is proposed for parallel mechanisms based on the dynamical reconfiguration technique. As a specific, application, the problem of optimizing the kinematics isotropy of a five-bar planar parallel mechanism is studied. The motion of a reconfigurable mechanism can be parted into two phases, the natural motion phase and the reconfiguration phase. The two motion phases can be studied by the same performance evaluation methodology. This points out from both theory and practices a novel method for improving the motion performance of the parallel mechanisms. Simulation by a symmetrical five-bar planar parallel manipulator shows some aspects of the investigations.展开更多
Aiming at the development of parallel hybrid electric vehicle (PHEV) powertrain, parameter matching and optimization are presented, According to the performance of PHEV, the optimization range of engine, motor, driv...Aiming at the development of parallel hybrid electric vehicle (PHEV) powertrain, parameter matching and optimization are presented, According to the performance of PHEV, the optimization range of engine, motor, driveline gear ratio and battery parameters are determined. And then a two-level optimization problem is formulated based on analytical target cascading (ATC). At the system level, the optimization of the whole vehicle fuel economy is carried out, while the tractive performance is defined as the constraints. The optimized parameters are cascaded to the subsystem as the optimization targets. At the subsystem level, the final drive and transmission design are optimized to make the ratios as close to the targets as possible. The optimization result shows that the fuel economy had improved significantly, while the tractive performance maintains the former level.展开更多
Pointing mechanism is widely used in aerospace field,and its pointing accuracy and stability have high requirements.The pointing mechanism will be affected by external interference when it works.In order to eliminate ...Pointing mechanism is widely used in aerospace field,and its pointing accuracy and stability have high requirements.The pointing mechanism will be affected by external interference when it works.In order to eliminate the impact of interference forces on the output accuracy of the mechanism,firstly,this paper proposes a design method for highprecision pointing mechanisms based on interference separation,aiming at the high-precision pointing requirements of pointing mechanisms.Based on the screw theory,a synthesis method for inner compensation mechanisms has been proposed.And a new type of double-layer parallel mechanism has been designed to compensate for interference forces.Then,the kinematics and dynamics of the mechanism are carried out.An evaluation index for compensating external interference forces is proposed.The interference compensation analysis is conducted for the pointing mechanism.The correctness of the proposed interference force compensation coefficient is verified.Finally,in order to find the optimal solution for the workspace and interference force compensation coefficient of the pointing mechanism,multi-objective optimization design of the structural parameters of the mechanism was carried out based on the particle swarm optimization algorithm.This provides a theoretical basis for the prototype design of the subsequent double-layer parallel mechanism.This double-layer parallel mechanism combines the advantages of large load-bearing capacity,large workspace,and high output accuracy.It can be better applied in the aerospace field where high-precision pointing and force interference compensation are integrated.展开更多
This paper aims to solve large-scale and complex isogeometric topology optimization problems that consumesignificant computational resources. A novel isogeometric topology optimization method with a hybrid parallelstr...This paper aims to solve large-scale and complex isogeometric topology optimization problems that consumesignificant computational resources. A novel isogeometric topology optimization method with a hybrid parallelstrategy of CPU/GPU is proposed, while the hybrid parallel strategies for stiffness matrix assembly, equationsolving, sensitivity analysis, and design variable update are discussed in detail. To ensure the high efficiency ofCPU/GPU computing, a workload balancing strategy is presented for optimally distributing the workload betweenCPU and GPU. To illustrate the advantages of the proposedmethod, three benchmark examples are tested to verifythe hybrid parallel strategy in this paper. The results show that the efficiency of the hybrid method is faster thanserial CPU and parallel GPU, while the speedups can be up to two orders of magnitude.展开更多
Aiming at parallel distributed constant false alarm rate (CFAR) detection employing K/N fusion rule,an optimization algorithm based on the genetic algorithm with interval encoding is proposed. N-1 local probabilitie...Aiming at parallel distributed constant false alarm rate (CFAR) detection employing K/N fusion rule,an optimization algorithm based on the genetic algorithm with interval encoding is proposed. N-1 local probabilities of false alarm are selected as optimization variables. And the encoding intervals for local false alarm probabilities are sequentially designed by the person-by-person optimization technique according to the constraints. By turning constrained optimization to unconstrained optimization,the problem of increasing iteration times due to the punishment technique frequently adopted in the genetic algorithm is thus overcome. Then this optimization scheme is applied to spacebased synthetic aperture radar (SAR) multi-angle collaborative detection,in which the nominal factor for each local detector is determined. The scheme is verified with simulations of cases including two,three and four independent SAR systems. Besides,detection performances with varying K and N are compared and analyzed.展开更多
Currently,energy conservation draws wide attention in industrial manufacturing systems.In recent years,many studies have aimed at saving energy consumption in the process of manufacturing and scheduling is regarded as...Currently,energy conservation draws wide attention in industrial manufacturing systems.In recent years,many studies have aimed at saving energy consumption in the process of manufacturing and scheduling is regarded as an effective approach.This paper puts forwards a multi-objective stochastic parallel machine scheduling problem with the consideration of deteriorating and learning effects.In it,the real processing time of jobs is calculated by using their processing speed and normal processing time.To describe this problem in a mathematical way,amultiobjective stochastic programming model aiming at realizing makespan and energy consumption minimization is formulated.Furthermore,we develop a multi-objective multi-verse optimization combined with a stochastic simulation method to deal with it.In this approach,the multi-verse optimization is adopted to find favorable solutions from the huge solution domain,while the stochastic simulation method is employed to assess them.By conducting comparison experiments on test problems,it can be verified that the developed approach has better performance in coping with the considered problem,compared to two classic multi-objective evolutionary algorithms.展开更多
In this paper, a parallel Surface Extraction from Binary Volumes with Higher-Order Smoothness (SEBVHOS) algorithm is proposed to accelerate the SEBVHOS execution. The original SEBVHOS algorithm is parallelized first, ...In this paper, a parallel Surface Extraction from Binary Volumes with Higher-Order Smoothness (SEBVHOS) algorithm is proposed to accelerate the SEBVHOS execution. The original SEBVHOS algorithm is parallelized first, and then several performance optimization techniques which are loop optimization, cache optimization, false sharing optimization, synchronization overhead op-timization, and thread affinity optimization, are used to improve the implementation's performance on multi-core systems. The performance of the parallel SEBVHOS algorithm is analyzed on a dual-core system. The experimental results show that the parallel SEBVHOS algorithm achieves an average of 1.86x speedup. More importantly, our method does not come with additional aliasing artifacts, com-paring to the original SEBVHOS algorithm.展开更多
The Global-Regional Integrated forecast System(GRIST)is the next-generation weather and climate integrated model dynamic framework developed by Chinese Academy of Meteorological Sciences.In this paper,we present sever...The Global-Regional Integrated forecast System(GRIST)is the next-generation weather and climate integrated model dynamic framework developed by Chinese Academy of Meteorological Sciences.In this paper,we present several changes made to the global nonhydrostatic dynamical(GND)core,which is part of the ongoing prototype of GRIST.The changes leveraging MPI and PnetCDF techniques were targeted at the parallelization and performance optimization to the original serial GND core.Meanwhile,some sophisticated data structures and interfaces were designed to adjust flexibly the size of boundary and halo domains according to the variable accuracy in parallel context.In addition,the I/O performance of PnetCDF decreases as the number of MPI processes increases in our experimental environment.Especially when the number exceeds 6000,it caused system-wide outages(SWO).Thus,a grouping solution was proposed to overcome that issue.Several experiments were carried out on the supercomputing platform based on Intel x86 CPUs in the National Supercomputing Center in Wuxi.The results demonstrated that the parallel GND core based on grouping solution achieves good strong scalability and improves the performance significantly,as well as avoiding the SWOs.展开更多
摘要In order to achieve comprehensive,highly efficient,and multi-objective precise optimization of fiber structural parameters and further enhance the transmission capacity of optical communication systems,a homogeneous weakly coupled seven-core fiber based on trench-assisted structures is designed.Particle Swarm Optimization(PSO)is introduced to replace traditional empirical designs or local scanning methods.First,a multi-objective fitness function incorporating constraints such as dispersion,cutoff wavelength,ef-fective mode field area,and coating loss is established.Then,the algorithm performs a global search to pre-cisely determine the optimal structural parameters under standard dimensional constraints.Simulation results demonstrate that with a fiber core pitch of 45μm,the optimized fiber achieves an ultra-low inter-core crosstalk of below−90 dB/km at a wavelength of 1550 nm.This design scheme not only effectively resolves the conflict between crosstalk suppression and spatial utilization in multi-core fibers but also proves the effi-ciency and reliability of the PSO algorithm in complex fiber structural design,providing an important theor-etical basis and technical support for the research and manufacturing of ultra-large-capacity optical commu-nication systems.
基金supported by the National Key Research and Development Program of China(Grant No.2023YFB3309104)the National Natural Science Foundation of China(Grant Nos.11821202 and 123721222)+1 种基金the Science Technology Plan of Liaoning Province(Grant No.2023JH2/101600044)the 111 Project of China(Grant No.B14013).
摘要Although supplying extensive design space,the curse of dimensionality restricts the widespread application of largescale topology optimization in practical engineering.Various acceleration techniques have been integrated with topology optimization,achieving significant attention and progress in large-scale problems.This work aims to investigate how much benefit can be obtained by combining parallel computing and machine learning techniques to enhance the efficiency of large-scale topology optimization algorithms.Accordingly,a parallel problem independent machine learning(PIML)-enhanced topology optimization method is proposed.The PIML model substantially reduces the dimension of the condensed stiffness matrix and its computational cost,and parallel computing reduces the workload per process and enables the application of a parallel multigrid solver.Besides,several techniques,such as matrix-free implementation,direct condensation of uniform coarse elements,and adjusting computational resource limits,have been developed to enhance computational efficiency.The weak scaling efficiency,strong scaling speedup,and maximum achievable efficiency of the proposed method are validated across multiple numerical examples,showing significant improvement in the tractable problem size and solution efficiency compared to traditional topology optimization algorithms.
基金funded by the Malaysian Ministry of Higher Education through the Fundamental Research Grant Scheme(FRGS/1/2024/ICT02/UCSI/02/1).
摘要Accurate estimation of photovoltaic(PV)parameters is essential for optimizing solar module perfor-mance and enhancing resource efficiency in renewable energy systems.This study presents a process innovation by introducing,for the first time,the Triangulation Topology Aggregation Optimizer(TTAO)integrated with parallel computing to address PV parameter estimation challenges.The effectiveness and robustness of TTAO are rigorously evaluated using two standard benchmark datasets(KC200GT and R.T.C.France solar cells)and a real-world dataset(Poly70W solar module)under single-,double-,and triple-diode configurations.Results show that TTAO consistently achieves superior accuracy by producing the lowest RMSE values and faster convergence compared to state-of-the-art metaheuristic algorithms.In addition,the integration of parallel computing significantly enhances computational efficiency,reducing execution time by up to 85%without compromising accuracy.Validation using real-world data further demonstrates TTAO’s adaptability and practical relevance in renewable energy systems,effectively bridging the gap between theoretical modeling and real-world implementation for PV system monitoring and optimization,contributing to climate mitigation through improved solar energy performance.
摘要Magnetic Resonance Imaging(MRI)has a pivotal role in medical image analysis,for its ability in supporting disease detection and diagnosis.Fuzzy C-Means(FCM)clustering is widely used for MRI segmentation due to its ability to handle image uncertainty.However,the latter still has countless limitations,including sensitivity to initialization,susceptibility to local optima,and high computational cost.To address these limitations,this study integrates Grey Wolf Optimization(GWO)with FCM to enhance cluster center selection,improving segmentation accuracy and robustness.Moreover,to further refine optimization,Fuzzy Entropy Clustering was utilized for its distinctive features from other traditional objective functions.Fuzzy entropy effectively quantifies uncertainty,leading to more well-defined clusters,improved noise robustness,and better preservation of anatomical structures in MRI images.Despite these advantages,the iterative nature of GWO and FCM introduces significant computational overhead,which restricts their applicability to high-resolution medical images.To overcome this bottleneck,we propose a Parallelized-GWO-based FCM(P-GWO-FCM)approach using GPU acceleration,where both GWO optimization and FCM updates(centroid computation and membership matrix updates)are parallelized.By concurrently executing these processes,our approach efficiently distributes the computational workload,significantly reducing execution time while maintaining high segmentation accuracy.The proposed parallel method,P-GWO-FCM,was evaluated on both simulated and clinical brain MR images,focusing on segmenting white matter,gray matter,and cerebrospinal fluid regions.The results indicate significant improvements in segmentation accuracy,achieving a Jaccard Similarity(JS)of 0.92,a Partition Coefficient Index(PCI)of 0.91,a Partition Entropy Index(PEI)of 0.25,and a Davies-Bouldin Index(DBI)of 0.30.Experimental comparisons demonstrate that P-GWO-FCM outperforms existing methods in both segmentation accuracy and computational efficiency,making it a promising solution for real-time medical image segmentation.
基金supported by National Key Research and Development Program of China under Grant 2024YFE0210800National Natural Science Foundation of China under Grant 62495062Beijing Natural Science Foundation under Grant L242017.
摘要The Dynamical Density Functional Theory(DDFT)algorithm,derived by associating classical Density Functional Theory(DFT)with the fundamental Smoluchowski dynamical equation,describes the evolution of inhomo-geneous fluid density distributions over time.It plays a significant role in studying the evolution of density distributions over time in inhomogeneous systems.The Sunway Bluelight II supercomputer,as a new generation of China’s developed supercomputer,possesses powerful computational capabilities.Porting and optimizing industrial software on this platform holds significant importance.For the optimization of the DDFT algorithm,based on the Sunway Bluelight II supercomputer and the unique hardware architecture of the SW39000 processor,this work proposes three acceleration strategies to enhance computational efficiency and performance,including direct parallel optimization,local-memory constrained optimization for CPEs,and multi-core groups collaboration and communication optimization.This method combines the characteristics of the program’s algorithm with the unique hardware architecture of the Sunway Bluelight II supercomputer,optimizing the storage and transmission structures to achieve a closer integration of software and hardware.For the first time,this paper presents Sunway-Dynamical Density Functional Theory(SW-DDFT).Experimental results show that SW-DDFT achieves a speedup of 6.67 times within a single-core group compared to the original DDFT implementation,with six core groups(a total of 384 CPEs),the maximum speedup can reach 28.64 times,and parallel efficiency can reach 71%,demonstrating excellent acceleration performance.
基金supported in part by the National Science and Technology Major Project(2022ZD0119902)the National Natural Science Foundation of China(52471372,623B2018,62203015,62233001)+4 种基金the Liaoning Revitalization Leading Talents Program(XLYC2402054)the Key Basic Research of Dalian(2023JJ11CG008)the Fundamental Research Funds for the Central Universities(3132023508)the Collaborative Research Fund of Hong Kong Research Grants Council(C1013-24G)the Cultivation Program for the Excellent Doctoral Dissertation of Dalian Maritime University(2023YBPY005).
摘要This paper addresses the parallel control of autonomous surface vehicles subject to external disturbances,state constraints,and input constraints in complex ocean environments with multiple obstacles.A safety-certified parallel model predictive control scheme with collision-avoiding capability is proposed for autonomous surface vehicles in the framework of parallel control.Specifically,an extended state observer is designed by leveraging historical and real-time data for concurrent learning to map the motion of autonomous surface vehicles from its physical system to its artificial counterpart.A parallel model predictive control law is developed on the basis of the artificial system for both physical and artificial autonomous surface vehicles to realize virtual-physical tracking control of vehicles subject to state and input constraints.To ensure safety,highorder discrete control barrier functions are encoded in the parallel model predictive control law as safety constraints such that collision avoidance with obstacles can be achieved.A recedinghorizon constrained optimization problem is constructed with the safety constraints encoded by control barrier functions for parallel model predictive control of autonomous surface vehicles and solved via neurodynamic optimization with projection neural networks.The effectiveness and characteristics of the proposed method are demonstrated via simulations for the safe trajectory tracking and automatic berthing of autonomous surface vehicles.
基金supported in part by the National Key Research and Development Program of China(2024YFF0908200)the National Natural Science Foundation of China(62302402,62272078)+1 种基金the Chongqing Natural Science Foundation(CSTB2024TIAD-KPX0018,CSTB2023NSCO-LZX006)the Southwest University Graduate Research Innovation Project(SWUB24050)
摘要High-dimensional and incomplete(HDI) matrices are commonly encountered in various big data-related applications for illustrating the complex interactions among numerous entities, like the user-item interactions in a commercial recommender system or the user-user interactions in a social network services system. The factorization of such an HDI matrix can embed the involved entities into the low-dimensional feature space for acquiring their principal representation, which is a vital task in various application scenes and is often established through the Latent Factor Analysis(LFA). Nevertheless, an HDI matrix can be huge when the corresponding application explodes to involve millions of users, items, or other interactive nodes. In this case, a parallel optimization algorithm is desired for raising the scalability and time efficiency of an LFA model. This paper provides a comprehensive review of the existing parallel optimization algorithms for the LFA model. Specifically, it performs: 1) discussion and summary of these algorithms based on computing architecture and mode, 2) empirical studies of representative models, and3) summary of the current challenges and future directions in this domain. This survey aims to offer an exhaustive review of Parallel Optimization Algorithms for High-Dimensional and Incomplete Matrix Factorization, thereby fostering further research in this field.
基金Project(2008AA01A201) supported the National High-tech Research and Development Program of ChinaProjects(60833004, 60633050) supported by the National Natural Science Foundation of China
摘要Developing parallel applications on heterogeneous processors is facing the challenges of 'memory wall',due to limited capacity of local storage,limited bandwidth and long latency for memory access. Aiming at this problem,a parallelization approach was proposed with six memory optimization schemes for CG,four schemes of them aiming at all kinds of sparse matrix-vector multiplication (SPMV) operation. Conducted on IBM QS20,the parallelization approach can reach up to 21 and 133 times speedups with size A and B,respectively,compared with single power processor element. Finally,the conclusion is drawn that the peak bandwidth of memory access on Cell BE can be obtained in SPMV,simple computation is more efficient on heterogeneous processors and loop-unrolling can hide local storage access latency while executing scalar operation on SIMD cores.
基金Supported by National Natural Science Foundation of China(Grant No.51175029)Beijing Municipal Natural Science Foundation of China(Grant No.3132019)
摘要Dimensional synthesis is one of the most difficult issues in the field of parallel robots with actuation redundancy. To deal with the optimal design of a redundantly actuated parallel robot used for ankle rehabilitation, a methodology of dimensional synthesis based on multi-objective optimization is presented. First, the dimensional synthesis of the redundant parallel robot is formulated as a nonlinear constrained multi-objective optimization problem. Then four objective functions, separately reflecting occupied space, input/output transmission and torque performances, and multi-criteria constraints, such as dimension, interference and kinematics, are defined. In consideration of the passive exercise of plantar/dorsiflexion requiring large output moment, a torque index is proposed. To cope with the actuation redundancy of the parallel robot, a new output transmission index is defined as well. The multi-objective optimization problem is solved by using a modified Differential Evolution(DE) algorithm, which is characterized by new selection and mutation strategies. Meanwhile, a special penalty method is presented to tackle the multi-criteria constraints. Finally, numerical experiments for different optimization algorithms are implemented. The computation results show that the proposed indices of output transmission and torque, and constraint handling are effective for the redundant parallel robot; the modified DE algorithm is superior to the other tested algorithms, in terms of the ability of global search and the number of non-dominated solutions. The proposed methodology of multi-objective optimization can be also applied to the dimensional synthesis of other redundantly actuated parallel robots only with rotational movements.
基金supported by National Natural Science Foundation of China (Grant No. 50605041, No. 50775125)National Basic Research Program of China (973 Program, Grant No. 2006CB705400)
摘要The dynamic dexterity is an important issue for manipulator design, some indices were proposed for analyzing dynamic dexterity, but they can evaluate the dynamic performance just at one pose in the workspaee of the manipulator, and can't be applied to dynamic design expediently. Much work has been done in the kinematic optimization, but the work in the dynamic optimization is much less. A global dynamic condition number index is proposed and applied to the dynamic optimization design the parallel manipulator. This paper deals with the dynamic manipulability and dynamic optimization of a two degree-of-freedom (DOF) parallel manipulator. The particular velocity and particular angular velocity matrices of each moving part about the part's pivot point are derived fi'om the kinematic formulation of the manipulator, and the inertial force and inertial movement are obtained utilizing Newton-Euler formulation, then the inverse dynamic model of the parallel manipulator is proposed based on the virtual work principle. The general inertial ellipsoid and dynamic manipulability ellipsoid are applied to evaluate the dynamic performance of the manipulator, a global dynamic condition number index based on the condition number of general inertial matrix in the workspace is proposed, and then the link lengths of the manipulator is redesigned to optimize the dynamic manipulability by this index. The dynamic manipulability of the origin mechanism and the optimized mechanism are compared, the result shows that the optimized one is much better. The global dynamic condition number index has good effect in evaluating the dynamic dexterity of the whole workspace, and is efficient in the dynamic optimal design of the parallel manipulator.
摘要In the development of modern DSP, more and more use of C/C++ as a development language has become a trend. Optimizationof C/C++ program has become an important link of the DSP software development. This article describes the structure features ofTMS320C6678 processor, illustrates the principle of efficient optimization method for C/C++, and analyzes the results.
基金funded by the National Natural Science Foundation of China (No. 41274140)the National Science and Technology Major Projects of China (No. 2017ZX05035003-001)
摘要This paper describes parallel simulation of the memory/computation-intensive acoustic wave equation with CPU template buffer optimization. Considering the 8-core CPU shared storage platform as an example,we obtain a one-time speed-up ratio of 6.7× compared with the serial program by using a coarse-grained OpenMP parallel scheme. Then,data is vectorized on the template buffer using the single instruction-multiple data(SIMD) technique to further exploit the computing potential of the CPUs. We apply an 8-channel parallel vector to simulate seismic wavefields with the 256-bit advanced vector extensions(AVX) instruction set. This increases the computing bandwidth,thus eliminating a significant volume of the computing instructions and obtaining a secondary speed-up ratio of 3–7×. In addition,we use 32-byte data alignment,shortest data direction vectorization,and loop tiling optimization algorithm to achieve faster program execution. Finally,we analyze the factors affecting the secondary speed-up of AVX through three-dimensional modeling experiments with the salt model.The results indicate that the memory,cache,and register can better cooperate with each other and the speed-up is increased by optimizing the AVX algorithm.
摘要The parallel mechanisms have the disadvantage of small workspace and complication in kinematics and dynamics. An optimizing design for the parallel mechanisms can improve the motion performance relatively, but not guarantee the design results which satisfy the various practical requirements simultaneously. In this paper, a dynamical and optimal synthesis method is proposed for parallel mechanisms based on the dynamical reconfiguration technique. As a specific, application, the problem of optimizing the kinematics isotropy of a five-bar planar parallel mechanism is studied. The motion of a reconfigurable mechanism can be parted into two phases, the natural motion phase and the reconfiguration phase. The two motion phases can be studied by the same performance evaluation methodology. This points out from both theory and practices a novel method for improving the motion performance of the parallel mechanisms. Simulation by a symmetrical five-bar planar parallel manipulator shows some aspects of the investigations.
摘要Aiming at the development of parallel hybrid electric vehicle (PHEV) powertrain, parameter matching and optimization are presented, According to the performance of PHEV, the optimization range of engine, motor, driveline gear ratio and battery parameters are determined. And then a two-level optimization problem is formulated based on analytical target cascading (ATC). At the system level, the optimization of the whole vehicle fuel economy is carried out, while the tractive performance is defined as the constraints. The optimized parameters are cascaded to the subsystem as the optimization targets. At the subsystem level, the final drive and transmission design are optimized to make the ratios as close to the targets as possible. The optimization result shows that the fuel economy had improved significantly, while the tractive performance maintains the former level.
基金Supported by the National Natural Science Foundation of China(52275032)the Natural Science Foundation of Hebei Province(E2022203077)+1 种基金the Hebei Provincial Science and Technology Plan(22371801D)the Hebei Provincial Science and Technology Research and Development Program-Central Guidance for Local Science and Technology Development Fund(246Z1818G).
摘要Pointing mechanism is widely used in aerospace field,and its pointing accuracy and stability have high requirements.The pointing mechanism will be affected by external interference when it works.In order to eliminate the impact of interference forces on the output accuracy of the mechanism,firstly,this paper proposes a design method for highprecision pointing mechanisms based on interference separation,aiming at the high-precision pointing requirements of pointing mechanisms.Based on the screw theory,a synthesis method for inner compensation mechanisms has been proposed.And a new type of double-layer parallel mechanism has been designed to compensate for interference forces.Then,the kinematics and dynamics of the mechanism are carried out.An evaluation index for compensating external interference forces is proposed.The interference compensation analysis is conducted for the pointing mechanism.The correctness of the proposed interference force compensation coefficient is verified.Finally,in order to find the optimal solution for the workspace and interference force compensation coefficient of the pointing mechanism,multi-objective optimization design of the structural parameters of the mechanism was carried out based on the particle swarm optimization algorithm.This provides a theoretical basis for the prototype design of the subsequent double-layer parallel mechanism.This double-layer parallel mechanism combines the advantages of large load-bearing capacity,large workspace,and high output accuracy.It can be better applied in the aerospace field where high-precision pointing and force interference compensation are integrated.
基金the National Key R&D Program of China(2020YFB1708300)the National Natural Science Foundation of China(52005192)the Project of Ministry of Industry and Information Technology(TC210804R-3).
摘要This paper aims to solve large-scale and complex isogeometric topology optimization problems that consumesignificant computational resources. A novel isogeometric topology optimization method with a hybrid parallelstrategy of CPU/GPU is proposed, while the hybrid parallel strategies for stiffness matrix assembly, equationsolving, sensitivity analysis, and design variable update are discussed in detail. To ensure the high efficiency ofCPU/GPU computing, a workload balancing strategy is presented for optimally distributing the workload betweenCPU and GPU. To illustrate the advantages of the proposedmethod, three benchmark examples are tested to verifythe hybrid parallel strategy in this paper. The results show that the efficiency of the hybrid method is faster thanserial CPU and parallel GPU, while the speedups can be up to two orders of magnitude.
基金New Century Program for Excellent Talents of Minis-try of Education of China (NECT-06-0166)The Eleventh Five-year Scientific and Technological Development Plan of National Defense Pre-study Foundation (A2120060006)
摘要Aiming at parallel distributed constant false alarm rate (CFAR) detection employing K/N fusion rule,an optimization algorithm based on the genetic algorithm with interval encoding is proposed. N-1 local probabilities of false alarm are selected as optimization variables. And the encoding intervals for local false alarm probabilities are sequentially designed by the person-by-person optimization technique according to the constraints. By turning constrained optimization to unconstrained optimization,the problem of increasing iteration times due to the punishment technique frequently adopted in the genetic algorithm is thus overcome. Then this optimization scheme is applied to spacebased synthetic aperture radar (SAR) multi-angle collaborative detection,in which the nominal factor for each local detector is determined. The scheme is verified with simulations of cases including two,three and four independent SAR systems. Besides,detection performances with varying K and N are compared and analyzed.
摘要Currently,energy conservation draws wide attention in industrial manufacturing systems.In recent years,many studies have aimed at saving energy consumption in the process of manufacturing and scheduling is regarded as an effective approach.This paper puts forwards a multi-objective stochastic parallel machine scheduling problem with the consideration of deteriorating and learning effects.In it,the real processing time of jobs is calculated by using their processing speed and normal processing time.To describe this problem in a mathematical way,amultiobjective stochastic programming model aiming at realizing makespan and energy consumption minimization is formulated.Furthermore,we develop a multi-objective multi-verse optimization combined with a stochastic simulation method to deal with it.In this approach,the multi-verse optimization is adopted to find favorable solutions from the huge solution domain,while the stochastic simulation method is employed to assess them.By conducting comparison experiments on test problems,it can be verified that the developed approach has better performance in coping with the considered problem,compared to two classic multi-objective evolutionary algorithms.
基金Supported by the National Natural Science Foundation of China(No.61071173)
摘要In this paper, a parallel Surface Extraction from Binary Volumes with Higher-Order Smoothness (SEBVHOS) algorithm is proposed to accelerate the SEBVHOS execution. The original SEBVHOS algorithm is parallelized first, and then several performance optimization techniques which are loop optimization, cache optimization, false sharing optimization, synchronization overhead op-timization, and thread affinity optimization, are used to improve the implementation's performance on multi-core systems. The performance of the parallel SEBVHOS algorithm is analyzed on a dual-core system. The experimental results show that the parallel SEBVHOS algorithm achieves an average of 1.86x speedup. More importantly, our method does not come with additional aliasing artifacts, com-paring to the original SEBVHOS algorithm.
基金This work was supported by the National Key Research and Development Program of China under Grant No.2017YFC1502203.
摘要The Global-Regional Integrated forecast System(GRIST)is the next-generation weather and climate integrated model dynamic framework developed by Chinese Academy of Meteorological Sciences.In this paper,we present several changes made to the global nonhydrostatic dynamical(GND)core,which is part of the ongoing prototype of GRIST.The changes leveraging MPI and PnetCDF techniques were targeted at the parallelization and performance optimization to the original serial GND core.Meanwhile,some sophisticated data structures and interfaces were designed to adjust flexibly the size of boundary and halo domains according to the variable accuracy in parallel context.In addition,the I/O performance of PnetCDF decreases as the number of MPI processes increases in our experimental environment.Especially when the number exceeds 6000,it caused system-wide outages(SWO).Thus,a grouping solution was proposed to overcome that issue.Several experiments were carried out on the supercomputing platform based on Intel x86 CPUs in the National Supercomputing Center in Wuxi.The results demonstrated that the parallel GND core based on grouping solution achieves good strong scalability and improves the performance significantly,as well as avoiding the SWOs.