The complexity of an elastic wavefield increases the nonlinearity of inversion, To some extent, multiscale inversion decreases the nonlinearity of inversion and prevents it from falling into local extremes. A multisca...The complexity of an elastic wavefield increases the nonlinearity of inversion, To some extent, multiscale inversion decreases the nonlinearity of inversion and prevents it from falling into local extremes. A multiscale strategy based on the simultaneous use of frequency groups and layer stripping method based on damped wave field improves the stability of inversion. A dual-level parallel algorithm is then used to decrease the computational cost and improve practicability. The seismic wave modeling of a single frequency and inversion in a frequency group are computed in parallel by multiple nodes based on multifrontal massively parallel sparse direct solver and MPI. Numerical tests using an overthrust model show that the proposed inversion algorithm can effectively improve the stability and accuracy of inversion by selecting the appropriate inversion frequency and damping factor in low- frequency seismic data.展开更多
Multiple sequence alignment (MSA) is the alignment among more than two molecular biological sequences, which is a fundamental method to analyze evolutionary events such as mutations, insertions, deletions, and re-ar...Multiple sequence alignment (MSA) is the alignment among more than two molecular biological sequences, which is a fundamental method to analyze evolutionary events such as mutations, insertions, deletions, and re-arrangements. In theory, a dynamic programming algorithm can be employed to produce the optimal MSA. However, this leads to an explosive increase in computing time and memory consumption as the number of sequences increases (Taylor, 1990). So far, MSA is still regarded as one of the most challenging problems in bioinformatics and computational biology (Chatzou et al., 2016).展开更多
According to the characteristics of Chinese marginal seas, the Marginal Sea Model of China(MSMC) has been developed independently in China. Because the model requires long simulation time, as a routine forecasting mod...According to the characteristics of Chinese marginal seas, the Marginal Sea Model of China(MSMC) has been developed independently in China. Because the model requires long simulation time, as a routine forecasting model, the parallelism of MSMC becomes necessary to be introduced to improve the performance of it. However, some methods used in MSMC, such as Successive Over Relaxation(SOR) algorithm, are not suitable for parallelism. In this paper, methods are developedto solve the parallel problem of the SOR algorithm following the steps as below. First, based on a 3D computing grid system, an automatic data partition method is implemented to dynamically divide the computing grid according to computing resources. Next, based on the characteristics of the numerical forecasting model, a parallel method is designed to solve the parallel problem of the SOR algorithm. Lastly, a communication optimization method is provided to avoid the cost of communication. In the communication optimization method, the non-blocking communication of Message Passing Interface(MPI) is used to implement the parallelism of MSMC with complex physical equations, and the process of communication is overlapped with the computations for improving the performance of parallel MSMC. The experiments show that the parallel MSMC runs 97.2 times faster than the serial MSMC, and root mean square error between the parallel MSMC and the serial MSMC is less than 0.01 for a 30-day simulation(172800 time steps), which meets the requirements of timeliness and accuracy for numerical ocean forecasting products.展开更多
Training deep neural networks(DNNs)requires a significant amount of time and resources to obtain acceptable results,which severely limits its deployment in resource-limited platforms.This paper proposes DarkFPGA,a nov...Training deep neural networks(DNNs)requires a significant amount of time and resources to obtain acceptable results,which severely limits its deployment in resource-limited platforms.This paper proposes DarkFPGA,a novel customizable framework to efficiently accelerate the entire DNN training on a single FPGA platform.First,we explore batch-level parallelism to enable efficient FPGA-based DNN training.Second,we devise a novel hardware architecture optimised by a batch-oriented data pattern and tiling techniques to effectively exploit parallelism.Moreover,an analytical model is developed to determine the optimal design parameters for the DarkFPGA accelerator with respect to a specific network specification and FPGA resource constraints.Our results show that the accelerator is able to perform about 10 times faster than CPU training and about a third of the energy consumption than GPU training using 8-bit integers for training VGG-like networks on the CIFAR dataset for the Maxeler MAX5 platform.展开更多
During the fabrication of quartz crystal resonators(QCRs),parallelism error is inevitably generated,which is rarely investigated.In order to reveal the influence of parallelism error on the working performance of QCRs...During the fabrication of quartz crystal resonators(QCRs),parallelism error is inevitably generated,which is rarely investigated.In order to reveal the influence of parallelism error on the working performance of QCRs,the coupled vibration of a non-parallel AT-cut quartz crystal plate with electrodes is systematically studied from the views of theoretical analysis and numerical simulations.The two-dimensional thermal incremental field equations are solved for the free vibration analysis via the coefficient-formed partial differential equation module of the COMSOL Multiphysics software,from which the frequency spectra,frequency–temperature curves,and mode shapes are discussed in detail.Additionally,the piezoelectric module is utilized to obtain the admittance response under different conditions.It is demonstrated that the parallelism error reduces the resonant frequency.Additionally,symmetry broken by the non-parallelism increases the probability of activity dip and is harmful to QCR’s thermal stability.However,if the top and bottom surfaces incline synchronously in the same direction,the influence of parallelism error is tiny.The conclusions achieved are helpful for the QCR design,and the methodology presented can also be applied to other wave devices.展开更多
Two groups of parallelism in Martin Luther King's I Have a Dream are selected to analyze the rhythmic features in terms of prominence and duration. By applying Praat to label the sentences, some findings are as fo...Two groups of parallelism in Martin Luther King's I Have a Dream are selected to analyze the rhythmic features in terms of prominence and duration. By applying Praat to label the sentences, some findings are as follows: In parallel sentences,the duration of each parallel structure is more or less equal due to the same number of prominence, although sentence length is different; the mean duration of syllables in long sentences is shorter as a result of vowel truncation.展开更多
An improvement detecting method was proposed according to the disadvantages of testing method of optical axes parallelism of shipboard photoelectrical theodolite (short for theodolite) based on image processing. Point...An improvement detecting method was proposed according to the disadvantages of testing method of optical axes parallelism of shipboard photoelectrical theodolite (short for theodolite) based on image processing. Pointolite replaced 0.2'' collimator to reduce the errors of crosshair images processing and improve the quality of image. What’s more, the high quality images could help to optimize the image processing method and the testing accuracy. The errors between the trial results interpreted by software and the results tested in dock were less than 10'', which indicated the improve method had some actual application values.展开更多
Parallelism在英语中具有修辞和语法双重功能。在英汉翻译对比研究中,学术界关注较多的是它与中文排比的构成差异及修辞功用,对它的语法功能研究较少。实践证明,英汉排比结构的对等转译不仅有助于保留源文本的语言特色和功能,还可以帮...Parallelism在英语中具有修辞和语法双重功能。在英汉翻译对比研究中,学术界关注较多的是它与中文排比的构成差异及修辞功用,对它的语法功能研究较少。实践证明,英汉排比结构的对等转译不仅有助于保留源文本的语言特色和功能,还可以帮助译文读者更好地把握原文结构,明了行文思路。Parallelism的修辞功能和语法功能同等重要,都应引起译界的足够重视。本文在功能对等理论的指导下,以美国作家M·斯科特·派克的心理学著作——The Road Less Traveled中的排比句的翻译为例,探讨如何通过直译法、转译法、分句法、逆序法、增词法和减词法来提高原文读者和译文读者反应的相似性,实现功能对等目的。展开更多
The simulation is an important means of performance evaluation of the computer architecture. Nowadays, the serial simulation of general purpose graphics processing unit(GPGPU) architecture is the main bottleneck for t...The simulation is an important means of performance evaluation of the computer architecture. Nowadays, the serial simulation of general purpose graphics processing unit(GPGPU) architecture is the main bottleneck for the simulation speed. To address this issue, we propose the intra-kernel parallelization on a multicore processor and the inter-kernel parallelization on a multiple-machine platform. We apply these two methods to the GPGPU-sim simulator. The intra-kernel parallelization method firstly parallelizes the serial simulation of multiple compute units in one cycle. Then it parallelizes the timing and functional simulation to reduce the performance loss caused by the synchronization between different compute units. The inter-kernel parallelization method divides multiple kernels of a CUDA program into several groups and distributes these groups across multiple simulation hosts to perform the simulation. Experimental results show that the intra-kernel parallelization method achieves a speed-up of up to 12 with a maximum error rate of 0.009 4% on a 32-core machine, and the inter-kernel parallelization method can accelerate the simulation by a factor of up to 3.9 with a maximum error rate of 0.11% on four simulation hosts. The orthogonality between these two methods allows us to combine them together on multiple multi-core hosts to get further performance improvements.展开更多
Define and theory of autocorrelation decision tree (ADT) is introduced. In spatial data mining, spatial parallel query are very expensive operations. A new parallel algorithm in terms of autocorrelation decision tre...Define and theory of autocorrelation decision tree (ADT) is introduced. In spatial data mining, spatial parallel query are very expensive operations. A new parallel algorithm in terms of autocorrelation decision tree is presented. And the new method reduces CPU- and I/O-time and improves the query efficiency of spatial data. For dynamic load balancing, there are better control and optimization. Experimental performance comparison shows that the improved algorithm can obtain a optimal accelerator with the same quantities of processors. There are more completely accesses on nodes. And an individual implement of intelligent information retrieval for spatial data mining is presented.展开更多
基金supported by the Natural Science Foundation of China(No.41374122)
摘要The complexity of an elastic wavefield increases the nonlinearity of inversion, To some extent, multiscale inversion decreases the nonlinearity of inversion and prevents it from falling into local extremes. A multiscale strategy based on the simultaneous use of frequency groups and layer stripping method based on damped wave field improves the stability of inversion. A dual-level parallel algorithm is then used to decrease the computational cost and improve practicability. The seismic wave modeling of a single frequency and inversion in a frequency group are computed in parallel by multiple nodes based on multifrontal massively parallel sparse direct solver and MPI. Numerical tests using an overthrust model show that the proposed inversion algorithm can effectively improve the stability and accuracy of inversion by selecting the appropriate inversion frequency and damping factor in low- frequency seismic data.
基金supported by the National Key R&D Program of China (Nos. 2017YFB0202600, 2016YFC1302500, 2016YFB0200400 and 2017YFB0202104)the National Natural Science Foundation of China (Nos. 61772543, U1435222, 61625202, 61272056 and 61771331)Guangdong Provincial Department of Science and Technology (No. 2016B090918122)
摘要Multiple sequence alignment (MSA) is the alignment among more than two molecular biological sequences, which is a fundamental method to analyze evolutionary events such as mutations, insertions, deletions, and re-arrangements. In theory, a dynamic programming algorithm can be employed to produce the optimal MSA. However, this leads to an explosive increase in computing time and memory consumption as the number of sequences increases (Taylor, 1990). So far, MSA is still regarded as one of the most challenging problems in bioinformatics and computational biology (Chatzou et al., 2016).
基金supported by the research of the key technology and exemplary applications about safety service system for marine fisheries under contract No. 201205006the foundation of Chinese Scholarship Council
摘要According to the characteristics of Chinese marginal seas, the Marginal Sea Model of China(MSMC) has been developed independently in China. Because the model requires long simulation time, as a routine forecasting model, the parallelism of MSMC becomes necessary to be introduced to improve the performance of it. However, some methods used in MSMC, such as Successive Over Relaxation(SOR) algorithm, are not suitable for parallelism. In this paper, methods are developedto solve the parallel problem of the SOR algorithm following the steps as below. First, based on a 3D computing grid system, an automatic data partition method is implemented to dynamically divide the computing grid according to computing resources. Next, based on the characteristics of the numerical forecasting model, a parallel method is designed to solve the parallel problem of the SOR algorithm. Lastly, a communication optimization method is provided to avoid the cost of communication. In the communication optimization method, the non-blocking communication of Message Passing Interface(MPI) is used to implement the parallelism of MSMC with complex physical equations, and the process of communication is overlapped with the computations for improving the performance of parallel MSMC. The experiments show that the parallel MSMC runs 97.2 times faster than the serial MSMC, and root mean square error between the parallel MSMC and the serial MSMC is less than 0.01 for a 30-day simulation(172800 time steps), which meets the requirements of timeliness and accuracy for numerical ocean forecasting products.
摘要Training deep neural networks(DNNs)requires a significant amount of time and resources to obtain acceptable results,which severely limits its deployment in resource-limited platforms.This paper proposes DarkFPGA,a novel customizable framework to efficiently accelerate the entire DNN training on a single FPGA platform.First,we explore batch-level parallelism to enable efficient FPGA-based DNN training.Second,we devise a novel hardware architecture optimised by a batch-oriented data pattern and tiling techniques to effectively exploit parallelism.Moreover,an analytical model is developed to determine the optimal design parameters for the DarkFPGA accelerator with respect to a specific network specification and FPGA resource constraints.Our results show that the accelerator is able to perform about 10 times faster than CPU training and about a third of the energy consumption than GPU training using 8-bit integers for training VGG-like networks on the CIFAR dataset for the Maxeler MAX5 platform.
基金supported by the National Natural Science Foundation of China(12061131013,11972276,12172171 and 12102183)the Fundamental Research Funds for the Central Universities(NE2020002 andNS2022011)+5 种基金JiangsuHigh-Level Innovative and Entrepreneurial Talents Introduction Plan(Shuangchuang Doctor Program,JSSCBS20210166)the National Natural Science Foundation of Jiangsu Province(BK20211176)the State Key Laboratory of Mechanics and Control of Mechanical Structures at NUAA(No.MCMS-I-0522G01)Local Science andTechnologyDevelopment Fund ProjectsGuided by the CentralGovernment(2021Szvup061)the Opening Projects from the Key Laboratory of Impact and Safety Engineering of Ningbo University(CJ202104)a project Funded by the Priority Academic Program Development of Jiangsu Higher Education Institutions(PAPD).
摘要During the fabrication of quartz crystal resonators(QCRs),parallelism error is inevitably generated,which is rarely investigated.In order to reveal the influence of parallelism error on the working performance of QCRs,the coupled vibration of a non-parallel AT-cut quartz crystal plate with electrodes is systematically studied from the views of theoretical analysis and numerical simulations.The two-dimensional thermal incremental field equations are solved for the free vibration analysis via the coefficient-formed partial differential equation module of the COMSOL Multiphysics software,from which the frequency spectra,frequency–temperature curves,and mode shapes are discussed in detail.Additionally,the piezoelectric module is utilized to obtain the admittance response under different conditions.It is demonstrated that the parallelism error reduces the resonant frequency.Additionally,symmetry broken by the non-parallelism increases the probability of activity dip and is harmful to QCR’s thermal stability.However,if the top and bottom surfaces incline synchronously in the same direction,the influence of parallelism error is tiny.The conclusions achieved are helpful for the QCR design,and the methodology presented can also be applied to other wave devices.
摘要Two groups of parallelism in Martin Luther King's I Have a Dream are selected to analyze the rhythmic features in terms of prominence and duration. By applying Praat to label the sentences, some findings are as follows: In parallel sentences,the duration of each parallel structure is more or less equal due to the same number of prominence, although sentence length is different; the mean duration of syllables in long sentences is shorter as a result of vowel truncation.
摘要An improvement detecting method was proposed according to the disadvantages of testing method of optical axes parallelism of shipboard photoelectrical theodolite (short for theodolite) based on image processing. Pointolite replaced 0.2'' collimator to reduce the errors of crosshair images processing and improve the quality of image. What’s more, the high quality images could help to optimize the image processing method and the testing accuracy. The errors between the trial results interpreted by software and the results tested in dock were less than 10'', which indicated the improve method had some actual application values.
摘要Parallelism在英语中具有修辞和语法双重功能。在英汉翻译对比研究中,学术界关注较多的是它与中文排比的构成差异及修辞功用,对它的语法功能研究较少。实践证明,英汉排比结构的对等转译不仅有助于保留源文本的语言特色和功能,还可以帮助译文读者更好地把握原文结构,明了行文思路。Parallelism的修辞功能和语法功能同等重要,都应引起译界的足够重视。本文在功能对等理论的指导下,以美国作家M·斯科特·派克的心理学著作——The Road Less Traveled中的排比句的翻译为例,探讨如何通过直译法、转译法、分句法、逆序法、增词法和减词法来提高原文读者和译文读者反应的相似性,实现功能对等目的。
基金the National Natural Science Foundation of China(Nos.61572508,61272144,61303065and 61202121)the National High Technology Research and Development Program(863)of China(No.2012AA010905)+2 种基金the Research Project of National University of Defense Technology(No.JC13-06-02)the Doctoral Fund of Ministry of Education of China(No.20134307120028)the Research Fund for the Doctoral Program of Higher Education of China(No.20114307120013)
摘要The simulation is an important means of performance evaluation of the computer architecture. Nowadays, the serial simulation of general purpose graphics processing unit(GPGPU) architecture is the main bottleneck for the simulation speed. To address this issue, we propose the intra-kernel parallelization on a multicore processor and the inter-kernel parallelization on a multiple-machine platform. We apply these two methods to the GPGPU-sim simulator. The intra-kernel parallelization method firstly parallelizes the serial simulation of multiple compute units in one cycle. Then it parallelizes the timing and functional simulation to reduce the performance loss caused by the synchronization between different compute units. The inter-kernel parallelization method divides multiple kernels of a CUDA program into several groups and distributes these groups across multiple simulation hosts to perform the simulation. Experimental results show that the intra-kernel parallelization method achieves a speed-up of up to 12 with a maximum error rate of 0.009 4% on a 32-core machine, and the inter-kernel parallelization method can accelerate the simulation by a factor of up to 3.9 with a maximum error rate of 0.11% on four simulation hosts. The orthogonality between these two methods allows us to combine them together on multiple multi-core hosts to get further performance improvements.
摘要Define and theory of autocorrelation decision tree (ADT) is introduced. In spatial data mining, spatial parallel query are very expensive operations. A new parallel algorithm in terms of autocorrelation decision tree is presented. And the new method reduces CPU- and I/O-time and improves the query efficiency of spatial data. For dynamic load balancing, there are better control and optimization. Experimental performance comparison shows that the improved algorithm can obtain a optimal accelerator with the same quantities of processors. There are more completely accesses on nodes. And an individual implement of intelligent information retrieval for spatial data mining is presented.