期刊文献+
共找到889篇文章
< 1 2 45 >
每页显示 20 50 100
Complex hexagonal close-packed dendritic growth during alloy solidification by graphics processing unit-accelerated three-dimensional phase-field simulations:demo for Mg–Gd alloy 认领 引用 被引量:3
1
作者 Sheng-Lan Yang Jing Zhong +5 位作者 Kai Wang Xun Kang Jian-Bao Gao Jiong Wang Qian Li Li-Jun Zhang 《Rare Metals》 SCIE EI CAS CSCD 2023年第10期3468-3484,共17页
In this study,insights into the effect of interfacial anisotropy on a complex hexagonal close-packed(hcp) dendritic growth during alloy solidification were gained by graphics processing unit(GPU)-accelerated three-dim... In this study,insights into the effect of interfacial anisotropy on a complex hexagonal close-packed(hcp) dendritic growth during alloy solidification were gained by graphics processing unit(GPU)-accelerated three-dimensional(3D) phase-field simulations,as demonstrated for a Mg-Gd alloy.An anisotropic phasefield model with finite interface dissipation was developed by incorporating the contribution of the anisotropy of interfacial energy into the total free energy functional.The modified spherical harmonic anisotropy function was then chosen for the hcp crystal.The GPU parallel computing algorithm was implemented in the present phase-field model,and a corresponding code was developed in the compute unified device architecture parallel computing platform.Benchmark tests indicated that the calculation efficiency of a single TESLA V100 GPU could be~80times that of open multi-processing(OpenMP) with eight central processing unit cores.By coupling the phase-field model with reliable thermodynamic and interfacial energy descriptions,the 3D phase-field simulation of α-Mg dendritic growth in the Mg-6Gd(in wt%) alloy during solidification was performed.Various two-dimensional dendrite morphologies were revealed by cutting the simulated 3D dendrite along different crystallographic planes.Typical sixfold equiaxed and butterflied microstructures observed in experiments were well reproduced. 展开更多
关键词 Interfacial anisotropy Dendrite solidification Phase-field model Graphics processing unit(GPU) Mg–Gd
暂未订购 下载PDF
Simulation of fluid-structure interaction in a microchannel using the lattice Boltzmann method and size-dependent beam element on a graphics processing unit 认领 引用 被引量:4
2
作者 Vahid Esfahanian Esmaeil Dehdashti Amir Mehdi Dehrouye-Semnani 《Chinese Physics B》 SCIE EI CAS CSCD 2014年第8期389-395,共7页
Fluid-structure interaction (FSI) problems in microchannels play a prominent role in many engineering applications. The present study is an effort toward the simulation of flow in microchannel considering FSI. The b... Fluid-structure interaction (FSI) problems in microchannels play a prominent role in many engineering applications. The present study is an effort toward the simulation of flow in microchannel considering FSI. The bottom boundary of the microchannel is simulated by size-dependent beam elements for the finite element method (FEM) based on a modified cou- ple stress theory. The lattice Boltzmann method (LBM) using the D2Q13 LB model is coupled to the FEM in order to solve the fluid part of the FSI problem. Because of the fact that the LBM generally needs only nearest neighbor information, the algorithm is an ideal candidate for parallel computing. The simulations are carried out on graphics processing units (GPUs) using computed unified device architecture (CUDA). In the present study, the governing equations are non-dimensionalized and the set of dimensionless groups is exhibited to show their effects on micro-beam displacement. The numerical results show that the displacements of the micro-beam predicted by the size-dependent beam element are smaller than those by the classical beam element. 展开更多
关键词 fluid-structure interaction graphics processing unit lattice Boltzmann method size-dependentbeam element
暂未订购 下载PDF
Optimization of a precise integration method for seismic modeling based on graphic processing unit 认领 引用 被引量:2
3
作者 Jingyu Li Genyang Tang Tianyue Hu 《Earthquake Science》 2010年第4期387-393,共7页
General purpose graphic processing unit (GPU) calculation technology is gradually widely used in various fields. Its mode of single instruction, multiple threads is capable of seismic numerical simulation which has ... General purpose graphic processing unit (GPU) calculation technology is gradually widely used in various fields. Its mode of single instruction, multiple threads is capable of seismic numerical simulation which has a huge quantity of data and calculation steps. In this study, we introduce a GPU-based parallel calculation method of a precise integration method (PIM) for seismic forward modeling. Compared with CPU single-core calculation, GPU parallel calculating perfectly keeps the features of PIM, which has small bandwidth, high accuracy and capability of modeling complex substructures, and GPU calculation brings high computational efficiency, which means that high-performing GPU parallel calculation can make seismic forward modeling closer to real seismic records. 展开更多
关键词 precise integration method seismic modeling general purpose GPU graphic processing unit
暂未订购 下载PDF
A graphics processing unit-based robust numerical model for solute transport driven by torrential flow condition 认领 引用 被引量:1
4
作者 Jing-ming HOU Bao-shan SHI +6 位作者 Qiu-hua LIANG Yu TONG Yong-de KANG Zhao-an ZHANG Gang-gang BAI Xu-jun GAO Xiao YANG 《Journal of Zhejiang University-SCIENCE A》 SCIE EI CAS CSCD 2021年第10期835-850,共16页
Solute transport simulations are important in water pollution events.This paper introduces a finite volume Godunovtype model for solving a 4×4 matrix form of the hyperbolic conservation laws consisting of 2D shal... Solute transport simulations are important in water pollution events.This paper introduces a finite volume Godunovtype model for solving a 4×4 matrix form of the hyperbolic conservation laws consisting of 2D shallow water equations and transport equations.The model adopts the Harten-Lax-van Leer-contact(HLLC)-approximate Riemann solution to calculate the cell interface fluxes.It can deal well with the changes in the dry and wet interfaces in an actual complex terrain,and it has a strong shock-wave capturing ability.Using monotonic upstream-centred scheme for conservation laws(MUSCL)linear reconstruction with finite slope and the Runge-Kutta time integration method can achieve second-order accuracy.At the same time,the introduction of graphics processing unit(GPU)-accelerated computing technology greatly increases the computing speed.The model is validated against multiple benchmarks,and the results are in good agreement with analytical solutions and other published numerical predictions.The third test case uses the GPU and central processing unit(CPU)calculation models which take 3.865 s and 13.865 s,respectively,indicating that the GPU calculation model can increase the calculation speed by 3.6 times.In the fourth test case,comparing the numerical model calculated by GPU with the traditional numerical model calculated by CPU,the calculation efficiencies of the numerical model calculated by GPU under different resolution grids are 9.8–44.6 times higher than those by CPU.Therefore,it has better potential than previous models for large-scale simulation of solute transport in water pollution incidents.It can provide a reliable theoretical basis and strong data support in the rapid assessment and early warning of water pollution accidents. 展开更多
关键词 Solute transport Shallow water equations Godunov-type scheme Harten-Lax-van Leer-contact(HLLC)Riemann solver Graphics processing unit(GPU)acceleration technology Torrential flow
暂未订购 下载PDF
Multi-relaxation-time lattice Boltzmann simulations of lid driven flows using graphics processing unit 认领 引用 被引量:1
5
作者 Chenggong LI J.P.Y.MAA 《Applied Mathematics and Mechanics(English Edition)》 SCIE EI CSCD 2017年第5期707-722,共16页
Large eddy simulation (LES) using the Smagorinsky eddy viscosity model is added to the two-dimensional nine velocity components (D2Q9) lattice Boltzmann equation (LBE) with multi-relaxation-time (MRT) to simul... Large eddy simulation (LES) using the Smagorinsky eddy viscosity model is added to the two-dimensional nine velocity components (D2Q9) lattice Boltzmann equation (LBE) with multi-relaxation-time (MRT) to simulate incompressible turbulent cavity flows with the Reynolds numbers up to 1 × 10^7. To improve the computation efficiency of LBM on the numerical simulations of turbulent flows, the massively parallel computing power from a graphic processing unit (GPU) with a computing unified device architecture (CUDA) is introduced into the MRT-LBE-LES model. The model performs well, compared with the results from others, with an increase of 76 times in computation efficiency. It appears that the higher the Reynolds numbers is, the smaller the Smagorinsky constant should be, if the lattice number is fixed. Also, for a selected high Reynolds number and a selected proper Smagorinsky constant, there is a minimum requirement for the lattice number so that the Smagorinsky eddy viscosity will not be excessively large. 展开更多
关键词 large eddy simulation (LES) multi-relaxation-time (MRT) lattice Boltzmann equation (LBE) two-dimensional nine velocity components (D2Q9) Smagorinskymodel graphic processing unit (GPU) computing unified device architecture (CUDA)
暂未订购 下载PDF
Exploiting Parallelism in the Simulation of General Purpose Graphics Processing Unit Program 认领 引用
6
作者 赵夏 马胜 +1 位作者 陈微 王志英 《Journal of Shanghai Jiaotong university(Science)》 EI 2016年第3期280-288,共9页
The simulation is an important means of performance evaluation of the computer architecture. Nowadays, the serial simulation of general purpose graphics processing unit(GPGPU) architecture is the main bottleneck for t... The simulation is an important means of performance evaluation of the computer architecture. Nowadays, the serial simulation of general purpose graphics processing unit(GPGPU) architecture is the main bottleneck for the simulation speed. To address this issue, we propose the intra-kernel parallelization on a multicore processor and the inter-kernel parallelization on a multiple-machine platform. We apply these two methods to the GPGPU-sim simulator. The intra-kernel parallelization method firstly parallelizes the serial simulation of multiple compute units in one cycle. Then it parallelizes the timing and functional simulation to reduce the performance loss caused by the synchronization between different compute units. The inter-kernel parallelization method divides multiple kernels of a CUDA program into several groups and distributes these groups across multiple simulation hosts to perform the simulation. Experimental results show that the intra-kernel parallelization method achieves a speed-up of up to 12 with a maximum error rate of 0.009 4% on a 32-core machine, and the inter-kernel parallelization method can accelerate the simulation by a factor of up to 3.9 with a maximum error rate of 0.11% on four simulation hosts. The orthogonality between these two methods allows us to combine them together on multiple multi-core hosts to get further performance improvements. 展开更多
关键词 general purpose graphics processing unit(GPGPU) multicore intra-kernel inter-kernel parallel
暂未订购 下载PDF
Graphic Processing Unit-Accelerated Neural Network Model for Biological Species Recognition 认领 引用
7
作者 温程璐 潘伟 +1 位作者 陈晓熹 祝青园 《Journal of Donghua University(English Edition)》 EI CAS 2012年第1期5-8,共4页
A graphic processing unit (GPU)-accelerated biological species recognition method using partially connected neural evolutionary network model is introduced in this paper. The partial connected neural evolutionary netw... A graphic processing unit (GPU)-accelerated biological species recognition method using partially connected neural evolutionary network model is introduced in this paper. The partial connected neural evolutionary network adopted in the paper can overcome the disadvantage of traditional neural network with small inputs. The whole image is considered as the input of the neural network, so the maximal features can be kept for recognition. To speed up the recognition process of the neural network, a fast implementation of the partially connected neural network was conducted on NVIDIA Tesla C1060 using the NVIDIA compute unified device architecture (CUDA) framework. Image sets of eight biological species were obtained to test the GPU implementation and counterpart serial CPU implementation, and experiment results showed GPU implementation works effectively on both recognition rate and speed, and gained 343 speedup over its counterpart CPU implementation. Comparing to feature-based recognition method on the same recognition task, the method also achieved an acceptable correct rate of 84.6% when testing on eight biological species. 展开更多
关键词 graphic processing unit(GPU) compute unified device architecture (CUDA) neural network species recognition
暂未订购 下载PDF
Graphic Processing Unit-Accelerated Mutual Information-Based 3D Image Rigid Registration 认领 引用
8
作者 李冠华 欧宗瑛 +1 位作者 苏铁明 韩军 《Transactions of Tianjin University》 EI CAS 2009年第5期375-380,共6页
Mutual information (MI)-based image registration is effective in registering medical images, but it is computationally expensive. This paper accelerates MI-based image registration by dividing computation of mutual ... Mutual information (MI)-based image registration is effective in registering medical images, but it is computationally expensive. This paper accelerates MI-based image registration by dividing computation of mutual information into spatial transformation and histogram-based calculation, and performing 3D spatial transformation and trilinear interpolation on graphic processing unit (GPU). The 3D floating image is downloaded to GPU as flat 3D texture, and then fetched and interpolated for each new voxel location in fragment shader. The transformed resuits are rendered to textures by using frame buffer object (FBO) extension, and then read to the main memory used for the remaining computation on CPU. Experimental results show that GPU-accelerated method can achieve speedup about an order of magnitude with better registration result compared with the software implementation on a single-core CPU. 展开更多
关键词 image registration mutual information graphic processing unit (GPU)
暂未订购 下载PDF
Compute Unified Device Architecture Implementation of Euler/Navier-Stokes Solver on Graphics Processing Unit Desktop Platform for 2-D Compressible Flows 认领 引用
9
作者 Zhang Jiale Chen Hongquan 《Transactions of Nanjing University of Aeronautics and Astronautics》 EI CSCD 2016年第5期536-545,共10页
Personal desktop platform with teraflops peak performance of thousands of cores is realized at the price of conventional workstations using the programmable graphics processing units(GPUs).A GPU-based parallel Euler/N... Personal desktop platform with teraflops peak performance of thousands of cores is realized at the price of conventional workstations using the programmable graphics processing units(GPUs).A GPU-based parallel Euler/Navier-Stokes solver is developed for 2-D compressible flows by using NVIDIA′s Compute Unified Device Architecture(CUDA)programming model in CUDA Fortran programming language.The techniques of implementation of CUDA kernels,double-layered thread hierarchy and variety memory hierarchy are presented to form the GPU-based algorithm of Euler/Navier-Stokes equations.The resulting parallel solver is validated by a set of typical test flow cases.The numerical results show that dozens of times speedup relative to a serial CPU implementation can be achieved using a single GPU desktop platform,which demonstrates that a GPU desktop can serve as a costeffective parallel computing platform to accelerate computational fluid dynamics(CFD)simulations substantially. 展开更多
关键词 graphics processing unit(GPU) GPU parallel computing compute unified device architecture(CUDA)Fortran finite volume method(FVM) acceleration
暂未订购 下载PDF
TIME-DOMAIN INTERPOLATION ON GRAPHICS PROCESSING UNIT 认领 引用 被引量:3
10
作者 XIQI LI GUOHUA SHI YUDONG ZHANG 《Journal of Innovative Optical Health Sciences》 SCIE EI 2011年第1期89-95,共7页
The signal processing speed of spectral domain optical coherence tomography(SD-OCT)has become a bottleneck in a lot of medical applications.Recently,a time-domain interpolation method was proposed.This method can get ... The signal processing speed of spectral domain optical coherence tomography(SD-OCT)has become a bottleneck in a lot of medical applications.Recently,a time-domain interpolation method was proposed.This method can get better signal-to-noise ratio(SNR)but much-reduced signal processing time in SD-OCT data processing as compared with the commonly used zeropadding interpolation method.Additionally,the resampled data can be obtained by a few data and coefficients in the cutoff window.Thus,a lot of interpolations can be performed simultaneously.So,this interpolation method is suitable for parallel computing.By using graphics processing unit(GPU)and the compute unified device architecture(CUDA)program model,time-domain interpolation can be accelerated significantly.The computing capability can be achieved more than 250,000 A-lines,200,000 A-lines,and 160,000 A-lines in a second for 2,048 pixel OCT when the cutoff length is L=11,L=21,and L=31,respectively.A frame SD-OCT data(400A-lines×2,048 pixel per line)is acquired and processed on GPU in real time.The results show that signal processing time of SD-OCT can befinished in 6.223 ms when the cutoff length L=21,which is much faster than that on central processing unit(CPU).Real-time signal processing of acquired data can be realized. 展开更多
关键词 Optical coherence tomography real-time signal processing graphics processing unit GPU CUDA
暂未订购 下载PDF
The inversion of density structure by graphic processing unit(GPU) and identification of igneous rocks in Xisha area 认领 引用 被引量:1
11
作者 Lei Yu Jian Zhang +2 位作者 Wei Lin Rongqiang Wei Shiguo Wu 《Earthquake Science》 2014年第1期117-125,共9页
Organic reefs, the targets of deep-water petro- leum exploration, developed widely in Xisha area. However, there are concealed igneous rocks undersea, to which organic rocks have nearly equal wave impedance. So the ig... Organic reefs, the targets of deep-water petro- leum exploration, developed widely in Xisha area. However, there are concealed igneous rocks undersea, to which organic rocks have nearly equal wave impedance. So the igneous rocks have become interference for future explo- ration by having similar seismic reflection characteristics. Yet, the density and magnetism of organic reefs are very different from igneous rocks. It has obvious advantages to identify organic reefs and igneous rocks by gravity and magnetic data. At first, frequency decomposition was applied to the free-air gravity anomaly in Xisha area to obtain the 2D subdivision of the gravity anomaly and magnetic anomaly in the vertical direction. Thus, the dis- tribution of igneous rocks in the horizontal direction can be acquired according to high-frequency field, low-frequency field, and its physical properties. Then, 3D forward model- ing of gravitational field was carried out to establish the density model of this area by reference to physical properties of rocks based on former researches. Furthermore, 3D inversion of gravity anomaly by genetic algorithm method of the graphic processing unit (GPU) parallel processing in Xisha target area was applied, and 3D density structure of this area was obtained. By this way, we can confine the igneous rocks to the certain depth according to the density of the igneous rocks. The frequency decomposition and 3D inversion of gravity anomaly by genetic algorithm method of the GPU parallel processing proved to be a useful method for recognizing igneous rocks to its 3D geological position. So organic reefs and igneous rocks can be identified, which provide a prescient information for further exploration. 展开更多
关键词 Xisha area Organic reefs and igneous rocks -Frequency decomposition of potential field 3D inversionof the graphic processing unit (GPU) parallel processing
暂未订购 下载PDF
Bypass-Enabled Thread Compaction for Divergent Control Flow in Graphics Processing Units 认领 引用
12
作者 LI Bingchao WEI Jizeng +1 位作者 GUO Wei SUN Jizhou 《Journal of Shanghai Jiaotong university(Science)》 EI 2021年第2期245-256,共12页
Graphics processing units(GPUs)employ the single instruction multiple data(SIMD)hardware to run threads in parallel and allow each thread to maintain an arbitrary control flow.Threads running concurrently within a war... Graphics processing units(GPUs)employ the single instruction multiple data(SIMD)hardware to run threads in parallel and allow each thread to maintain an arbitrary control flow.Threads running concurrently within a warp may jump to different paths after conditional branches.Such divergent control flow makes some lanes idle and hence reduces the SIMD utilization of GPUs.To alleviate the waste of SIMD lanes,threads from multiple warps can be collected together to improve the SIMD lane utilization by compacting threads into idle lanes.However,this mechanism induces extra barrier synchronizations since warps have to be stalled to wait for other warps for compactions,resulting in that no warps are scheduled in some cases.In this paper,we propose an approach to reduce the overhead of barrier synchronizat ions induced by compactions,In our approach,a compaction is bypassed by warps whose threads all jump to the same path after branches.Moreover,warps waiting for a compaction can also bypass this compaction when no warps are ready for issuing.In addition,a compaction is canceled if idle lanes can not be reduced via this compaction.The experimental results demonstrate that our approach provides an average improvement of 21%over the baseline GPU for applications with massive divergent branches,while recovering the performance loss induced by compactions by 13%on average for applications with many non-divergent control flows. 展开更多
关键词 graphics processing unit(GPU) single instruction ultiple data(SIMD) thread warps bypass
暂未订购 下载PDF
Graphic Processing Unit Based Phase Retrieval and CT Reconstruction for Differential X-Ray Phase Contrast Imaging 认领 引用
13
作者 陈晓庆 王宇杰 孙建奇 《Journal of Shanghai Jiaotong university(Science)》 EI 2014年第5期550-554,共5页
Compared with the conventional X-ray absorption imaging, the X-ray phase-contrast imaging shows higher contrast on samples with low attenuation coefficient like blood vessels and soft tissues. Among the modalities of ... Compared with the conventional X-ray absorption imaging, the X-ray phase-contrast imaging shows higher contrast on samples with low attenuation coefficient like blood vessels and soft tissues. Among the modalities of phase-contrast imaging, the grating-based phase contrast imaging has been widely accepted owing to the advantage of wide range of sample selections and exemption of coherent source. However, the downside is the substantially larger amount of data generated from the phase-stepping method which slows down the reconstruction process. Graphic processing unit(GPU) has the advantage of allowing parallel computing which is very useful for large quantity data processing. In this paper, a compute unified device architecture(CUDA) C program based on GPU is introduced to accelerate the phase retrieval and filtered back projection(FBP) algorithm for grating-based tomography. Depending on the size of the data, the CUDA C program shows different amount of speed-up over the standard C program on the same Visual Studio 2010 platform. Meanwhile, the speed-up ratio increases as the size of data increases. 展开更多
关键词 grating-based phase contrast imaging parallel computing graphic processing unit(GPU) compute unified device architecture(CUDA) filtered back projection(FBP)
暂未订购 下载PDF
NGP-ERGAS: Revisit Instant Neural Graphics Primitives with the Relative Dimensionless Global Error in Synthesis 认领 引用
14
作者 Dongheng Ye Heping Li +2 位作者 Ning An Jian Cheng Liang Wang 《Computers, Materials & Continua》 SCIE EI 2025年第8期3731-3747,共17页
The newly emerging neural radiance fields(NeRF)methods can implicitly fulfill three-dimensional(3D)reconstruction via training a neural network to render novel-view images of a given scene with given posed images.The ... The newly emerging neural radiance fields(NeRF)methods can implicitly fulfill three-dimensional(3D)reconstruction via training a neural network to render novel-view images of a given scene with given posed images.The Instant Neural Graphics Primitives(Instant-NGP)method further improves the position encoding of NeRF.It obtains state-of-the-art efficiency.However,only a local pixel-wised loss is considered when training the Instant-NGP while overlooking the nonlocal structural information between pixels.Despite a good quantitative result,it leads to a poor visual effect,especially the completeness.Inspired by the stochastic structural similarity(S3IM)method that exploits nonlocal structural information of groups of pixels,this paper proposes a new method to improve the completeness of fast novel view synthesis.The proposed method first extends the thread-wised processing of the Instant-NGP to the processing in a customthread block(i.e.,a group of threads).Then,the relative dimensionless global error in synthesis,i.e.,Erreur Relative Globale Adimensionnelle de Synthese(ERGAS),of a group of pixels corresponding to a group of threads is computed and incorporated into the loss function.Extensive experiments validate the proposed method.It can obtain better quantitative results than the original Instant-NGP with fewer iteration steps.PSNR is increased by 1%.Amazing qualitative results are obtained,especially for delicate structures and details such as lines and continuous structures.With the dramatic improvements in the visual effects,our method can boost the practicability of implicit 3D reconstruction in applications such as self-driving and augmented reality. 展开更多
关键词 Neural radiance fields novel view synthesis 3D reconstruction graphic processing unit
暂未订购 下载PDF
适用于大规模光伏电站的电磁暂态细粒度并行仿真方法 认领 引用 被引量:2
15
作者 曹书豪 许建中 +2 位作者 周宇权 陈浩 冯谟可 《电工技术学报》 EI CSCD 北大核心 2026年第8期2658-2671,共14页
大规模光伏电站的电磁暂态精细化仿真是观测其详细内部特性的重要手段,支撑内部故障溯源、场站与系统间振荡等研究。现有对光伏电站的仿真未深入考虑建模方式、并行仿真方法与硬件间的交互关系,难以兼顾规模、精度和效率,为此该文提出... 大规模光伏电站的电磁暂态精细化仿真是观测其详细内部特性的重要手段,支撑内部故障溯源、场站与系统间振荡等研究。现有对光伏电站的仿真未深入考虑建模方式、并行仿真方法与硬件间的交互关系,难以兼顾规模、精度和效率,为此该文提出一种深度融合等值建模和图形处理单元(GPU)技术的电磁暂态细粒度并行仿真方法。首先,根据GPU软硬件的架构特点,设计模块化光伏系统并行计算程序基本逻辑。其次,考虑换流器闭锁模式,利用开关函数模型实现光伏单元交、直流侧解耦划分,基于割集矩阵实现交流侧端口等效,将光伏单元整合为集群等效模型参与系统解算。最后,利用GPU单指令多数据特点,在各光伏单元独立解算任务的粗粒度并行基础上,叠加元件内部更新过程的细粒度并行,将海量模块化光伏单元同阶矩阵进行高效批处理,进而实现大规模光伏电站的元件级细粒度并行仿真。对比GPU仿真平台与PSCAD/EMTDC的仿真结果,验证了所提并行方法的仿真精度和加速比。 展开更多
关键词 大规模光伏电站 等值建模 细粒度 并行仿真 图形处理单元(GPU)
暂未订购 下载PDF
基于GPU的大规模新能源电网EMT仿真并行加速方法 认领 引用
16
作者 于智同 杨明皓 +3 位作者 宋炎侃 黄少伟 陈颖 沈沉 《电力自动化设备》 EI CSCD 北大核心 2026年第8期123-131,共9页
高比例新能源并网使含海量电力电子设备的大规模电磁暂态仿真在计算效率与规模扩展性上面临瓶颈。图形处理器具备高吞吐并行优势,但直接应用时面临控制拓扑异构导致的线程束发散、图式计算带来的访存不连续和动态解析引起的指令冗余三... 高比例新能源并网使含海量电力电子设备的大规模电磁暂态仿真在计算效率与规模扩展性上面临瓶颈。图形处理器具备高吞吐并行优势,但直接应用时面临控制拓扑异构导致的线程束发散、图式计算带来的访存不连续和动态解析引起的指令冗余三大难点。为此,提出一种面向大规模新能源电网的异构并行加速方法。引入图同构思想,利用图聚类算法实现控制系统的拓扑级识别与聚合,消除异构控制器导致的指令流分歧;构建面向图形处理器的自动代码生成框架,将聚合后的模型转化为无分支、访存规整的计算内核,提升执行效率。算例测试表明,在包含数千台新能源设备的仿真场景中,所提方法相比传统仿真方法可获得近3倍的加速比,且具备优异的规模可扩展性。 展开更多
关键词 电磁暂态仿真 图形处理器 细粒度并行 图同构 自动代码生成 异构计算
暂未订购 下载PDF
面向GPU的稀疏对角矩阵自适应SpMV优化方法 认领 引用
17
作者 王宇华 何俊飞 +2 位作者 张宇琪 兰海燕 曹林琳 《计算机工程》 CAS CSCD 北大核心 2026年第3期332-345,共14页
稀疏矩阵向量乘(SpMV)是稀疏线性系统的计算核心和瓶颈,其运算效率会影响迭代求解器的整体性能,其优化研究一直是科学计算和工程应用领域中的研究热点之一。偏微分方程的离散化会产生稀疏对角矩阵,由于其多样的非零元分布,导致没有一种... 稀疏矩阵向量乘(SpMV)是稀疏线性系统的计算核心和瓶颈,其运算效率会影响迭代求解器的整体性能,其优化研究一直是科学计算和工程应用领域中的研究热点之一。偏微分方程的离散化会产生稀疏对角矩阵,由于其多样的非零元分布,导致没有一种方法能够在所有矩阵中取得最优时间性能。针对上述问题,提出一种面向图形处理单元(GPU)的稀疏对角矩阵自适应SpMV优化方法AST(Adaptive SpMV Tuning)。该方法通过设计特征空间,构建特征提取器,提取矩阵结构精细特征,通过深入分析特征和SpMV方法的相关性,建立可扩展的候选方法集合,形成特征和最优方法的映射关系,构建性能预测工具,实现矩阵最优方法的高效预测。实验结果表明,AST能够取得85.8%的预测准确率,平均时间性能损失为0.09,相比于DIA(Diagonal)、HDIA(Hacked DIA)、HDC(Hybrid of DIA and Compressed Sparse Row)、DIA-Adaptive和DRM(Divide-Rearrange and Merge),能够获得平均20.19、1.86、3.06、3.72和1.53倍的内核运行时间加速和1.05、1.28、12.45、1.94和0.97倍的浮点运算性能加速。 展开更多
关键词 稀疏矩阵向量乘 稀疏对角矩阵 图形处理单元 自适应优化方法 矩阵结构特征
暂未订购 下载PDF
基于单帧多通道配准的穿墙雷达运动目标实时检测方法设计与实现 认领 引用
18
作者 曾小路 刘国振 +2 位作者 钟世超 刘飞杨 杨小鹏 《系统工程与电子技术》 EI CSCD 北大核心 2026年第8期2560-2580,共21页
穿墙雷达运动目标检测可实时获取封闭环境中的运动目标信息,在城市反恐维稳等方面发挥着重要的作用。当前研究大多集中在穿墙雷达运动目标检测算法方面,多数算法复杂度较高,难以满足实时检测的应用化需求。针对以上问题,提出一种单帧多... 穿墙雷达运动目标检测可实时获取封闭环境中的运动目标信息,在城市反恐维稳等方面发挥着重要的作用。当前研究大多集中在穿墙雷达运动目标检测算法方面,多数算法复杂度较高,难以满足实时检测的应用化需求。针对以上问题,提出一种单帧多通道配准的穿墙雷达运动目标实时检测方法及其高效实现。首先,提出一种基于单帧多通道数据频域相位补偿与对消的杂波抑制方法,有效减少传统基于多帧数据对消方法中的数据等待延迟。其次,设计一种基于图像处理单元(graphics processing unit,GPU)多数据流异步并行和线程层次结构的穿墙雷达运动目标实时检测实现方法。该方法包括基于GPU的杂波抑制多流异步并行技术、后向投影成像算法粗细粒度并行处理技术和二维恒虚警率检测算法块间并行块内规约并行处理技术。仿真与实测实验表明,所提方法能在1.078 ms内完成基于单帧、8通道和8192个距离向采样点回波数据的墙后运动目标实时检测,相较于常规基于中央处理器处理方法,处理效率提升92倍,表明了所提方法的有效性和鲁棒性。 展开更多
关键词 穿墙雷达 移动目标 实时检测 图形处理器 并行
暂未订购 下载PDF
非对称容差关系粗糙近似集的GPU加速方法 认领 引用
19
作者 吴正江 王梦松 武星晨 《计算机工程》 CAS CSCD 北大核心 2026年第8期247-259,共13页
大数据时代信息来源纷繁复杂,收集数据形成的信息系统很难保证其完整性。在越来越大的不完备信息系统(IIS)中,使用基于非对称容差关系的粗糙集理论进行知识蒸馏,提升近似集的计算速度成为其应用之前必须要解决的问题。针对非对称容差关... 大数据时代信息来源纷繁复杂,收集数据形成的信息系统很难保证其完整性。在越来越大的不完备信息系统(IIS)中,使用基于非对称容差关系的粗糙集理论进行知识蒸馏,提升近似集的计算速度成为其应用之前必须要解决的问题。针对非对称容差关系的冗余容差类问题,提出非对称容差关系的改进方案,设计不完备信息系统中上、下近似集的布尔矩阵表示方法,计算非对称容差关系的粗糙近似集的矩阵分块算法,并在图像处理单元(GPU)上实现近似集计算过程的加速,提高近似集的计算效率。另外,针对GPU存储空间有限的现状,构建不完备信息系统中对象之间的层次结构及其算法,将全局关系矩阵转换成多个局部关系矩阵,缓解计算最近容差类过程中GPU存储压力。在UCI数据集和生成数据集上的实验结果表明,基于最近容差关系的容差类数量相较于基准情况明显减少,同时,矩阵分块算法在GPU上实现了对近似集计算过程的有效加速,相比CPU串行计算和分布式并行计算,GPU分块算法执行速度平均提高了16.69倍和3.89倍。 展开更多
关键词 图像处理单元 近似集 最近容差关系 矩阵分块 不完备信息系统
暂未订购 下载PDF
面向GPU的低能耗数据传输的组重映射编码方法 认领 引用
20
作者 章铁飞 邢建国 《计算机工程与科学》 CSCD 北大核心 2026年第5期803-809,共7页
现代图形处理单元(GPU)的高性能计算能力,依赖于高带宽的图形DDR(GDDR)接口。高带宽的数据传输速率导致高能耗,特别是GDDR的伪开漏(POD)I/O接口中传输逻辑1值的不对称能耗。通过减少数据传输过程中高能耗的逻辑1值,可以缓解数据传输时... 现代图形处理单元(GPU)的高性能计算能力,依赖于高带宽的图形DDR(GDDR)接口。高带宽的数据传输速率导致高能耗,特别是GDDR的伪开漏(POD)I/O接口中传输逻辑1值的不对称能耗。通过减少数据传输过程中高能耗的逻辑1值,可以缓解数据传输时的高能耗问题。提出一种基于逻辑1值数量的组重映射编码方法。首先,将待传输数据按4位划分为基本单元,根据单元包含的逻辑1值数量再分组,然后将包含逻辑1值较多且数量较多的组映射编码为包含逻辑1值较少且数量较少的组,以最小化全局的逻辑1值数量。在现代GPU架构上评估,结果显示组重映射编码方法可以有效减少各种应用程序在数据传输时的逻辑1值的数量,平均降低比例达到26%,证明了方法的有效性。 展开更多
关键词 数据传输能耗 GDDR I/O接口 组重映射编码 图形处理器
暂未订购 下载PDF
上一页 1 2 45 下一页 到第
在线咨询 使用帮助 返回顶部 意见反馈