The satellite laser ranging (SLR) data quality from the COMPASS was analyzed, and the difference between curve recognition in computer vision and pre-process of SLR data finally proposed a new algorithm for SLR was ...The satellite laser ranging (SLR) data quality from the COMPASS was analyzed, and the difference between curve recognition in computer vision and pre-process of SLR data finally proposed a new algorithm for SLR was discussed data based on curve recognition from points cloud is proposed. The results obtained by the new algorithm are 85 % (or even higher) consistent with that of the screen displaying method, furthermore, the new method can process SLR data automatically, which makes it possible to be used in the development of the COMPASS navigation system.展开更多
3D laser scanning technology is widely used in underground openings for high-precision,rapid,and nondestructive structural evaluations.Segmenting large 3D point cloud datasets,particularly in coal mine roadways with m...3D laser scanning technology is widely used in underground openings for high-precision,rapid,and nondestructive structural evaluations.Segmenting large 3D point cloud datasets,particularly in coal mine roadways with multi-scale targets,remains challenging.This paper proposes an enhanced segmentation method integrating improved PointNet++with a coverage-voted strategy.The coverage-voted strategy reduces data while preserving multi-scale target topology.The segmentation is achieved using an enhanced PointNet++algorithm with a normalization preprocessing head,resulting in a 94%accuracy for common supporting components.Ablation experiments show that the preprocessing head and coverage strategies increase segmentation accuracy by 20%and 2%,respectively,and improve Intersection over Union(IoU)for bearing plate segmentation by 58%and 20%.The accuracy of the current pretraining segmentation model may be affected by variations in surface support components,but it can be readily enhanced through re-optimization with additional labeled point cloud data.This proposed method,combined with a previously developed machine learning model that links rock bolt load and the deformation field of its bearing plate,provides a robust technique for simultaneously measuring the load of multiple rock bolts in a single laser scan.展开更多
The precision of shape representation and the dimensionality of the design space significantly influence the cost and outcomes of aerodynamic optimization.The design space can be represented more compactly by maintain...The precision of shape representation and the dimensionality of the design space significantly influence the cost and outcomes of aerodynamic optimization.The design space can be represented more compactly by maintaining geometric precision while reducing dimensions,hence enhancing the cost-effectiveness of the optimization process.This research presents a new point cloud Autoencoder Based on Uncertainty Feature Enhancement(AE-BUFE)architecture,designed to attain efficient and precise generalized representations of 3D aircraft through uncertainty analysis of the deformation relationships among surface grid points.The deep learning architecture consists of two components:the uncertainty index-based feature enhancement module and the point cloud autoencoder module.It learns the shape features of the point cloud geometric representation to establish a low-dimensional latent space.To assess and evaluate the efficiency of the method,a comparison was conducted with the prevailing point cloud autoencoder architecture and the proper orthogonal decomposition linear dimensionality reduction method under conditions of complex shape deformation.The results show that the new architecture significantly improves the extraction effect of the low-dimensional latent space.Then,this paper developed the surrogate-based optimization framework based on the AE-BUFE parameterization method and completed a multi-objective aerodynamic optimization design for a wide-speed-range vehicle considering volume and moment constraints.While ensuring the take-off and landing performance,the aerodynamic performance is improved under transonic and hypersonic conditions,which verifies the efficiency and engineering practicability of this method.展开更多
Recently,large-scale deep learning models have been increasingly adopted for point cloud classification.However,thesemethods typically require collecting extensive datasets frommultiple clients,which may lead to priva...Recently,large-scale deep learning models have been increasingly adopted for point cloud classification.However,thesemethods typically require collecting extensive datasets frommultiple clients,which may lead to privacy leaks.Federated learning provides an effective solution to data leakage by eliminating the need for data transmission,relying instead on the exchange of model parameters.However,the uneven distribution of client data can still affect the model’s ability to generalize effectively.To address these challenges,we propose a new framework for point cloud classification called Federated Dynamic Aggregation Selection Strategy-based Multi-Receptive Field Fusion Classification Framework(FDASS-MRFCF).Specifically,we tackle these challenges with two key innovations:(1)During the client local training phase,we propose a Multi-Receptive Field Fusion Classification Model(MRFCM),which captures local and global structures in point cloud data through dynamic convolution and multi-scale feature fusion,enhancing the robustness of point cloud classification.(2)In the server aggregation phase,we introduce a Federated Dynamic Aggregation Selection Strategy(FDASS),which employs a hybrid strategy to average client model parameters,skip aggregation,or reallocate local models to different clients,thereby balancing global consistency and local diversity.We evaluate our framework using the ModelNet40 and ShapeNetPart benchmarks,demonstrating its effectiveness.The proposed method is expected to significantly advance the field of point cloud classification in a secure environment.展开更多
In recent years,drought-induced soil cracking has become increasingly prevalent,posing significant challenges to geotechnical engineering applications due to the associated hazards.Previous studies have investigated s...In recent years,drought-induced soil cracking has become increasingly prevalent,posing significant challenges to geotechnical engineering applications due to the associated hazards.Previous studies have investigated soil cracking behaviors and mechanisms via multi-scale methods.However,research on the spatial distribution characteristics of soil cracks and their depth evolution mechanisms is still rare.This paper analyzes the three-dimensional(3D)deformation of fiveclayey materials.In this study,3D surface elevation data of soil samples were acquired using a Gocator 3110 structured-light sensor.The reconstruction process involved point cloud acquisition and rasterization,Delaunay triangulation meshing,and calculation of crack depth.Different cracking behaviors are investigated,such as cracking in opening mode and shearing mode,coalescence and bifurcation of cracks,etc.The mineralogical analysis is carried out to reveal different cracking behaviors of clays.It is concluded that the crack depth of montmorillonite is larger than that of kaolin,reflectingits strong shrinkage characteristic,which makes the crack evolve deeper than the other materials.This technique enables multi-dimensional observation of soil crack propagation.The mineralogical analysis is carried out to reveal the different cracking behaviors of clays.It is concluded that microstructure and mineralogy are key factors influencingcracking behavior.Understanding soil cracking behaviors is significantfor engineering protection,geohazard prevention,and geological monitoring.展开更多
The virtual preassembly of super-high steel bridge towers faces a challenge in the efficient and precise extraction of complex cross-sectional features.Factors such as fabrication errors,gravity-induced deformations,a...The virtual preassembly of super-high steel bridge towers faces a challenge in the efficient and precise extraction of complex cross-sectional features.Factors such as fabrication errors,gravity-induced deformations,and temperature fluctuations can compromise the accuracy of contour extraction.To address these limitations,an improved Alpha-shape-based point cloud contour extraction method is proposed.The proposed approach uses a hierarchical strategy to process three-dimensional laser scanning point clouds.The processed data are then subjected to curvatureadaptive voxel filtering to reduce acquisition noise.In addition,an enhanced iterative closest point(ICP)variant with correspondence validation accurately aligns the discrete point cloud segments.The proposed curvature-responsive Alpha-shape framework enables multiscale contour delineation through topology-adaptive threshold modulation,which resolves boundary ambiguities in geometrically complex cross-sections.The method was experimentally validated using field-acquired measurement datasets from the Zhangjinggao Yangtze River Bridge tower segments,confirming its capability to reconstruct noncanonical cross-sectional geometries.Three contour extraction methods,including Poisson reconstruction,the conventional Alpha-shape algorithm,and random sample consensus with ICP(RANSAC-ICP),were compared to evaluate the performance of the proposed Alpha-shape algorithm.The results demonstrate that the proposed method achieves superior contour extraction accuracy and data reduction efficiency,highlighting its effectiveness in contour extraction tasks.展开更多
Three-dimensional(3D)point cloud semantic segmentation is a core task in indoor scene understanding,providing detailed semantic information about spatial structures and object categories in indoor environments.Althoug...Three-dimensional(3D)point cloud semantic segmentation is a core task in indoor scene understanding,providing detailed semantic information about spatial structures and object categories in indoor environments.Although methods based on deep learning have made steady progress in recent years,accurately segmenting complex indoor scenes remains challenging due to the unordered nature of point clouds and variations across large scales.Most existing networks have limited capability for multi-scale feature aggregation and struggle to balance local geometric details with global semantic context.These issues are further exacerbated by hierarchical downsampling,which often leads to the loss of fine-grained structural information.Moreover,feature interaction restricted to local neighborhoods may limit the capture of non-local semantic dependencies in complex indoor scenes.To address these limitations,we propose PointNMSA(PointNeXt with Non-local Multi-Scale Aggregation),an improved semantic segmentation network built upon the PointNeXt backbone.A Multi-Scale Feature Enhancement(MSFE)module is introduced in the decoding stage to fuse features from different encoding levels,and further refines the fused features to produce more stable multi-scale representations,which preserves geometric details across scales.In addition,a Convolution-Attention Mixing(CA-Mix)module is designed to jointly integrate local spatial structures and non-local contextual dependencies via dual-stream aggregation and multi-dimensional attention fusion,thereby enabling more discriminative feature representations.Experiments on the Stanford Large-Scale 3D Indoor Spaces(S3DIS)benchmark demonstrate the effectiveness of PointNMSA.On the Area 5 test split,PointNMSA achieves a mean intersection over union(mIoU)of 65.10%,outperforming the PointNeXt baseline by 1.59%,while introducing only a modest increase in computational cost(latency from 42.24 to 45.18 ms and parameters from 3.16 to 8.67M).Despite the noticeable growth in parameter count,the increase in inference latency remains relatively limited,indicating a favorable trade-off between segmentation accuracy and computational efficiency.Additional cross-dataset experiments on ScanNet further verify that PointNMSA maintains stable gains under different indoor scene distributions.Such performance gains suggest that PointNMSA provides a more robust and generalizable solution for semantic segmentation in large-scale indoor environments with complex structural layouts.展开更多
Data-driven approaches have shown great advantage in rapidly and accurately predicting pressure coefficient distributions,which is of crucial importance to efficient aircraft design.Nevertheless,most data-driven appro...Data-driven approaches have shown great advantage in rapidly and accurately predicting pressure coefficient distributions,which is of crucial importance to efficient aircraft design.Nevertheless,most data-driven approaches still encounter limitations in characterizing diverse aerodynamic configurations and adapting to varying grid densities,which have hindered their engineering applicability.In response to these challenges,this work adopts point clouds,a specific type of geometric data structure that is inherently suitable for uniformly characterizing diverse 2D/3D geometric shapes as the input for deep learning-based prediction of pressure coefficient distribution.By augmenting the dimensions of point cloud coordinates for local feature enhancement and utilizing the symmetric function“max pooling”to extract global features,the proposed aerodynamic model establishes the mapping between point cloud coordinates and pressure coefficients.Basic aerodynamic configurations like airfoils and wings are employed as test cases,the results demonstrate that the proposed model achieves both high accuracy and robust generalizability across variable geometries.For class-shape transformation-perturbed airfoils,the prediction error can be reduced to one-third of that of the conventional parameterization-based model.For airfoils selected in the University of Illinois Urbana-Champaign airfoil dataset,among which airfoil profiles are widely distributed,the average error of the proposed approach remains approximately 1.5%,whereas the parameterization-based model may fail.For wings,the prediction error still stays below 2.5%.Finally,the model exhibits strong robustness and generalizability across different point cloud densities.In conclusion,this work makes a breakthrough in predicting pressure coefficient distribution for variable geometric configurations,establishing the foundational framework for designing a large model capable of predicting distributed aerodynamic loads in aerospace applications.展开更多
Accurate extraction of rock mass discontinuity parameters is crucial for stability assessment and engineering safety.High-resolution remote sensing facilitates automated extraction,but its effectiveness relies heavily...Accurate extraction of rock mass discontinuity parameters is crucial for stability assessment and engineering safety.High-resolution remote sensing facilitates automated extraction,but its effectiveness relies heavily on precise normal estimation to ensure geometric reliability.Conventional methods struggle to preserve sharp features such as edges and corners,thereby reducing accuracy.To address this,we propose a normal estimation method based on local geometric adjustment that enhances feature extraction while maintaining sharp geometries.The approach consists of four steps:(1)classifying points,(2)applying normal and axial projections,(3)fitting segmentation lines via least squares,and(4)refining normals by optimizing local neighborhoods.The proposed method was evaluated on computer-aided design(CAD)models,real objects,and rock mass point clouds,and benchmarked against eight representative algorithms,including principal component analysis(PCA),2-Jet PCA,Voronoi-based PCA,PCPNet,neural gradient function(NeuralGF),low rank representation(LRR),normal estimation via shifted neighborhood(NSN)and pair consistency voting(PCV).Experimental results demonstrate that our method achieves superior accuracy and efficiency,significantly improving structural plane extraction and ensuring better preservation of sharp geometric features.展开更多
3D single object tracking(SOT)based on point clouds is a fundamental task for environmental perception in autonomous driving and dynamic scene understanding in robotics.Recent technological advancements in this field ...3D single object tracking(SOT)based on point clouds is a fundamental task for environmental perception in autonomous driving and dynamic scene understanding in robotics.Recent technological advancements in this field have significantly bolstered the environmental interaction capabilities of intelligent systems.This field faces persistent challenges,including feature degradation induced by point cloud sparsity,representation drift caused by non-rigid deformation,and occlusion in complex scenarios.Traditional appearance matching methods,particularly those relying on Siamese networks,are severely constrained by point cloud characteristics,often failing under rapid motions or structural ambiguities among similar objects.In response,the research paradigm has progressively evolved toward motion-centric modeling approaches.These emerging frameworks utilize spatio-temporal joint modeling and geometric shape completion to attain notable performance gains.Furthermore,the incorporation of attention mechanisms and State Space Model(SSM)has enabled more effective multi-scale spatio-temporal feature association,which is particularly beneficial for long-term tracking scenarios.To the best of our knowledge,this is the first comprehensive survey dedicated to 3D single object tracking in point clouds.We provide a detailed analysis of current tracking methods,scrutinizing their limitations regarding multi-object interference and analyzing the trade-off between accuracy and computational efficiency.Finally,we discuss potential future directions,including the development of lightweight models for edge deployment and the integration of cross-modal fusion strategies.展开更多
To support the process of grasping objects on a tabletop for the blind or robotic arm,it is necessary to address fundamental computer vision tasks,such as detecting,recognizing,and locating objects in space,and determ...To support the process of grasping objects on a tabletop for the blind or robotic arm,it is necessary to address fundamental computer vision tasks,such as detecting,recognizing,and locating objects in space,and determining the position of the grasping information.These results can then be used to guide the visually impaired or to execute grasping tasks with a robotic arm.In this paper,we collected,annotated,and published the benchmark TQUGraspingObject dataset for testing,validation,and evaluation of deep learning(DL)models for detecting,recognizing,and localizing grasping objects in 2D and 3D space,especially 3D point cloud data.Our dataset is collected in a shared room,with common everyday objects placed on the tabletop in jumbled positions by Intel RealSense D435(IR-D435).This dataset includes more than 63k RGB-D pairs and related data such as normalized 3D object point cloud,3D object point cloud segmented,coordinate system normalizationmatrix,3D object point cloud normalized,and hand pose for grasping each object.At the same time,we also conducted experiments on fourDL networks with the best performance:SSD-MobileNetV3,ResNet50-Transformer,ResNet101-Transformer,and YOLOv12.The results present that YOLOv12 has the most suitable results in detecting and recognizing objects in images.All data,annotations,toolkit,source code,point cloud data,and results are publicly available on our project website:http://gffzz188fe103f8f1460asfccnxf5ob05c6bw0.ffgz.tsg.suse.edu.cn/HuaTThanhIT2327Tqu/datasetv2.展开更多
In the field of aircraft design and maintenance,with the innovation of cabin cable three-dimensional(3D)scanning and sensor technology,high-precision cabin point cloud data has become the key to improving the accuracy...In the field of aircraft design and maintenance,with the innovation of cabin cable three-dimensional(3D)scanning and sensor technology,high-precision cabin point cloud data has become the key to improving the accuracy of cabin navigation and building a realistic virtual reality environment.In the face of largescale point cloud data,how to efficiently and uniformly construct a realistic virtual reality environment has become a challenge.In this paper,we propose a new low-parametric point cloud upsampling network(LPNet),which is based on the no-learn model to learn the complementary geometric knowledge between point clouds based on some simple data transformations,to efficiently retain the geometric properties of point clouds,and then input the results into the up-sampling module,and simply insert a few layers of multilayer perceptron(MLP)to efficiently generate high-resolution point clouds.It is able to efficiently generate high-resolution point clouds,showing great flexibility and realizing the efficient use of computational resources.展开更多
Machine vision-based detection methods have been widely applied in the detection of aircraft skin damage.During drone inspection processes,a key step is to spatially locate high-resolution detailed images of aircraft ...Machine vision-based detection methods have been widely applied in the detection of aircraft skin damage.During drone inspection processes,a key step is to spatially locate high-resolution detailed images of aircraft skin from multiple angles onto a threedimensional point cloud model of the aircraft.This relies on the rigid registration of image center position coordinate point cloud with the aircraft 3D point cloud.To address the issues of low accuracy and poor robustness encountered by existing registration algorithms when dealing with heterogeneous point clouds with significant differences in density and low overlap,this paper presents a novel cross-source point cloud registration network.The network integrates multi-scale information from the point cloud and employs an attention mechanism to identify representative overlapping points.First,the network achieves initial correspondences using the multi-scale geometric features and positional information of the point cloud.Then,an overlapping feature guidance module predicts the overlapping score of the point cloud.By utilizing information interaction through the attention mechanism,the network combines point overlapping scores with fused features to filter out representative overlapping points,achieving precise correspondences in the point cloud.The network employs weighted singular value decomposition(SVD)to estimate two sets of transformation matrices,yielding the relative pose parameters of the point cloud.Experiments were conducted in an unsupervised manner.The experimental results on the ModelNet40 dataset and the aero object dataset aircraft measurement data showed that,compared to other existing traditional and learning-based methods,this approach demonstrated excellent performance in terms of registration accuracy and robustness.展开更多
This survey reviews feed-forward,point-cloud 3D reconstruction methods from DUSt3R to VGGT and their recent variants.Here,feed-forward primarily refers to predicting dense geometry and,when applicable,camera poses thr...This survey reviews feed-forward,point-cloud 3D reconstruction methods from DUSt3R to VGGT and their recent variants.Here,feed-forward primarily refers to predicting dense geometry and,when applicable,camera poses through learned network inference,without relying on classical per-scene SfM+MVS optimization as the main inference mechanism.We first formalize the reconstruction task in pose-aware and pose-free settings,and contrast feed-forward point-map regression with classical Structure-from-Motion and Multi-View Stereo pipelines.Building on this,we organize existing methods into three stages:early pairwise models typified by DUSt3R,DUSt3R-style extensions that enhance multi-view consistency,streaming,efficiency,and dynamic-scene handling,and large unified transformers such as VGGT that process tens to hundreds of views jointly,while noting differences in their inference paradigms.We analyze these models along shared axes,including scene representation,correspondence reasoning,pose regression,fusion strategies,and the role of large-scale training data.We summarize widely used 3D datasets and evaluation metrics,and provide a case study on the DTU benchmark for multi-view depth and point map estimation,highlighting accuracy-efficiency trade-offs between optimization-based and feed-forward approaches.Finally,we discuss open challenges in data scarcity,sparse-view reconstruction,non-Lambertian structures,dynamic scenes,long-context processing,and resource-efficient deployment,and outline future directions that combine feed-forward architectures with differentiable rendering,generative priors,and safety mechanisms to enable scalable and trustworthy 3D reconstruction systems.展开更多
Wing-fuselage assembly is a critical process in aircraft manufacturing,and the gap distribution of the wing-fuselage assembly frames directly determines the stress state of the connection interface and the structural ...Wing-fuselage assembly is a critical process in aircraft manufacturing,and the gap distribution of the wing-fuselage assembly frames directly determines the stress state of the connection interface and the structural load-carrying capacity after assembly.However,most existing pose adjustment methods either do not consider the global gap distribution or optimize the gap only based on theoretical computer-aided design(CAD)models,making it difficult to effectively control the actual gap distribution under real manufacturing errors.In addition,traditional positioners adopt a master-slave driving mode,in which passive axes suffer from following errors due to the lack of independent control,thereby limiting the execution accuracy of pose adjustment.To address these problems,this paper proposes a fullyactuated pose adjustment method for wing-fuselage assembly via point-cloud gap optimization.First,point cloud data of the wing-fuselage assembly frames in a unified assembly coordinate system are obtained through laser scanning combined with Enhanced Reference System(ERS)reference point registration.Assembly gaps are then defined along prescribed region-wise assembly directions,and a multi-surface global gap evaluation model is established.Second,a two-stage optimization strategy combining coarse pre-alignment and fine pose adjustment is adopted,in which insertion-depth correction is decoupled from rotational gap-uniformity optimization to solve the optimal target assembly pose.Finally,the optimized pose is converted into multi-axis synchronous motion commands for fully-actuated positioners through fifth-order polynomial trajectory planning and inverse kinematic mapping,and the force-position data of the positioners are effectively monitored throughout the pose adjustment process.The proposed method is compared with a traditional manual assembly method on a wing-fuselage assembly experimental platform.The results show that the proposed method reduces the in-plane gap standard deviations of the upper and right assembly surfaces from 2.19 mm and 1.34 mm to 1.70 mm and 1.04 mm,respectively,significantly improving gap uniformity.Meanwhile,the positioner motion remains continuous and smooth during pose adjustment,and no abnormal force fluctuations occur at the support points.The proposed method provides an integrated and executable solution for highprecision wing-fuselage assembly by linking measured-point-cloud-based global gap optimization,fully-actuated pose adjustment,and real-time force-position monitoring,thereby supporting assembly safety assessment and quality traceability.展开更多
The location where a robot grasps an object is closely related to the task type.For the same object,different user requirements may necessitate different grasping strategies.Visual affordance serves as a reliable sour...The location where a robot grasps an object is closely related to the task type.For the same object,different user requirements may necessitate different grasping strategies.Visual affordance serves as a reliable source of prior knowledge for manipulation.Existing methods learn affordance from images or videos,but planar affordance lacks the spatial information required for 6-degree-of-freedom(6-DoF)manipulation.Furthermore,current approaches are limited to affordances associated with predefined categories and cannot directly infer affordances from user instructions.To address such limitations,we propose a novel task:instruction-driven three-dimensional(3D)object affordance segmentation.To support this research,we introduce an instruction–affordance dataset(IAD),a challenging dataset consisting of 7190 object instances across 20 common object categories,paired with 624 manipulation instructions that specify the corresponding affordances.To evaluate generalization to novel commands,our dataset includes both seen and unseen settings.Building on this,we design an instruction-driven 3D affordance segmentation(IDAS)network,which extracts point cloud features and integrates instruction features layer by layer.Given a user instruction,our method segments suggested manipulation regions on the object’s point cloud,thereby guiding the selection of optimal grasp poses.Experimental results show that our method outperforms other related approaches under both seen and unseen settings,demonstrating generalization ability to diverse user commands and unknown affordances.展开更多
Point Cloud Registration(PCR)is a basic task in computer vision,mobile robotics,and autonomous driving.PCR primarily faces challenges,including insufficient registration performance in low-overlap scenarios and high c...Point Cloud Registration(PCR)is a basic task in computer vision,mobile robotics,and autonomous driving.PCR primarily faces challenges,including insufficient registration performance in low-overlap scenarios and high computational resource consumption in large-scale point cloud scenarios.Most recent PCR methods are transformer-based.Methods like transformers have quadratic computational complexity O(n2d),,leading to rapid increases in computational cost with large-scale point cloud data.To address these problems,an iterative PCR method named Attention and Mamba Based Iterative Registration Network(AMBIR)is proposed,overcoming the shortcomings of the current PCR method on low-overlap and large-scale scenarios.Specifically,an iterative network architecture is introduced that learns overlap experience from prior registration results,thereby enhancing registration performance by leveraging knowledge from the preceding step.Additionally,to convert 3-D point cloud data into linear sequences suitable for the Mamba encoder,the Prior-Informed Co-aligned Serialization is proposed to ensure that points with adjacent indices after serialization are spatial neighbors,thereby improving the efficiency and robustness of the subsequent registration process.Lastly,a Consistency-Aware Mamba Encoder is introduced to leverage its linear computational complexity,making the method more suitable for large-scale point clouds.This method simultaneously overcomes the shortcomings of existing methods,including insufficient registration performance in low-overlap and large-scale point cloud scenarios.It performs well on the 3DMatch dataset,3DLoMatch low-overlap dataset,and KITTI large-scale scene dataset,demonstrating high practical value.展开更多
In view of the large amount of data and dense pixel points in point cloud files,this article proposes a multiple point cloud file encryption algorithm based on principal component analysis(PCA)and fractional Fourier t...In view of the large amount of data and dense pixel points in point cloud files,this article proposes a multiple point cloud file encryption algorithm based on principal component analysis(PCA)and fractional Fourier transform(FrFT).In this method,a point cloud data matrix(PCDM)is generated by extracting the coordinates and color information of the point cloud,then using PCA to reduce the dimension of a sequence of PCDMs,which are spliced and scrambled to produce a feature vector matrix and a dimension-reduced matrix(DRM)for encryption and reconstruction.Then using the hyperchaotic Lorenz system to generate the random phase masks and the orders of the FrFT.These two parameters will be used as keys to encrypt the point cloud feature vector matrix.The simulation results verify that the encryption algorithm can quickly encrypt multiple point cloud files,and the quality of the point cloud files obtained by decryption and reconstruction is good.The algorithm also has a large enough key space and highly sensitive keys,which means it has good security and strong robustness to different attacks.展开更多
In the realms of computer vision and remote sensing,the matching of images and point clouds poses significant challenges due to modality discrepancies.This study introduces a crossmodal consistency network,detector-fr...In the realms of computer vision and remote sensing,the matching of images and point clouds poses significant challenges due to modality discrepancies.This study introduces a crossmodal consistency network,detector-free image and point cloud matching via diffusion-guided crossmodal consistency,2D3D-DiffMatch,leveraging diffusion prior information to enhance feature extraction consistency and alignment across modalities.To strengthen cross-modal consistency in complex scenes,diffusion priors generated by a pre-trained diffusion model are used to guide the feature extraction process toward semantically and geometrically consistent representations.These representations are further refined through a hierarchical fusion process,in which the most consistent diffusion features are adaptively selected using centered kernel alignment(CKA)and integrated with multi-scale backbone features,thereby mitigating the impact of modality gaps.Furthermore,to address feature-space misalignment between images and point clouds,we propose a crossmodal feature consistency loss that adaptively constrains correspondences,separates positive and negative pairs,and optimizes the agreement of positive pairs,enabling high-quality,detector-free matching.Experimental results on the 7Scenes and RGB-D Scenes V2 Datasets demonstrate superior registration recall rates of 81.2%and 61.0%,respectively,outperforming state-of-the-art methods and exhibiting robustness in challenging scenarios.This research advances the collaborative processing of multi-modal data,offering a robust solution for image–point cloud matching in challenging scenarios.展开更多
Evaluating rock mass quality using three-dimensional(3D)point clouds is crucial for discontinuity extraction and is widely applied in various industrial sectors.However,the utilization of this method in geological sur...Evaluating rock mass quality using three-dimensional(3D)point clouds is crucial for discontinuity extraction and is widely applied in various industrial sectors.However,the utilization of this method in geological surveys remains limited.Notable limitations of current research include the scarcity of validation using simple geometric shapes for discontinuity extraction methods,and the lack of studies that target both planar and linear discontinuity.To address these gaps,this study proposes a workflow for identifying discontinuity planes and traces in rock outcrops from photogrammetric 3D modeling,employing the Compass and Facets plugins in the open-source CloudCompare software.Prior to field application,the efficacy of the extraction methods was first evaluated using experimental datasets of a cube and an isosceles triangular prism generated under laboratory-controlled conditions.This validation demonstrated exceptional accuracy,with the dip and dip direction(DDD)of extracted structures consistently within±2°of the actual values.Following this rigorous laboratory validation,this methodology was applied to a more complex natural rock outcrop(Miocene–Pliocene deposits in Japan),demonstrating its applicability in realistic geological settings for identifying structures.The results showed that the dip and dip direction trends of the extracted bedding planes and faults were consistent with field measurements,achieving a time reduction of approximately 40%compared to traditional methods.In conclusion,through strictly controlled initial verification and subsequent successful application to a complex natural setting,this study confirmed that the proposed workflow can effectively and efficiently extract discontinuous geological structures from point clouds.展开更多
摘要The satellite laser ranging (SLR) data quality from the COMPASS was analyzed, and the difference between curve recognition in computer vision and pre-process of SLR data finally proposed a new algorithm for SLR was discussed data based on curve recognition from points cloud is proposed. The results obtained by the new algorithm are 85 % (or even higher) consistent with that of the screen displaying method, furthermore, the new method can process SLR data automatically, which makes it possible to be used in the development of the COMPASS navigation system.
基金supported by the National Natural Science Foundation of China(Grant Nos.52304139,52325403)the CCTEG Coal Mining Research Institute funding(Grant No.KCYJY-2024-MS-10).
摘要3D laser scanning technology is widely used in underground openings for high-precision,rapid,and nondestructive structural evaluations.Segmenting large 3D point cloud datasets,particularly in coal mine roadways with multi-scale targets,remains challenging.This paper proposes an enhanced segmentation method integrating improved PointNet++with a coverage-voted strategy.The coverage-voted strategy reduces data while preserving multi-scale target topology.The segmentation is achieved using an enhanced PointNet++algorithm with a normalization preprocessing head,resulting in a 94%accuracy for common supporting components.Ablation experiments show that the preprocessing head and coverage strategies increase segmentation accuracy by 20%and 2%,respectively,and improve Intersection over Union(IoU)for bearing plate segmentation by 58%and 20%.The accuracy of the current pretraining segmentation model may be affected by variations in surface support components,but it can be readily enhanced through re-optimization with additional labeled point cloud data.This proposed method,combined with a previously developed machine learning model that links rock bolt load and the deformation field of its bearing plate,provides a robust technique for simultaneously measuring the load of multiple rock bolts in a single laser scan.
基金the support from the National Natural Science Foundation of China(No.52472384)the Fundamental Research Funds for the Central Universities,China(No.G2024KY0615)+1 种基金sponsored by the foundations of the National Key Laboratory of Unmanned Aerial Vehicle Technology in Northwestern Polytechnical University,China(No.WR202411-2)the National Key Laboratory of Aircraft Configuration Design,China(No.JBGS-2024-01)。
摘要The precision of shape representation and the dimensionality of the design space significantly influence the cost and outcomes of aerodynamic optimization.The design space can be represented more compactly by maintaining geometric precision while reducing dimensions,hence enhancing the cost-effectiveness of the optimization process.This research presents a new point cloud Autoencoder Based on Uncertainty Feature Enhancement(AE-BUFE)architecture,designed to attain efficient and precise generalized representations of 3D aircraft through uncertainty analysis of the deformation relationships among surface grid points.The deep learning architecture consists of two components:the uncertainty index-based feature enhancement module and the point cloud autoencoder module.It learns the shape features of the point cloud geometric representation to establish a low-dimensional latent space.To assess and evaluate the efficiency of the method,a comparison was conducted with the prevailing point cloud autoencoder architecture and the proper orthogonal decomposition linear dimensionality reduction method under conditions of complex shape deformation.The results show that the new architecture significantly improves the extraction effect of the low-dimensional latent space.Then,this paper developed the surrogate-based optimization framework based on the AE-BUFE parameterization method and completed a multi-objective aerodynamic optimization design for a wide-speed-range vehicle considering volume and moment constraints.While ensuring the take-off and landing performance,the aerodynamic performance is improved under transonic and hypersonic conditions,which verifies the efficiency and engineering practicability of this method.
基金supported in part by the National Key Research and Development Program of Chinaunder(Grant 2021YFB3101100)in part by the National Natural Science Foundation of Chinaunder(Grant 42461057),(Grant 62272123),and(Grant 42371470)+1 种基金in part by the Fundamental Research Program of Shanxi Province under(Grant 202303021212164)in part by the Postgraduate Education Innovation Program of Shanxi Province under(Grant 2024KY474).
摘要Recently,large-scale deep learning models have been increasingly adopted for point cloud classification.However,thesemethods typically require collecting extensive datasets frommultiple clients,which may lead to privacy leaks.Federated learning provides an effective solution to data leakage by eliminating the need for data transmission,relying instead on the exchange of model parameters.However,the uneven distribution of client data can still affect the model’s ability to generalize effectively.To address these challenges,we propose a new framework for point cloud classification called Federated Dynamic Aggregation Selection Strategy-based Multi-Receptive Field Fusion Classification Framework(FDASS-MRFCF).Specifically,we tackle these challenges with two key innovations:(1)During the client local training phase,we propose a Multi-Receptive Field Fusion Classification Model(MRFCM),which captures local and global structures in point cloud data through dynamic convolution and multi-scale feature fusion,enhancing the robustness of point cloud classification.(2)In the server aggregation phase,we introduce a Federated Dynamic Aggregation Selection Strategy(FDASS),which employs a hybrid strategy to average client model parameters,skip aggregation,or reallocate local models to different clients,thereby balancing global consistency and local diversity.We evaluate our framework using the ModelNet40 and ShapeNetPart benchmarks,demonstrating its effectiveness.The proposed method is expected to significantly advance the field of point cloud classification in a secure environment.
基金the financialsupport provided by the Technology Innovation Center for Land Engineering and Human Settlements by Shaanxi Land Engineering Construction Group Co.,Ltd(Grant No.201912131-C2)Young Scientists Project under the Natural Science Basic Research Program of Shaanxi Province(Grant No.2025JC-YBQN-382)Xi'an Key Industrial Chain Technology Breakthrough General Project(Grant No.2024JHCLYB-0128).
摘要In recent years,drought-induced soil cracking has become increasingly prevalent,posing significant challenges to geotechnical engineering applications due to the associated hazards.Previous studies have investigated soil cracking behaviors and mechanisms via multi-scale methods.However,research on the spatial distribution characteristics of soil cracks and their depth evolution mechanisms is still rare.This paper analyzes the three-dimensional(3D)deformation of fiveclayey materials.In this study,3D surface elevation data of soil samples were acquired using a Gocator 3110 structured-light sensor.The reconstruction process involved point cloud acquisition and rasterization,Delaunay triangulation meshing,and calculation of crack depth.Different cracking behaviors are investigated,such as cracking in opening mode and shearing mode,coalescence and bifurcation of cracks,etc.The mineralogical analysis is carried out to reveal different cracking behaviors of clays.It is concluded that the crack depth of montmorillonite is larger than that of kaolin,reflectingits strong shrinkage characteristic,which makes the crack evolve deeper than the other materials.This technique enables multi-dimensional observation of soil crack propagation.The mineralogical analysis is carried out to reveal the different cracking behaviors of clays.It is concluded that microstructure and mineralogy are key factors influencingcracking behavior.Understanding soil cracking behaviors is significantfor engineering protection,geohazard prevention,and geological monitoring.
基金The National Natural Science Foundation of China(No.52338011)the Start-up Research Fund of Southeast University(No.RF1028624058)+1 种基金the Southeast University Interdisciplinary Research Program for Young Scholarsthe National Key Research and Development Program of China(No.2024YFC3014103).
摘要The virtual preassembly of super-high steel bridge towers faces a challenge in the efficient and precise extraction of complex cross-sectional features.Factors such as fabrication errors,gravity-induced deformations,and temperature fluctuations can compromise the accuracy of contour extraction.To address these limitations,an improved Alpha-shape-based point cloud contour extraction method is proposed.The proposed approach uses a hierarchical strategy to process three-dimensional laser scanning point clouds.The processed data are then subjected to curvatureadaptive voxel filtering to reduce acquisition noise.In addition,an enhanced iterative closest point(ICP)variant with correspondence validation accurately aligns the discrete point cloud segments.The proposed curvature-responsive Alpha-shape framework enables multiscale contour delineation through topology-adaptive threshold modulation,which resolves boundary ambiguities in geometrically complex cross-sections.The method was experimentally validated using field-acquired measurement datasets from the Zhangjinggao Yangtze River Bridge tower segments,confirming its capability to reconstruct noncanonical cross-sectional geometries.Three contour extraction methods,including Poisson reconstruction,the conventional Alpha-shape algorithm,and random sample consensus with ICP(RANSAC-ICP),were compared to evaluate the performance of the proposed Alpha-shape algorithm.The results demonstrate that the proposed method achieves superior contour extraction accuracy and data reduction efficiency,highlighting its effectiveness in contour extraction tasks.
摘要Three-dimensional(3D)point cloud semantic segmentation is a core task in indoor scene understanding,providing detailed semantic information about spatial structures and object categories in indoor environments.Although methods based on deep learning have made steady progress in recent years,accurately segmenting complex indoor scenes remains challenging due to the unordered nature of point clouds and variations across large scales.Most existing networks have limited capability for multi-scale feature aggregation and struggle to balance local geometric details with global semantic context.These issues are further exacerbated by hierarchical downsampling,which often leads to the loss of fine-grained structural information.Moreover,feature interaction restricted to local neighborhoods may limit the capture of non-local semantic dependencies in complex indoor scenes.To address these limitations,we propose PointNMSA(PointNeXt with Non-local Multi-Scale Aggregation),an improved semantic segmentation network built upon the PointNeXt backbone.A Multi-Scale Feature Enhancement(MSFE)module is introduced in the decoding stage to fuse features from different encoding levels,and further refines the fused features to produce more stable multi-scale representations,which preserves geometric details across scales.In addition,a Convolution-Attention Mixing(CA-Mix)module is designed to jointly integrate local spatial structures and non-local contextual dependencies via dual-stream aggregation and multi-dimensional attention fusion,thereby enabling more discriminative feature representations.Experiments on the Stanford Large-Scale 3D Indoor Spaces(S3DIS)benchmark demonstrate the effectiveness of PointNMSA.On the Area 5 test split,PointNMSA achieves a mean intersection over union(mIoU)of 65.10%,outperforming the PointNeXt baseline by 1.59%,while introducing only a modest increase in computational cost(latency from 42.24 to 45.18 ms and parameters from 3.16 to 8.67M).Despite the noticeable growth in parameter count,the increase in inference latency remains relatively limited,indicating a favorable trade-off between segmentation accuracy and computational efficiency.Additional cross-dataset experiments on ScanNet further verify that PointNMSA maintains stable gains under different indoor scene distributions.Such performance gains suggest that PointNMSA provides a more robust and generalizable solution for semantic segmentation in large-scale indoor environments with complex structural layouts.
基金supported by the National Natural Science Foundation of China(Grant Nos.U2441211,U23B6009,and 2022YFB4300200).
摘要Data-driven approaches have shown great advantage in rapidly and accurately predicting pressure coefficient distributions,which is of crucial importance to efficient aircraft design.Nevertheless,most data-driven approaches still encounter limitations in characterizing diverse aerodynamic configurations and adapting to varying grid densities,which have hindered their engineering applicability.In response to these challenges,this work adopts point clouds,a specific type of geometric data structure that is inherently suitable for uniformly characterizing diverse 2D/3D geometric shapes as the input for deep learning-based prediction of pressure coefficient distribution.By augmenting the dimensions of point cloud coordinates for local feature enhancement and utilizing the symmetric function“max pooling”to extract global features,the proposed aerodynamic model establishes the mapping between point cloud coordinates and pressure coefficients.Basic aerodynamic configurations like airfoils and wings are employed as test cases,the results demonstrate that the proposed model achieves both high accuracy and robust generalizability across variable geometries.For class-shape transformation-perturbed airfoils,the prediction error can be reduced to one-third of that of the conventional parameterization-based model.For airfoils selected in the University of Illinois Urbana-Champaign airfoil dataset,among which airfoil profiles are widely distributed,the average error of the proposed approach remains approximately 1.5%,whereas the parameterization-based model may fail.For wings,the prediction error still stays below 2.5%.Finally,the model exhibits strong robustness and generalizability across different point cloud densities.In conclusion,this work makes a breakthrough in predicting pressure coefficient distribution for variable geometric configurations,establishing the foundational framework for designing a large model capable of predicting distributed aerodynamic loads in aerospace applications.
基金supported by the National Natural Science Foundation of China(Grant No.42507210)the Fundamental Research Funds for the Central Universities(Grant No.2025XJSB01).
摘要Accurate extraction of rock mass discontinuity parameters is crucial for stability assessment and engineering safety.High-resolution remote sensing facilitates automated extraction,but its effectiveness relies heavily on precise normal estimation to ensure geometric reliability.Conventional methods struggle to preserve sharp features such as edges and corners,thereby reducing accuracy.To address this,we propose a normal estimation method based on local geometric adjustment that enhances feature extraction while maintaining sharp geometries.The approach consists of four steps:(1)classifying points,(2)applying normal and axial projections,(3)fitting segmentation lines via least squares,and(4)refining normals by optimizing local neighborhoods.The proposed method was evaluated on computer-aided design(CAD)models,real objects,and rock mass point clouds,and benchmarked against eight representative algorithms,including principal component analysis(PCA),2-Jet PCA,Voronoi-based PCA,PCPNet,neural gradient function(NeuralGF),low rank representation(LRR),normal estimation via shifted neighborhood(NSN)and pair consistency voting(PCV).Experimental results demonstrate that our method achieves superior accuracy and efficiency,significantly improving structural plane extraction and ensuring better preservation of sharp geometric features.
基金supported by the National Natural Science Foundation of China(Nos.62306049,92471207 and W2421089)the General Program of Chongqing Natural Science Foundation(No.CSTB2023NSCQMSX0665)the Fundamental Research Funds for the Central Universities(No.2024CDJXY008).
摘要3D single object tracking(SOT)based on point clouds is a fundamental task for environmental perception in autonomous driving and dynamic scene understanding in robotics.Recent technological advancements in this field have significantly bolstered the environmental interaction capabilities of intelligent systems.This field faces persistent challenges,including feature degradation induced by point cloud sparsity,representation drift caused by non-rigid deformation,and occlusion in complex scenarios.Traditional appearance matching methods,particularly those relying on Siamese networks,are severely constrained by point cloud characteristics,often failing under rapid motions or structural ambiguities among similar objects.In response,the research paradigm has progressively evolved toward motion-centric modeling approaches.These emerging frameworks utilize spatio-temporal joint modeling and geometric shape completion to attain notable performance gains.Furthermore,the incorporation of attention mechanisms and State Space Model(SSM)has enabled more effective multi-scale spatio-temporal feature association,which is particularly beneficial for long-term tracking scenarios.To the best of our knowledge,this is the first comprehensive survey dedicated to 3D single object tracking in point clouds.We provide a detailed analysis of current tracking methods,scrutinizing their limitations regarding multi-object interference and analyzing the trade-off between accuracy and computational efficiency.Finally,we discuss potential future directions,including the development of lightweight models for edge deployment and the integration of cross-modal fusion strategies.
摘要To support the process of grasping objects on a tabletop for the blind or robotic arm,it is necessary to address fundamental computer vision tasks,such as detecting,recognizing,and locating objects in space,and determining the position of the grasping information.These results can then be used to guide the visually impaired or to execute grasping tasks with a robotic arm.In this paper,we collected,annotated,and published the benchmark TQUGraspingObject dataset for testing,validation,and evaluation of deep learning(DL)models for detecting,recognizing,and localizing grasping objects in 2D and 3D space,especially 3D point cloud data.Our dataset is collected in a shared room,with common everyday objects placed on the tabletop in jumbled positions by Intel RealSense D435(IR-D435).This dataset includes more than 63k RGB-D pairs and related data such as normalized 3D object point cloud,3D object point cloud segmented,coordinate system normalizationmatrix,3D object point cloud normalized,and hand pose for grasping each object.At the same time,we also conducted experiments on fourDL networks with the best performance:SSD-MobileNetV3,ResNet50-Transformer,ResNet101-Transformer,and YOLOv12.The results present that YOLOv12 has the most suitable results in detecting and recognizing objects in images.All data,annotations,toolkit,source code,point cloud data,and results are publicly available on our project website:http://gffzz188fe103f8f1460asfccnxf5ob05c6bw0.ffgz.tsg.suse.edu.cn/HuaTThanhIT2327Tqu/datasetv2.
摘要In the field of aircraft design and maintenance,with the innovation of cabin cable three-dimensional(3D)scanning and sensor technology,high-precision cabin point cloud data has become the key to improving the accuracy of cabin navigation and building a realistic virtual reality environment.In the face of largescale point cloud data,how to efficiently and uniformly construct a realistic virtual reality environment has become a challenge.In this paper,we propose a new low-parametric point cloud upsampling network(LPNet),which is based on the no-learn model to learn the complementary geometric knowledge between point clouds based on some simple data transformations,to efficiently retain the geometric properties of point clouds,and then input the results into the up-sampling module,and simply insert a few layers of multilayer perceptron(MLP)to efficiently generate high-resolution point clouds.It is able to efficiently generate high-resolution point clouds,showing great flexibility and realizing the efficient use of computational resources.
摘要Machine vision-based detection methods have been widely applied in the detection of aircraft skin damage.During drone inspection processes,a key step is to spatially locate high-resolution detailed images of aircraft skin from multiple angles onto a threedimensional point cloud model of the aircraft.This relies on the rigid registration of image center position coordinate point cloud with the aircraft 3D point cloud.To address the issues of low accuracy and poor robustness encountered by existing registration algorithms when dealing with heterogeneous point clouds with significant differences in density and low overlap,this paper presents a novel cross-source point cloud registration network.The network integrates multi-scale information from the point cloud and employs an attention mechanism to identify representative overlapping points.First,the network achieves initial correspondences using the multi-scale geometric features and positional information of the point cloud.Then,an overlapping feature guidance module predicts the overlapping score of the point cloud.By utilizing information interaction through the attention mechanism,the network combines point overlapping scores with fused features to filter out representative overlapping points,achieving precise correspondences in the point cloud.The network employs weighted singular value decomposition(SVD)to estimate two sets of transformation matrices,yielding the relative pose parameters of the point cloud.Experiments were conducted in an unsupervised manner.The experimental results on the ModelNet40 dataset and the aero object dataset aircraft measurement data showed that,compared to other existing traditional and learning-based methods,this approach demonstrated excellent performance in terms of registration accuracy and robustness.
基金Supported by the National Natural Science Foundation of China(Grant Nos.62595774,62302297,72192821,62272447,62472285,62472282)the Fundamental Research Funds for the Central Universities(Grant Nos.YG 2023 QNB 17,YG 2024 QNA 44)+1 种基金the National Key R&D Program of China(Grant No.2024 YFE 0115500)the Beijing Natural Science Foundation(Grant No.L 222117).
摘要This survey reviews feed-forward,point-cloud 3D reconstruction methods from DUSt3R to VGGT and their recent variants.Here,feed-forward primarily refers to predicting dense geometry and,when applicable,camera poses through learned network inference,without relying on classical per-scene SfM+MVS optimization as the main inference mechanism.We first formalize the reconstruction task in pose-aware and pose-free settings,and contrast feed-forward point-map regression with classical Structure-from-Motion and Multi-View Stereo pipelines.Building on this,we organize existing methods into three stages:early pairwise models typified by DUSt3R,DUSt3R-style extensions that enhance multi-view consistency,streaming,efficiency,and dynamic-scene handling,and large unified transformers such as VGGT that process tens to hundreds of views jointly,while noting differences in their inference paradigms.We analyze these models along shared axes,including scene representation,correspondence reasoning,pose regression,fusion strategies,and the role of large-scale training data.We summarize widely used 3D datasets and evaluation metrics,and provide a case study on the DTU benchmark for multi-view depth and point map estimation,highlighting accuracy-efficiency trade-offs between optimization-based and feed-forward approaches.Finally,we discuss open challenges in data scarcity,sparse-view reconstruction,non-Lambertian structures,dynamic scenes,long-context processing,and resource-efficient deployment,and outline future directions that combine feed-forward architectures with differentiable rendering,generative priors,and safety mechanisms to enable scalable and trustworthy 3D reconstruction systems.
基金supported by the Young Elite Scientists Sponsorship Program by CAST(Grant No.YESS20240732)in part by the National Science Foundation for Young Scientists of China(Grant No.52405581)the Innovation Foundation of National Commercial Aircraft Manufacturing Engineering Technology Research Center(Grant No.COMAC-SFGS-2026-349)。
摘要Wing-fuselage assembly is a critical process in aircraft manufacturing,and the gap distribution of the wing-fuselage assembly frames directly determines the stress state of the connection interface and the structural load-carrying capacity after assembly.However,most existing pose adjustment methods either do not consider the global gap distribution or optimize the gap only based on theoretical computer-aided design(CAD)models,making it difficult to effectively control the actual gap distribution under real manufacturing errors.In addition,traditional positioners adopt a master-slave driving mode,in which passive axes suffer from following errors due to the lack of independent control,thereby limiting the execution accuracy of pose adjustment.To address these problems,this paper proposes a fullyactuated pose adjustment method for wing-fuselage assembly via point-cloud gap optimization.First,point cloud data of the wing-fuselage assembly frames in a unified assembly coordinate system are obtained through laser scanning combined with Enhanced Reference System(ERS)reference point registration.Assembly gaps are then defined along prescribed region-wise assembly directions,and a multi-surface global gap evaluation model is established.Second,a two-stage optimization strategy combining coarse pre-alignment and fine pose adjustment is adopted,in which insertion-depth correction is decoupled from rotational gap-uniformity optimization to solve the optimal target assembly pose.Finally,the optimized pose is converted into multi-axis synchronous motion commands for fully-actuated positioners through fifth-order polynomial trajectory planning and inverse kinematic mapping,and the force-position data of the positioners are effectively monitored throughout the pose adjustment process.The proposed method is compared with a traditional manual assembly method on a wing-fuselage assembly experimental platform.The results show that the proposed method reduces the in-plane gap standard deviations of the upper and right assembly surfaces from 2.19 mm and 1.34 mm to 1.70 mm and 1.04 mm,respectively,significantly improving gap uniformity.Meanwhile,the positioner motion remains continuous and smooth during pose adjustment,and no abnormal force fluctuations occur at the support points.The proposed method provides an integrated and executable solution for highprecision wing-fuselage assembly by linking measured-point-cloud-based global gap optimization,fully-actuated pose adjustment,and real-time force-position monitoring,thereby supporting assembly safety assessment and quality traceability.
基金supported by the National Natural Science Foundation of China(No.61973192).
摘要The location where a robot grasps an object is closely related to the task type.For the same object,different user requirements may necessitate different grasping strategies.Visual affordance serves as a reliable source of prior knowledge for manipulation.Existing methods learn affordance from images or videos,but planar affordance lacks the spatial information required for 6-degree-of-freedom(6-DoF)manipulation.Furthermore,current approaches are limited to affordances associated with predefined categories and cannot directly infer affordances from user instructions.To address such limitations,we propose a novel task:instruction-driven three-dimensional(3D)object affordance segmentation.To support this research,we introduce an instruction–affordance dataset(IAD),a challenging dataset consisting of 7190 object instances across 20 common object categories,paired with 624 manipulation instructions that specify the corresponding affordances.To evaluate generalization to novel commands,our dataset includes both seen and unseen settings.Building on this,we design an instruction-driven 3D affordance segmentation(IDAS)network,which extracts point cloud features and integrates instruction features layer by layer.Given a user instruction,our method segments suggested manipulation regions on the object’s point cloud,thereby guiding the selection of optimal grasp poses.Experimental results show that our method outperforms other related approaches under both seen and unseen settings,demonstrating generalization ability to diverse user commands and unknown affordances.
基金supported by the National Natural Science Foundation of China(Grant No.12141304).
摘要Point Cloud Registration(PCR)is a basic task in computer vision,mobile robotics,and autonomous driving.PCR primarily faces challenges,including insufficient registration performance in low-overlap scenarios and high computational resource consumption in large-scale point cloud scenarios.Most recent PCR methods are transformer-based.Methods like transformers have quadratic computational complexity O(n2d),,leading to rapid increases in computational cost with large-scale point cloud data.To address these problems,an iterative PCR method named Attention and Mamba Based Iterative Registration Network(AMBIR)is proposed,overcoming the shortcomings of the current PCR method on low-overlap and large-scale scenarios.Specifically,an iterative network architecture is introduced that learns overlap experience from prior registration results,thereby enhancing registration performance by leveraging knowledge from the preceding step.Additionally,to convert 3-D point cloud data into linear sequences suitable for the Mamba encoder,the Prior-Informed Co-aligned Serialization is proposed to ensure that points with adjacent indices after serialization are spatial neighbors,thereby improving the efficiency and robustness of the subsequent registration process.Lastly,a Consistency-Aware Mamba Encoder is introduced to leverage its linear computational complexity,making the method more suitable for large-scale point clouds.This method simultaneously overcomes the shortcomings of existing methods,including insufficient registration performance in low-overlap and large-scale point cloud scenarios.It performs well on the 3DMatch dataset,3DLoMatch low-overlap dataset,and KITTI large-scale scene dataset,demonstrating high practical value.
基金supported by National Natural Science Foundation of China(61771155).
摘要In view of the large amount of data and dense pixel points in point cloud files,this article proposes a multiple point cloud file encryption algorithm based on principal component analysis(PCA)and fractional Fourier transform(FrFT).In this method,a point cloud data matrix(PCDM)is generated by extracting the coordinates and color information of the point cloud,then using PCA to reduce the dimension of a sequence of PCDMs,which are spliced and scrambled to produce a feature vector matrix and a dimension-reduced matrix(DRM)for encryption and reconstruction.Then using the hyperchaotic Lorenz system to generate the random phase masks and the orders of the FrFT.These two parameters will be used as keys to encrypt the point cloud feature vector matrix.The simulation results verify that the encryption algorithm can quickly encrypt multiple point cloud files,and the quality of the point cloud files obtained by decryption and reconstruction is good.The algorithm also has a large enough key space and highly sensitive keys,which means it has good security and strong robustness to different attacks.
基金supported by the National Natural Science Foundation of China(No.62561034)Yunnan Provincial International Joint Laboratory of Agricultural Remote Sensing and Digital Technology(No.202503AP140020)+3 种基金“Xingdian”Talent Support Program Project(No.KKRD202221041)Yunnan Fundamental Research Projects(No.202301AT070463)Yunnan International Joint Laboratory for Integrated Sky-Ground Intelligent Monitoring of Mountain Hazards(No.202403AP140002)Yunnan Plateau Remote Sensing Innovation Team(No.202505AS350001).
摘要In the realms of computer vision and remote sensing,the matching of images and point clouds poses significant challenges due to modality discrepancies.This study introduces a crossmodal consistency network,detector-free image and point cloud matching via diffusion-guided crossmodal consistency,2D3D-DiffMatch,leveraging diffusion prior information to enhance feature extraction consistency and alignment across modalities.To strengthen cross-modal consistency in complex scenes,diffusion priors generated by a pre-trained diffusion model are used to guide the feature extraction process toward semantically and geometrically consistent representations.These representations are further refined through a hierarchical fusion process,in which the most consistent diffusion features are adaptively selected using centered kernel alignment(CKA)and integrated with multi-scale backbone features,thereby mitigating the impact of modality gaps.Furthermore,to address feature-space misalignment between images and point clouds,we propose a crossmodal feature consistency loss that adaptively constrains correspondences,separates positive and negative pairs,and optimizes the agreement of positive pairs,enabling high-quality,detector-free matching.Experimental results on the 7Scenes and RGB-D Scenes V2 Datasets demonstrate superior registration recall rates of 81.2%and 61.0%,respectively,outperforming state-of-the-art methods and exhibiting robustness in challenging scenarios.This research advances the collaborative processing of multi-modal data,offering a robust solution for image–point cloud matching in challenging scenarios.
摘要Evaluating rock mass quality using three-dimensional(3D)point clouds is crucial for discontinuity extraction and is widely applied in various industrial sectors.However,the utilization of this method in geological surveys remains limited.Notable limitations of current research include the scarcity of validation using simple geometric shapes for discontinuity extraction methods,and the lack of studies that target both planar and linear discontinuity.To address these gaps,this study proposes a workflow for identifying discontinuity planes and traces in rock outcrops from photogrammetric 3D modeling,employing the Compass and Facets plugins in the open-source CloudCompare software.Prior to field application,the efficacy of the extraction methods was first evaluated using experimental datasets of a cube and an isosceles triangular prism generated under laboratory-controlled conditions.This validation demonstrated exceptional accuracy,with the dip and dip direction(DDD)of extracted structures consistently within±2°of the actual values.Following this rigorous laboratory validation,this methodology was applied to a more complex natural rock outcrop(Miocene–Pliocene deposits in Japan),demonstrating its applicability in realistic geological settings for identifying structures.The results showed that the dip and dip direction trends of the extracted bedding planes and faults were consistent with field measurements,achieving a time reduction of approximately 40%compared to traditional methods.In conclusion,through strictly controlled initial verification and subsequent successful application to a complex natural setting,this study confirmed that the proposed workflow can effectively and efficiently extract discontinuous geological structures from point clouds.