Three-dimensional(3D)point cloud semantic segmentation is a core task in indoor scene understanding,providing detailed semantic information about spatial structures and object categories in indoor environments.Althoug...Three-dimensional(3D)point cloud semantic segmentation is a core task in indoor scene understanding,providing detailed semantic information about spatial structures and object categories in indoor environments.Although methods based on deep learning have made steady progress in recent years,accurately segmenting complex indoor scenes remains challenging due to the unordered nature of point clouds and variations across large scales.Most existing networks have limited capability for multi-scale feature aggregation and struggle to balance local geometric details with global semantic context.These issues are further exacerbated by hierarchical downsampling,which often leads to the loss of fine-grained structural information.Moreover,feature interaction restricted to local neighborhoods may limit the capture of non-local semantic dependencies in complex indoor scenes.To address these limitations,we propose PointNMSA(PointNeXt with Non-local Multi-Scale Aggregation),an improved semantic segmentation network built upon the PointNeXt backbone.A Multi-Scale Feature Enhancement(MSFE)module is introduced in the decoding stage to fuse features from different encoding levels,and further refines the fused features to produce more stable multi-scale representations,which preserves geometric details across scales.In addition,a Convolution-Attention Mixing(CA-Mix)module is designed to jointly integrate local spatial structures and non-local contextual dependencies via dual-stream aggregation and multi-dimensional attention fusion,thereby enabling more discriminative feature representations.Experiments on the Stanford Large-Scale 3D Indoor Spaces(S3DIS)benchmark demonstrate the effectiveness of PointNMSA.On the Area 5 test split,PointNMSA achieves a mean intersection over union(mIoU)of 65.10%,outperforming the PointNeXt baseline by 1.59%,while introducing only a modest increase in computational cost(latency from 42.24 to 45.18 ms and parameters from 3.16 to 8.67M).Despite the noticeable growth in parameter count,the increase in inference latency remains relatively limited,indicating a favorable trade-off between segmentation accuracy and computational efficiency.Additional cross-dataset experiments on ScanNet further verify that PointNMSA maintains stable gains under different indoor scene distributions.Such performance gains suggest that PointNMSA provides a more robust and generalizable solution for semantic segmentation in large-scale indoor environments with complex structural layouts.展开更多
Accurately extracting plant point clouds from complex agricultural environments is essential for high-throughput phenotyping in smart farming.However,existing methods face significant challenges when processing large-...Accurately extracting plant point clouds from complex agricultural environments is essential for high-throughput phenotyping in smart farming.However,existing methods face significant challenges when processing large-scale agricultural point clouds owing to high noise levels,dense spatial distribution,and blurred structural boundaries between plant and non-plant regions.To address these issues,this study proposes PlaneSegNet,a voxel-based semantic segmentation network that incorporates an innovative plane attention module.This module aggregates projection features from the XZ and YZ planes,enhancing the model's ability to detect vertical geometric variations and thereby improving segmentation performance in boundary regions.Extensive experiments across representative agricultural scenarios at multiple scales,including open-field populations,greenhouse cultivation environments,and large-scale rural landscapes,demonstrate that PlaneSegNet significantly outperforms traditional geometry-based approaches and deep-learning models in plant and non-plant separation.By directly generating high-quality plant-only point clouds,PlaneSegNet significantly reduces reliance on manual pre-processing,offering a practical and generalisable solution for automated plant extraction across a wide range of agricultural applications.展开更多
摘要Three-dimensional(3D)point cloud semantic segmentation is a core task in indoor scene understanding,providing detailed semantic information about spatial structures and object categories in indoor environments.Although methods based on deep learning have made steady progress in recent years,accurately segmenting complex indoor scenes remains challenging due to the unordered nature of point clouds and variations across large scales.Most existing networks have limited capability for multi-scale feature aggregation and struggle to balance local geometric details with global semantic context.These issues are further exacerbated by hierarchical downsampling,which often leads to the loss of fine-grained structural information.Moreover,feature interaction restricted to local neighborhoods may limit the capture of non-local semantic dependencies in complex indoor scenes.To address these limitations,we propose PointNMSA(PointNeXt with Non-local Multi-Scale Aggregation),an improved semantic segmentation network built upon the PointNeXt backbone.A Multi-Scale Feature Enhancement(MSFE)module is introduced in the decoding stage to fuse features from different encoding levels,and further refines the fused features to produce more stable multi-scale representations,which preserves geometric details across scales.In addition,a Convolution-Attention Mixing(CA-Mix)module is designed to jointly integrate local spatial structures and non-local contextual dependencies via dual-stream aggregation and multi-dimensional attention fusion,thereby enabling more discriminative feature representations.Experiments on the Stanford Large-Scale 3D Indoor Spaces(S3DIS)benchmark demonstrate the effectiveness of PointNMSA.On the Area 5 test split,PointNMSA achieves a mean intersection over union(mIoU)of 65.10%,outperforming the PointNeXt baseline by 1.59%,while introducing only a modest increase in computational cost(latency from 42.24 to 45.18 ms and parameters from 3.16 to 8.67M).Despite the noticeable growth in parameter count,the increase in inference latency remains relatively limited,indicating a favorable trade-off between segmentation accuracy and computational efficiency.Additional cross-dataset experiments on ScanNet further verify that PointNMSA maintains stable gains under different indoor scene distributions.Such performance gains suggest that PointNMSA provides a more robust and generalizable solution for semantic segmentation in large-scale indoor environments with complex structural layouts.
基金supported by Funds from the National key research and development program[2022YFD2002303-01].
摘要Accurately extracting plant point clouds from complex agricultural environments is essential for high-throughput phenotyping in smart farming.However,existing methods face significant challenges when processing large-scale agricultural point clouds owing to high noise levels,dense spatial distribution,and blurred structural boundaries between plant and non-plant regions.To address these issues,this study proposes PlaneSegNet,a voxel-based semantic segmentation network that incorporates an innovative plane attention module.This module aggregates projection features from the XZ and YZ planes,enhancing the model's ability to detect vertical geometric variations and thereby improving segmentation performance in boundary regions.Extensive experiments across representative agricultural scenarios at multiple scales,including open-field populations,greenhouse cultivation environments,and large-scale rural landscapes,demonstrate that PlaneSegNet significantly outperforms traditional geometry-based approaches and deep-learning models in plant and non-plant separation.By directly generating high-quality plant-only point clouds,PlaneSegNet significantly reduces reliance on manual pre-processing,offering a practical and generalisable solution for automated plant extraction across a wide range of agricultural applications.