Humans can perceive our complex world through multi-sensory fusion.Under limited visual conditions,people can sense a variety of tactile signals to identify objects accurately and rapidly.However,replicating this uniq...Humans can perceive our complex world through multi-sensory fusion.Under limited visual conditions,people can sense a variety of tactile signals to identify objects accurately and rapidly.However,replicating this unique capability in robots remains a significant challenge.Here,we present a new form of ultralight multifunctional tactile nano-layered carbon aerogel sensor that provides pressure,temperature,material recognition and 3D location capabilities,which is combined with multimodal supervised learning algorithms for object recognition.The sensor exhibits human-like pressure(0.04–100 kPa)and temperature(21.5–66.2℃)detection,millisecond response times(11 ms),a pressure sensitivity of 92.22 kPa−1and triboelectric durability of over 6000 cycles.The devised algorithm has universality and can accommodate a range of application scenarios.The tactile system can identify common foods in a kitchen scene with 94.63%accuracy and explore the topographic and geomorphic features of a Mars scene with 100%accuracy.This sensing approach empowers robots with versatile tactile perception to advance future society toward heightened sensing,recognition and intelligence.展开更多
To address crop depredation by intelligent species(e.t,macaques)and the habituation from traditional methods,this study proposes an intelligent,closed-loop,adaptive laser deterrence system.A core contribution is an ef...To address crop depredation by intelligent species(e.t,macaques)and the habituation from traditional methods,this study proposes an intelligent,closed-loop,adaptive laser deterrence system.A core contribution is an efficient multi-stage Semi-Supervised Learning(SSL)and incremental fine-tuning(IFT)framework,which reduced manual annotation by~60%and training time by~68%.This framework was benchmarked against YOLOv8n,v10n,and v11n.Our analysis revealed that YOLOv12n’s high Signal-to-Noise Ratio(SNR)(47.1%retention)pseudo-labels made it the onlymodel to gain performance(+0.010mAP)fromSSL,allowing it to overtake competitors.Subsequently,in the IFT stress test,YOLOv12n proved most robust(a minimal−0.019 mAP decline),whereas YOLOv10n suffered catastrophic failure(−0.233mAP),highlighting its incompatibility with IFT.Thefinalmodel achieved high performance(mAP@0.5 of 0.947 for macaques,0.946 for laser spots).In Multi-Object Tracking(MOT),this study quantitatively confirms that Bottom-Up Tracking by Sorting(BoT-SORT)(1.88 s avg.tracklet lifetime)significantly outperforms ByteTrack(0.81 s)in identity preservation for visually similar macaques.System integration achieved 480 Frames Per Second(FPS)real-time inference on edge devices.A quadratic polynomial fittingmodel ensured high-precision aiming(RMSE<2 pixels;best 1.2 pixels)by compensating for distortion.To fundamentally solve habituation,an adaptive strategy driven by a Deep Deterministic Policy Gradient(DDPG)framework was introduced.By using a habituation penalty term(Rhabituation)to force unpredictable sequences,theDDPGstrategy achieved a stable 88%average Intrusion Frequency Reduction Rate(IFRR)in field experiments,suppressing habituation in highly intelligent species.This study develops an efficient,precise,low-cost,and habituation-resistant automated wildlife defense system.展开更多
The hard structural plane exerts a significantcontrolling effect on high-stress geological hazards in deep tunnels.Rapid and accurate acquisition of information on the hard structural planes is crucial for hazard asse...The hard structural plane exerts a significantcontrolling effect on high-stress geological hazards in deep tunnels.Rapid and accurate acquisition of information on the hard structural planes is crucial for hazard assessment,early warning,and control.First,the challenges of machine vision-based recognition of hard structural planes in deep tunnels were analyzed.Then,an adaptive illumination correction algorithm was developed to mitigate the adverse effects of lighting on the recognition of hard structural planes.Subsequently,a hard structural plane recognition algorithm integrating joint-local completion and YOLOv8 was established based on the developmental characteristics of hard structural planes,in order to address the challenge of discontinuous exposure of hard structural planes on the tunnel face.Finally,the model's performance was validated based on a deep tunnel project.The results indicate that by applying a preprocessing method combining a two-dimensional gamma function with adaptive nonlinear enhancement for uneven illumination correction,the adverse effects of lighting on hard structural plane recognition were significantlymitigated.The precision(P)increased by 7.03%,while the accuracy(Acc)improved by 4.43%.An identificationmethod incorporating automatic completion of discontinuous joints was established,which allows for the complete extraction of hard structural plane traces.The recognition accuracy increased by 8.49%.Through application of the proposed intelligent identificationmethod at the section DK196+600–700 of a deep tunnel,81.25%of the hard structural planes were completely recognized.The results can address the challenge of rapid,accurate,and noncontact recognition of hard structural planes in deep tunnels,providing a foundation for intelligent hazard assessment.展开更多
The laminar sedimentary structures of saline lacustrine mixed rocks affect both organic matter enrichment and reservoir storage performance.However,due to the small-scale nature of laminae,large-scale identification u...The laminar sedimentary structures of saline lacustrine mixed rocks affect both organic matter enrichment and reservoir storage performance.However,due to the small-scale nature of laminae,large-scale identification using well-logging data during reservoir exploration and development remains challenging.It is necessary to introduce a research method to identify and characterize the development of different types of laminae.Based on analyses of typical cores,XRD data,and welllogging curves from the upper member of the Xiaganchaigou Formation in the Yingxi area,five main types of laminae and six lamina combinations were classified.A Transformer-based intelligent recognition method was then applied to identify these lamina combinations from well-log data,with the Random Forest algorithm used as a comparative benchmark.Verification results show that the Transformer model achieves a higher total accuracy of 84%in lamina combination recognition.This study proposes a new approach for the conventional well-log characterization of laminae,in which the classification is established from the perspective of laminae genesis.It reflects the development patterns of lamina combinations driven by paleoenvironmental changes,and selects an appropriate intelligent recognition method to address the challenges in well-log characterization of such reservoirs.In terms of engineering applications,this study can accurately indicate the positions of high-quality reservoirs within sedimentary cycles during field development.It provides a sedimentary facies-controlled basis for the three-dimensional characterization of reservoir quality,thereby offering a valuable reference for the exploration and development of reservoirs formed under similar sedimentary conditions.展开更多
Wide-spectral and polarization-sensitive photodetectors are vital for applications in imaging,communication,and intelligent sensing.Although two-dimensional(2D)materials have shown great promise in enhancing the perfo...Wide-spectral and polarization-sensitive photodetectors are vital for applications in imaging,communication,and intelligent sensing.Although two-dimensional(2D)materials have shown great promise in enhancing the performance of these devices,conventional methods for spectral discrimination often rely on complex designs,such as external filters or multisensor systems,increasing system cost and complexity.Developing simplified devices that integrate spectral and polarization detection remains a key challenge.Here,we demonstrated a 2D MoTe2/GeSe-based photodetector with wide-spectral photoresponse(400 to 1064 nm)and polarization sensitivity,achieving a responsivity of 1.35 A W−1and a polarization ratio of 2.23 under 808 nm illumination.The device exhibited a unique 90°polarization reversal between green(532 nm)and red(808 nm),providing a novel mechanism for spectral discrimination.First-principles calculations reveal the polarization reversal phenomenon based on the heterostructure’s optical anisotropy.Furthermore,integration with a convolutional neural network enables intelligent traffic signal recognition using polarization-sensitive images.This work highlights the potential of MoTe2/GeSe heterostructures for next-generation photodetectors,offering compact,multifunctional solutions with integrated spectral and polarization discrimination capabilities.展开更多
Based on sensory evaluation,physicochemical indexes,and deep learning,a visual grading standard and intelligent recognition model for the cooking doneness of fried golden pompano(Trachinotus ovatus)were constructed.Ac...Based on sensory evaluation,physicochemical indexes,and deep learning,a visual grading standard and intelligent recognition model for the cooking doneness of fried golden pompano(Trachinotus ovatus)were constructed.According to the cluster analysis of sensory evaluation and physicochemical quality of fish with different frying times(0–1320 s),the fish could be classified into four levels(raw,medium rare,fully cooked,and overcooked).The correlation analysis highlighted a significant positive relationship between the level of doneness in cooking and the accompanying visual changes observed(a*and b*)(r=0.93 and 0.92,respectively),thus a visual database was established.VGGNet-19,ResNet-50,and DenseNet-121 visual learning models were introduced for cooking doneness recognition,with accuracies of 83.53%,86.75%,and 90.00%,respectively.By fine-tuning DenseNet-121,GP-Net was constructed with an accuracy improvement of nearly 6%.Overall,GP-Net could achieve real-time and rapid recognition of the cooking doneness of fried fish,promoting the industrial production of prepared fish dishes.展开更多
Water accumulation and ice formation in traffic tunnels pose prominent safety hazards(e.g.,reduced road friction,increased traffic accidents)and threaten structural integrity(e.g.,damage to waterproof layers and linin...Water accumulation and ice formation in traffic tunnels pose prominent safety hazards(e.g.,reduced road friction,increased traffic accidents)and threaten structural integrity(e.g.,damage to waterproof layers and lining structures).Therefore,the intelligent identification of these two hazards is crucial for safeguarding traffic safety and optimizing tunnel maintenance strategies.The intelligent identification system integrates computer vision,deep learning,and multi-source sensor data fusion technologies.Current state-of-the-art practices adopt deep learning models for target segmentation and detection,combined with robust image preprocessing and post-processing techniques.This technology exhibits significant practical application value,and its continuous innovation and development are expected to substantially enhance the level of tunnel safety management and structural durability preservation.展开更多
The staggered distribution of joints and fissures in space constitutes the weak part of any rock mass.The identification of rock mass structural planes and the extraction of characteristic parameters are the basis of ...The staggered distribution of joints and fissures in space constitutes the weak part of any rock mass.The identification of rock mass structural planes and the extraction of characteristic parameters are the basis of rock-mass integrity evaluation,which is very important for analysis of slope stability.The laser scanning technique can be used to acquire the coordinate information pertaining to each point of the structural plane,but large amount of point cloud data,uneven density distribution,and noise point interference make the identification efficiency and accuracy of different types of structural planes limited by point cloud data analysis technology.A new point cloud identification and segmentation algorithm for rock mass structural surfaces is proposed.Based on the distribution states of the original point cloud in different neighborhoods in space,the point clouds are characterized by multi-dimensional eigenvalues and calculated by the robust randomized Hough transform(RRHT).The normal vector difference and the final eigenvalue are proposed for characteristic distinction,and the identification of rock mass structural surfaces is completed through regional growth,which strengthens the difference expression of point clouds.In addition,nearest Voxel downsampling is also introduced in the RRHT calculation,which further reduces the number of sources of neighborhood noises,thereby improving the accuracy and stability of the calculation.The advantages of the method have been verified by laboratory models.The results showed that the proposed method can better achieve the segmentation and statistics of structural planes with interfaces and sharp boundaries.The method works well in the identification of joints,fissures,and other structural planes on Mangshezhai slope in the Three Gorges Reservoir area,China.It can provide a stable and effective technique for the identification and segmentation of rock mass structural planes,which is beneficial in engineering practice.展开更多
Despite extensive development of flexible pressure sensors,it is still difficult for them to simultaneously achieve high precision and a large response to subtle pressures.To address these challenges,this work demonst...Despite extensive development of flexible pressure sensors,it is still difficult for them to simultaneously achieve high precision and a large response to subtle pressures.To address these challenges,this work demonstrates a flexible pressure sensing platform that features the reduced graphene oxide aerogel sandwiched between a polydimethylsiloxane encapsulation layer and a thin polyimide film with interdigital electrodes.The resulting pressure sensor exhibits a high sensitivity of 698.96 kPa-1 and a low limit of detection(~1 Pa),and outstanding stability over 20,000 loading/unloading cycles.Besides monitoring various physiological signals and human motions,the flexible pressure sensors can be configured into an array layout as a smart artificial electronic skin to recognize the spatial pressure distribution.The flexible pressure sensor can also be integrated with signal processing and wireless communication modules as a teleoperation system for gesture recognition,force feedback control,and kitchen food recognition,highlighting future potential toward smart robotics and human–machine interfaces.展开更多
Infrared image recognition plays an important role in the inspection of power equipment.Existing technologies dedicated to this purpose often require manually selected features,which are not transferable and interpret...Infrared image recognition plays an important role in the inspection of power equipment.Existing technologies dedicated to this purpose often require manually selected features,which are not transferable and interpretable,and have limited training data.To address these limitations,this paper proposes an automatic infrared image recognition framework,which includes an object recognition module based on a deep self-attention network and a temperature distribution identification module based on a multi-factor similarity calculation.First,the features of an input image are extracted and embedded using a multi-head attention encoding-decoding mechanism.Thereafter,the embedded features are used to predict the equipment component category and location.In the located area,preliminary segmentation is performed.Finally,similar areas are gradually merged,and the temperature distribution of the equipment is obtained to identify a fault.Our experiments indicate that the proposed method demonstrates significantly improved accuracy compared with other related methods and,hence,provides a good reference for the automation of power equipment inspection.展开更多
With the advent of the Industry 5.0 era,the Internet of Things(IoT)devices face unprecedented proliferation,requiring higher communications rates and lower transmission delays.Considering its high spectrum efficiency,...With the advent of the Industry 5.0 era,the Internet of Things(IoT)devices face unprecedented proliferation,requiring higher communications rates and lower transmission delays.Considering its high spectrum efficiency,the promising filter bank multicarrier(FBMC)technique using offset quadrature amplitude modulation(OQAM)has been applied to Beyond 5G(B5G)industry IoT networks.However,due to the broadcasting nature of wireless channels,the FBMC-OQAMindustry IoT network is inevitably vulnerable to adversary attacks frommalicious IoT nodes.The FBMC-OQAMindustry cognitive radio network(ICRNet)is proposed to ensure security at the physical layer to tackle the above challenge.As a pivotal step of ICRNet,blind modulation recognition(BMR)can detect and recognize the modulation type of malicious signals.The previous works need to accomplish the BMR task of FBMC-OQAM signals in ICRNet nodes.A novel FBMC BMR algorithm is proposed with the transform channel convolution network(TCCNet)rather than a complicated two-dimensional convolution.Firstly,this is achieved by designing a low-complexity binary constellation diagram(BCD)gridding matrix as the input of TCCNet.Then,a transform channel convolution strategy is developed to convert the image-like BCD matrix into a serieslike data format,accelerating the BMR process while keeping discriminative features.Monte Carlo experimental results demonstrate that the proposed TCCNet obtains a performance gain of 8%and 40%over the traditional inphase/quadrature(I/Q)-based and constellation diagram(CD)-based methods at a signal noise ratio(SNR)of 12 dB,respectively.Moreover,the proposed TCCNet can achieve around 29.682 and 2.356 times faster than existing CD-Alex Network(CD-AlexNet)and I/Q-Convolutional Long Deep Neural Network(I/Q-CLDNN)algorithms,respectively.展开更多
The monitoring system designed in this paper is on account of YOLOv5(You Only Look Once)to monitor foreign objects on railway tracks and can broadcast the monitoring information to the locomotive in real time.First,th...The monitoring system designed in this paper is on account of YOLOv5(You Only Look Once)to monitor foreign objects on railway tracks and can broadcast the monitoring information to the locomotive in real time.First,the general structure of the system is determined through demand analysis and feasibility analysis,the foreign object intrusion recognition algorithm is designed,and the data set required for foreign object intrusion recognition is made.Secondly,according to the functional demands,the system selects a suitable neural web,and the programming is reasonable.At last,the system is simulated to validate its functionality(identification and classification of track intrusion and determination of a safe operating zone).展开更多
Triboelectricity-driven acoustic transducers with various merits have demonstrated significant potential in energy harvesting and self-powered sensing.The transducers generally require additionally a spacer and a corr...Triboelectricity-driven acoustic transducers with various merits have demonstrated significant potential in energy harvesting and self-powered sensing.The transducers generally require additionally a spacer and a corresponding exquisite process for smooth operation,which provides an unnecessary interface between the elements.The exploration of a novel manufacturing approach for triboelectricity-driven acoustic transducers is warranted to resolve this issue.Here,Triboelectricitydriven Oscillating Nano-Electricity generator(TONE)developed via mechanically guided fourdimensional(4D)printing is introduced for acoustic energy harvesting and self-powered voice recognition.The mechanically buckled structure of the TONE facilitates its smooth oscillation by sound wave without the use of an additional spacer,enabling the TONE to exhibit outputs of 156 V and 10μA.The output characteristics of the TONE are analyzed based on the acoustic-structuraltriboelectric interaction mechanism.The TONE demonstrates practical versatility by providing power to commercial electronics from controlled/daily sound and being utilized in artificial intelligence-based human voice recognition sensors.展开更多
Solar cell defects exhibit significant variations and multiple types,with some defect data being difficult to acquire or having small scales,posing challenges in terms of small sample and small target in defect detect...Solar cell defects exhibit significant variations and multiple types,with some defect data being difficult to acquire or having small scales,posing challenges in terms of small sample and small target in defect detection for solar cells.In order to address this issue,this paper proposes a multi-step approach for detecting the complex defects of solar cells.First,individual cell plates are extracted from electroluminescence images for block-by-block detection.Then,StyleGAN2-Ada is utilized for generative adversarial networks data augmentation to expand the number of defect samples in small sample defects.Finally,the fake dataset is combined with real dataset,and the improved YOLOv5 model is trained on this mixed dataset.Experimental results demonstrate that the proposed method achieves a superior performance in detecting the defects with small sample and small target,with the final recall rate reaching 99.7%,an increase of 3.9% compared with the unimproved model.Additionally,the precision and mean average precision are increased by 3.4% and 3.5%,respectively.Moreover,the experiments demonstrate that the improved network training on the mixed dataset can effectively enhance the detection performance of the model.The combination of these approaches significantly improves the network’s ability to detect solar cell defects.展开更多
The Lower Ganchaigou Formation in the Yingxi area of the Qaidam Basin is a typical lacustrine mixed rock reservoir in western China.It is characterized by strong interlayer heterogeneity,development of diverse lithofa...The Lower Ganchaigou Formation in the Yingxi area of the Qaidam Basin is a typical lacustrine mixed rock reservoir in western China.It is characterized by strong interlayer heterogeneity,development of diverse lithofacies types,and complex response features in logging curves.These complexities make lithofacies identification of the Ganchaigou Formation particularly challenging for non-coring wells,demanding a more efficient and accurate approach.Based on lithology and structural patterns,a lithofacies classification scheme was established.Three intelligent logging identification methods based on improved long short-term memory(LSTM)networks were constructed for lithofacies identification.The accuracy of these methods was evaluated,and the most suitable intelligent logging identification method for the reservoir lithofacies in the Yingxi area was selected.In the Upper Xiaganchaigou Formation(E32 section)of the Yingxi area,a total of eight lithofacies types were identified:laminated lime-dolostone,stratified lime-dolostone,laminated dolostonelime,stratified dolostone-lime,laminated lime-dolomitic shale,massive mudstone,sandstone,and gypsum.The overall recognition accuracies of the LSTM,Bi-LSTM,and Attention-based Bi-LSTM intelligent identification models are 81%,85%,and 87%,respectively.The overall recognition accuracies of the three intelligent algorithms are relatively high,with the Attention-based Bi-LSTM model achieving the highest accuracy.This model demonstrates superior applicability for intelligent lithofacies identification in lacustrine mixed rock reservoirs,particularly those dominated by carbonates in the Yingxi area.It effectively interprets the lithofacies types of non-coring wells in the study area and provides a valuable reference for interpreting lithofacies logs in similar depositional environments.展开更多
Recognizing handwritten characters remains a critical and formidable challenge within the realm of computervision. Although considerable strides have been made in enhancing English handwritten character recognitionthr...Recognizing handwritten characters remains a critical and formidable challenge within the realm of computervision. Although considerable strides have been made in enhancing English handwritten character recognitionthrough various techniques, deciphering Arabic handwritten characters is particularly intricate. This complexityarises from the diverse array of writing styles among individuals, coupled with the various shapes that a singlecharacter can take when positioned differently within document images, rendering the task more perplexing. Inthis study, a novel segmentation method for Arabic handwritten scripts is suggested. This work aims to locatethe local minima of the vertical and diagonal word image densities to precisely identify the segmentation pointsbetween the cursive letters. The proposed method starts with pre-processing the word image without affectingits main features, then calculates the directions pixel density of the word image by scanning it vertically and fromangles 30° to 90° to count the pixel density fromall directions and address the problem of overlapping letters, whichis a commonly attitude in writing Arabic texts by many people. Local minima and thresholds are also determinedto identify the ideal segmentation area. The proposed technique is tested on samples obtained fromtwo datasets: Aself-curated image dataset and the IFN/ENIT dataset. The results demonstrate that the proposed method achievesa significant improvement in the proportions of cursive segmentation of 92.96% on our dataset, as well as 89.37%on the IFN/ENIT dataset.展开更多
Designing stretchable and skin-conformal self-powered sensors for intelligent sensing and posture recognition is challenging.Here,based on a multi-force mixing and vulcanization process,as well as synergistically piez...Designing stretchable and skin-conformal self-powered sensors for intelligent sensing and posture recognition is challenging.Here,based on a multi-force mixing and vulcanization process,as well as synergistically piezoelectricity of BaTiO3and polyacrylonitrile,an all-in-one,stretchable,and self-powered elastomer-based piezo-pressure sensor(ASPS)with high sensitivity is reported.The ASPS presents excellent sensitivity(0.93 V/104 Pa of voltage and 4.92 nA/104 Pa of current at a pressure of 10-200 kPa)and high durability(over 10,000 cycles).Moreover,the ASPS exhibits a wide measurement range,good linearity,rapid response time,and stable frequency response.All components were fabricated using silicone,affording satisfactory skinconformality for sensing postures.Through cooperation with a homemade circuit and artificial intelligence algorithm,an information processing strategy was proposed to realize intelligent sensing and recognition.The home-made circuit achieves the acquisition and wireless transmission of ASPS signals(transmission distance up to 50 m),and the algorithm realizes the classification and identification of ASPS signals(accuracy up to 99.5%).This study proposes not only a novel fabrication method for developing self-powered sensors,but also a new information processing strategy for intelligent sensing and recognition,which offers significant application potential in human-machine interaction,physiological analysis,and medical research.展开更多
The meandering channel deposit of the upper member of Neogene Guantao Formation in Shengli Chengdao extra-shallow sea oilfield is characterized by rapid change in sedimentary facies.In addition,affected by surface tid...The meandering channel deposit of the upper member of Neogene Guantao Formation in Shengli Chengdao extra-shallow sea oilfield is characterized by rapid change in sedimentary facies.In addition,affected by surface tides and sea water reverberation,the double sensor seismic data processed by conventional methods has low signal-to-noise ratio and low resolution,and thus cannot meet the needs of seismic description and oil-bearing fluid identification of thin reservoirs less than 10 meters thick in this area.The two-step high resolution frequency bandwidth expanding processing technology was used to improve the signal-to-noise ratio and resolution of the seismic data,as a result,the dominant frequency of the seismic data was enhanced from 30 Hz to 50 Hz,and the sand body thickness resolution was enhanced from 10 m to 6 m.On the basis of fine layer control by seismic data,three types of seismic facies models,floodplain,natural levee and point bar,were defined,and the intelligent horizon-facies controlled recognition technology was worked out,which had a prediction error of reservoir thickness of less than 1.5 m.Clearly,the description accuracy of meandering channel sand bodies has been improved.The probability semi-quantitative oiliness identification method of fluid by prestack multi-parameters has been worked out by integrating Poisson’s ratio,fluid factor,product of Lame parameter and density,and other prestack elastic parameters,and the method has a coincidence rate of fluid identification of more than 90%,providing solid technical support for the exploration and development of thin reservoirs in Shengli Chengdao extra-shallow sea oilfield,which is expected to provide reference for the exploration and development of similar oilfields in China.展开更多
IEEE/CAA JOURNAL OF AUTOMATICA SINICA publishes high-quality papers in English on original theoretical and experimental research and development in all areas of automation.The coverage of this journal includes but is ...IEEE/CAA JOURNAL OF AUTOMATICA SINICA publishes high-quality papers in English on original theoretical and experimental research and development in all areas of automation.The coverage of this journal includes but is not limited to:1)Automatic control;2)Systems theory and engineering;3)Automation engineering and applications;4)Computer-aided technology for automation systems;5)Robotics;6)Artificial intelligence and intelligent control;7)Pattern recognition and intelligent systems;8)Information processing and information systems;9)Network based automation.展开更多
To improve the performance of the forest fire smoke detection model and achieve a better balance between detection accuracy and speed, an improved YOLOv4 detection model (MoAm-YOLOv4) that combines a lightweight netwo...To improve the performance of the forest fire smoke detection model and achieve a better balance between detection accuracy and speed, an improved YOLOv4 detection model (MoAm-YOLOv4) that combines a lightweight network and attention mechanism was proposed. Based on the YOLOv4 algorithm, the backbone network CSPDarknet53 was replaced with a lightweight network MobilenetV1 to reduce the model’s size. An attention mechanism was added to the three channels before the output to increase its ability to extract forest fire smoke effectively. The algorithm used the K-means clustering algorithm to cluster the smoke dataset, and obtained candidate frames that were close to the smoke images;the dataset was expanded to 2000 images by the random flip expansion method to avoid overfitting in training. The experimental results show that the improved YOLOv4 algorithm has excellent detection effect. Its mAP can reach 93.45%, precision can get 93.28%, and the model size is only 45.58 MB. Compared with YOLOv4 algorithm, MoAm-YOLOv4 improves the accuracy by 1.3% and reduces the model size by 80% while sacrificing only 0.27% mAP, showing reasonable practicability.展开更多
基金the National Natural Science Foundation of China(Grant No.52072041)the Beijing Natural Science Foundation(Grant No.JQ21007)+2 种基金the University of Chinese Academy of Sciences(Grant No.Y8540XX2D2)the Robotics Rhino-Bird Focused Research Project(No.2020-01-002)the Tencent Robotics X Laboratory.
摘要Humans can perceive our complex world through multi-sensory fusion.Under limited visual conditions,people can sense a variety of tactile signals to identify objects accurately and rapidly.However,replicating this unique capability in robots remains a significant challenge.Here,we present a new form of ultralight multifunctional tactile nano-layered carbon aerogel sensor that provides pressure,temperature,material recognition and 3D location capabilities,which is combined with multimodal supervised learning algorithms for object recognition.The sensor exhibits human-like pressure(0.04–100 kPa)and temperature(21.5–66.2℃)detection,millisecond response times(11 ms),a pressure sensitivity of 92.22 kPa−1and triboelectric durability of over 6000 cycles.The devised algorithm has universality and can accommodate a range of application scenarios.The tactile system can identify common foods in a kitchen scene with 94.63%accuracy and explore the topographic and geomorphic features of a Mars scene with 100%accuracy.This sensing approach empowers robots with versatile tactile perception to advance future society toward heightened sensing,recognition and intelligence.
基金Part of the research funding was provided by Tatung University.
摘要To address crop depredation by intelligent species(e.t,macaques)and the habituation from traditional methods,this study proposes an intelligent,closed-loop,adaptive laser deterrence system.A core contribution is an efficient multi-stage Semi-Supervised Learning(SSL)and incremental fine-tuning(IFT)framework,which reduced manual annotation by~60%and training time by~68%.This framework was benchmarked against YOLOv8n,v10n,and v11n.Our analysis revealed that YOLOv12n’s high Signal-to-Noise Ratio(SNR)(47.1%retention)pseudo-labels made it the onlymodel to gain performance(+0.010mAP)fromSSL,allowing it to overtake competitors.Subsequently,in the IFT stress test,YOLOv12n proved most robust(a minimal−0.019 mAP decline),whereas YOLOv10n suffered catastrophic failure(−0.233mAP),highlighting its incompatibility with IFT.Thefinalmodel achieved high performance(mAP@0.5 of 0.947 for macaques,0.946 for laser spots).In Multi-Object Tracking(MOT),this study quantitatively confirms that Bottom-Up Tracking by Sorting(BoT-SORT)(1.88 s avg.tracklet lifetime)significantly outperforms ByteTrack(0.81 s)in identity preservation for visually similar macaques.System integration achieved 480 Frames Per Second(FPS)real-time inference on edge devices.A quadratic polynomial fittingmodel ensured high-precision aiming(RMSE<2 pixels;best 1.2 pixels)by compensating for distortion.To fundamentally solve habituation,an adaptive strategy driven by a Deep Deterministic Policy Gradient(DDPG)framework was introduced.By using a habituation penalty term(Rhabituation)to force unpredictable sequences,theDDPGstrategy achieved a stable 88%average Intrusion Frequency Reduction Rate(IFRR)in field experiments,suppressing habituation in highly intelligent species.This study develops an efficient,precise,low-cost,and habituation-resistant automated wildlife defense system.
基金funding support from the National Natural Science Foundation of China(Grant No.52222810)China Postdoctoral Science Foundation(Grant No.2024M760371).
摘要The hard structural plane exerts a significantcontrolling effect on high-stress geological hazards in deep tunnels.Rapid and accurate acquisition of information on the hard structural planes is crucial for hazard assessment,early warning,and control.First,the challenges of machine vision-based recognition of hard structural planes in deep tunnels were analyzed.Then,an adaptive illumination correction algorithm was developed to mitigate the adverse effects of lighting on the recognition of hard structural planes.Subsequently,a hard structural plane recognition algorithm integrating joint-local completion and YOLOv8 was established based on the developmental characteristics of hard structural planes,in order to address the challenge of discontinuous exposure of hard structural planes on the tunnel face.Finally,the model's performance was validated based on a deep tunnel project.The results indicate that by applying a preprocessing method combining a two-dimensional gamma function with adaptive nonlinear enhancement for uneven illumination correction,the adverse effects of lighting on hard structural plane recognition were significantlymitigated.The precision(P)increased by 7.03%,while the accuracy(Acc)improved by 4.43%.An identificationmethod incorporating automatic completion of discontinuous joints was established,which allows for the complete extraction of hard structural plane traces.The recognition accuracy increased by 8.49%.Through application of the proposed intelligent identificationmethod at the section DK196+600–700 of a deep tunnel,81.25%of the hard structural planes were completely recognized.The results can address the challenge of rapid,accurate,and noncontact recognition of hard structural planes in deep tunnels,providing a foundation for intelligent hazard assessment.
基金financially supported by National Major Science and Technology Projects of China(No.2024ZD1406601)National Natural Science Foundation of China(Nos.42272186,42302128,42472179,42202109)。
摘要The laminar sedimentary structures of saline lacustrine mixed rocks affect both organic matter enrichment and reservoir storage performance.However,due to the small-scale nature of laminae,large-scale identification using well-logging data during reservoir exploration and development remains challenging.It is necessary to introduce a research method to identify and characterize the development of different types of laminae.Based on analyses of typical cores,XRD data,and welllogging curves from the upper member of the Xiaganchaigou Formation in the Yingxi area,five main types of laminae and six lamina combinations were classified.A Transformer-based intelligent recognition method was then applied to identify these lamina combinations from well-log data,with the Random Forest algorithm used as a comparative benchmark.Verification results show that the Transformer model achieves a higher total accuracy of 84%in lamina combination recognition.This study proposes a new approach for the conventional well-log characterization of laminae,in which the classification is established from the perspective of laminae genesis.It reflects the development patterns of lamina combinations driven by paleoenvironmental changes,and selects an appropriate intelligent recognition method to address the challenges in well-log characterization of such reservoirs.In terms of engineering applications,this study can accurately indicate the positions of high-quality reservoirs within sedimentary cycles during field development.It provides a sedimentary facies-controlled basis for the three-dimensional characterization of reservoir quality,thereby offering a valuable reference for the exploration and development of reservoirs formed under similar sedimentary conditions.
基金supported by the National Key Research and Development Program of China(Grant No.2024YFA1409700)Beijing Natural Science Foundation(Z220005)+3 种基金the National Natural Science Foundation of China(Grant No.62125404,U24A20285,12304540,62334007)CAS Project for Young Scientists in Basic Research(No.YSBR-053)the Talent Fund of Beijing Jiaotong University(2024XKRC091)the Training Program for Innovation and Entrepreneurship for Undergraduate(No.2024100041979).
摘要Wide-spectral and polarization-sensitive photodetectors are vital for applications in imaging,communication,and intelligent sensing.Although two-dimensional(2D)materials have shown great promise in enhancing the performance of these devices,conventional methods for spectral discrimination often rely on complex designs,such as external filters or multisensor systems,increasing system cost and complexity.Developing simplified devices that integrate spectral and polarization detection remains a key challenge.Here,we demonstrated a 2D MoTe2/GeSe-based photodetector with wide-spectral photoresponse(400 to 1064 nm)and polarization sensitivity,achieving a responsivity of 1.35 A W−1and a polarization ratio of 2.23 under 808 nm illumination.The device exhibited a unique 90°polarization reversal between green(532 nm)and red(808 nm),providing a novel mechanism for spectral discrimination.First-principles calculations reveal the polarization reversal phenomenon based on the heterostructure’s optical anisotropy.Furthermore,integration with a convolutional neural network enables intelligent traffic signal recognition using polarization-sensitive images.This work highlights the potential of MoTe2/GeSe heterostructures for next-generation photodetectors,offering compact,multifunctional solutions with integrated spectral and polarization discrimination capabilities.
基金supported by Special Fund for Scientific and Technological Innovation Strategy of Guangdong Province[grant number:2022A05036]Youth Fund of the National Natural Science Foundation of China[grant number:32302165]the Guangdong Basic and Applied Basic Research Foundation[grant number:2024A1515010217].
摘要Based on sensory evaluation,physicochemical indexes,and deep learning,a visual grading standard and intelligent recognition model for the cooking doneness of fried golden pompano(Trachinotus ovatus)were constructed.According to the cluster analysis of sensory evaluation and physicochemical quality of fish with different frying times(0–1320 s),the fish could be classified into four levels(raw,medium rare,fully cooked,and overcooked).The correlation analysis highlighted a significant positive relationship between the level of doneness in cooking and the accompanying visual changes observed(a*and b*)(r=0.93 and 0.92,respectively),thus a visual database was established.VGGNet-19,ResNet-50,and DenseNet-121 visual learning models were introduced for cooking doneness recognition,with accuracies of 83.53%,86.75%,and 90.00%,respectively.By fine-tuning DenseNet-121,GP-Net was constructed with an accuracy improvement of nearly 6%.Overall,GP-Net could achieve real-time and rapid recognition of the cooking doneness of fried fish,promoting the industrial production of prepared fish dishes.
摘要Water accumulation and ice formation in traffic tunnels pose prominent safety hazards(e.g.,reduced road friction,increased traffic accidents)and threaten structural integrity(e.g.,damage to waterproof layers and lining structures).Therefore,the intelligent identification of these two hazards is crucial for safeguarding traffic safety and optimizing tunnel maintenance strategies.The intelligent identification system integrates computer vision,deep learning,and multi-source sensor data fusion technologies.Current state-of-the-art practices adopt deep learning models for target segmentation and detection,combined with robust image preprocessing and post-processing techniques.This technology exhibits significant practical application value,and its continuous innovation and development are expected to substantially enhance the level of tunnel safety management and structural durability preservation.
基金the National Natural Science Foundation of China(51909136)the Open Research Fund of Key Laboratory of Geological Hazards on Three Gorges Reservoir Area(China Three Gorges University),Ministry of Education,Grant No.2022KDZ21Fund of National Major Water Conservancy Project Construction(0001212022CC60001)。
摘要The staggered distribution of joints and fissures in space constitutes the weak part of any rock mass.The identification of rock mass structural planes and the extraction of characteristic parameters are the basis of rock-mass integrity evaluation,which is very important for analysis of slope stability.The laser scanning technique can be used to acquire the coordinate information pertaining to each point of the structural plane,but large amount of point cloud data,uneven density distribution,and noise point interference make the identification efficiency and accuracy of different types of structural planes limited by point cloud data analysis technology.A new point cloud identification and segmentation algorithm for rock mass structural surfaces is proposed.Based on the distribution states of the original point cloud in different neighborhoods in space,the point clouds are characterized by multi-dimensional eigenvalues and calculated by the robust randomized Hough transform(RRHT).The normal vector difference and the final eigenvalue are proposed for characteristic distinction,and the identification of rock mass structural surfaces is completed through regional growth,which strengthens the difference expression of point clouds.In addition,nearest Voxel downsampling is also introduced in the RRHT calculation,which further reduces the number of sources of neighborhood noises,thereby improving the accuracy and stability of the calculation.The advantages of the method have been verified by laboratory models.The results showed that the proposed method can better achieve the segmentation and statistics of structural planes with interfaces and sharp boundaries.The method works well in the identification of joints,fissures,and other structural planes on Mangshezhai slope in the Three Gorges Reservoir area,China.It can provide a stable and effective technique for the identification and segmentation of rock mass structural planes,which is beneficial in engineering practice.
基金the support provided by the National Natural Science Foundation of China(52475591)the China Postdoctoral Science Foundation(2024T170651)+5 种基金the Science Research Project of Hebei Education Department(JCZX2025004)the Natural Science Foundation of Hebei Provincial(H2023202904)the Innovation Financing Program for Postgraduates of Hebei Province(CXZZBS2025046)the support provided by NIH(Award No.R21EB030140)NSF(Grant Nos.2309323,2319139,and 2243979)Penn State University。
摘要Despite extensive development of flexible pressure sensors,it is still difficult for them to simultaneously achieve high precision and a large response to subtle pressures.To address these challenges,this work demonstrates a flexible pressure sensing platform that features the reduced graphene oxide aerogel sandwiched between a polydimethylsiloxane encapsulation layer and a thin polyimide film with interdigital electrodes.The resulting pressure sensor exhibits a high sensitivity of 698.96 kPa-1 and a low limit of detection(~1 Pa),and outstanding stability over 20,000 loading/unloading cycles.Besides monitoring various physiological signals and human motions,the flexible pressure sensors can be configured into an array layout as a smart artificial electronic skin to recognize the spatial pressure distribution.The flexible pressure sensor can also be integrated with signal processing and wireless communication modules as a teleoperation system for gesture recognition,force feedback control,and kitchen food recognition,highlighting future potential toward smart robotics and human–machine interfaces.
基金This work was supported by National Key R&D Program of China(2019YFE0102900).
摘要Infrared image recognition plays an important role in the inspection of power equipment.Existing technologies dedicated to this purpose often require manually selected features,which are not transferable and interpretable,and have limited training data.To address these limitations,this paper proposes an automatic infrared image recognition framework,which includes an object recognition module based on a deep self-attention network and a temperature distribution identification module based on a multi-factor similarity calculation.First,the features of an input image are extracted and embedded using a multi-head attention encoding-decoding mechanism.Thereafter,the embedded features are used to predict the equipment component category and location.In the located area,preliminary segmentation is performed.Finally,similar areas are gradually merged,and the temperature distribution of the equipment is obtained to identify a fault.Our experiments indicate that the proposed method demonstrates significantly improved accuracy compared with other related methods and,hence,provides a good reference for the automation of power equipment inspection.
基金supported by the National Natural Science Foundation of China(Nos.61671095,61371164)the Project of Key Laboratory of Signal and Information Processing of Chongqing(No.CSTC2009CA2003).
摘要With the advent of the Industry 5.0 era,the Internet of Things(IoT)devices face unprecedented proliferation,requiring higher communications rates and lower transmission delays.Considering its high spectrum efficiency,the promising filter bank multicarrier(FBMC)technique using offset quadrature amplitude modulation(OQAM)has been applied to Beyond 5G(B5G)industry IoT networks.However,due to the broadcasting nature of wireless channels,the FBMC-OQAMindustry IoT network is inevitably vulnerable to adversary attacks frommalicious IoT nodes.The FBMC-OQAMindustry cognitive radio network(ICRNet)is proposed to ensure security at the physical layer to tackle the above challenge.As a pivotal step of ICRNet,blind modulation recognition(BMR)can detect and recognize the modulation type of malicious signals.The previous works need to accomplish the BMR task of FBMC-OQAM signals in ICRNet nodes.A novel FBMC BMR algorithm is proposed with the transform channel convolution network(TCCNet)rather than a complicated two-dimensional convolution.Firstly,this is achieved by designing a low-complexity binary constellation diagram(BCD)gridding matrix as the input of TCCNet.Then,a transform channel convolution strategy is developed to convert the image-like BCD matrix into a serieslike data format,accelerating the BMR process while keeping discriminative features.Monte Carlo experimental results demonstrate that the proposed TCCNet obtains a performance gain of 8%and 40%over the traditional inphase/quadrature(I/Q)-based and constellation diagram(CD)-based methods at a signal noise ratio(SNR)of 12 dB,respectively.Moreover,the proposed TCCNet can achieve around 29.682 and 2.356 times faster than existing CD-Alex Network(CD-AlexNet)and I/Q-Convolutional Long Deep Neural Network(I/Q-CLDNN)algorithms,respectively.
摘要The monitoring system designed in this paper is on account of YOLOv5(You Only Look Once)to monitor foreign objects on railway tracks and can broadcast the monitoring information to the locomotive in real time.First,the general structure of the system is determined through demand analysis and feasibility analysis,the foreign object intrusion recognition algorithm is designed,and the data set required for foreign object intrusion recognition is made.Secondly,according to the functional demands,the system selects a suitable neural web,and the programming is reasonable.At last,the system is simulated to validate its functionality(identification and classification of track intrusion and determination of a safe operating zone).
基金supported by the National Research Foundation of Korea(NRF)grant funded by the Korea government(MSIT)(No.RS-2024-00344920)supported by the Human Resources Development of the Korea Institute of Energy Technology Evaluation and Planning(KETEP)grant funded by the Ministry of Trade,Industry and Energy of Korea(No.RS-2023-00244330)。
摘要Triboelectricity-driven acoustic transducers with various merits have demonstrated significant potential in energy harvesting and self-powered sensing.The transducers generally require additionally a spacer and a corresponding exquisite process for smooth operation,which provides an unnecessary interface between the elements.The exploration of a novel manufacturing approach for triboelectricity-driven acoustic transducers is warranted to resolve this issue.Here,Triboelectricitydriven Oscillating Nano-Electricity generator(TONE)developed via mechanically guided fourdimensional(4D)printing is introduced for acoustic energy harvesting and self-powered voice recognition.The mechanically buckled structure of the TONE facilitates its smooth oscillation by sound wave without the use of an additional spacer,enabling the TONE to exhibit outputs of 156 V and 10μA.The output characteristics of the TONE are analyzed based on the acoustic-structuraltriboelectric interaction mechanism.The TONE demonstrates practical versatility by providing power to commercial electronics from controlled/daily sound and being utilized in artificial intelligence-based human voice recognition sensors.
摘要Solar cell defects exhibit significant variations and multiple types,with some defect data being difficult to acquire or having small scales,posing challenges in terms of small sample and small target in defect detection for solar cells.In order to address this issue,this paper proposes a multi-step approach for detecting the complex defects of solar cells.First,individual cell plates are extracted from electroluminescence images for block-by-block detection.Then,StyleGAN2-Ada is utilized for generative adversarial networks data augmentation to expand the number of defect samples in small sample defects.Finally,the fake dataset is combined with real dataset,and the improved YOLOv5 model is trained on this mixed dataset.Experimental results demonstrate that the proposed method achieves a superior performance in detecting the defects with small sample and small target,with the final recall rate reaching 99.7%,an increase of 3.9% compared with the unimproved model.Additionally,the precision and mean average precision are increased by 3.4% and 3.5%,respectively.Moreover,the experiments demonstrate that the improved network training on the mixed dataset can effectively enhance the detection performance of the model.The combination of these approaches significantly improves the network’s ability to detect solar cell defects.
基金supported by the the National Natural Science Foundation of China(No.42272186,42302128,42202109 and 42472179)
摘要The Lower Ganchaigou Formation in the Yingxi area of the Qaidam Basin is a typical lacustrine mixed rock reservoir in western China.It is characterized by strong interlayer heterogeneity,development of diverse lithofacies types,and complex response features in logging curves.These complexities make lithofacies identification of the Ganchaigou Formation particularly challenging for non-coring wells,demanding a more efficient and accurate approach.Based on lithology and structural patterns,a lithofacies classification scheme was established.Three intelligent logging identification methods based on improved long short-term memory(LSTM)networks were constructed for lithofacies identification.The accuracy of these methods was evaluated,and the most suitable intelligent logging identification method for the reservoir lithofacies in the Yingxi area was selected.In the Upper Xiaganchaigou Formation(E32 section)of the Yingxi area,a total of eight lithofacies types were identified:laminated lime-dolostone,stratified lime-dolostone,laminated dolostonelime,stratified dolostone-lime,laminated lime-dolomitic shale,massive mudstone,sandstone,and gypsum.The overall recognition accuracies of the LSTM,Bi-LSTM,and Attention-based Bi-LSTM intelligent identification models are 81%,85%,and 87%,respectively.The overall recognition accuracies of the three intelligent algorithms are relatively high,with the Attention-based Bi-LSTM model achieving the highest accuracy.This model demonstrates superior applicability for intelligent lithofacies identification in lacustrine mixed rock reservoirs,particularly those dominated by carbonates in the Yingxi area.It effectively interprets the lithofacies types of non-coring wells in the study area and provides a valuable reference for interpreting lithofacies logs in similar depositional environments.
摘要Recognizing handwritten characters remains a critical and formidable challenge within the realm of computervision. Although considerable strides have been made in enhancing English handwritten character recognitionthrough various techniques, deciphering Arabic handwritten characters is particularly intricate. This complexityarises from the diverse array of writing styles among individuals, coupled with the various shapes that a singlecharacter can take when positioned differently within document images, rendering the task more perplexing. Inthis study, a novel segmentation method for Arabic handwritten scripts is suggested. This work aims to locatethe local minima of the vertical and diagonal word image densities to precisely identify the segmentation pointsbetween the cursive letters. The proposed method starts with pre-processing the word image without affectingits main features, then calculates the directions pixel density of the word image by scanning it vertically and fromangles 30° to 90° to count the pixel density fromall directions and address the problem of overlapping letters, whichis a commonly attitude in writing Arabic texts by many people. Local minima and thresholds are also determinedto identify the ideal segmentation area. The proposed technique is tested on samples obtained fromtwo datasets: Aself-curated image dataset and the IFN/ENIT dataset. The results demonstrate that the proposed method achievesa significant improvement in the proportions of cursive segmentation of 92.96% on our dataset, as well as 89.37%on the IFN/ENIT dataset.
基金supported by the National Natural Science Foundation of China(Nos.62101513,51975542,52175554,and 62171414)China Postdoctoral Science Foundation(Nos.2022TQ0230 and 2022M712324)+2 种基金Shanxi“1331 Project”Key Subject Construction(No.1331KSC)the Fundamental Research Program of Shanxi Province(No.20210302124170)Young Academic Leaders of North University of China(No.11045501).
摘要Designing stretchable and skin-conformal self-powered sensors for intelligent sensing and posture recognition is challenging.Here,based on a multi-force mixing and vulcanization process,as well as synergistically piezoelectricity of BaTiO3and polyacrylonitrile,an all-in-one,stretchable,and self-powered elastomer-based piezo-pressure sensor(ASPS)with high sensitivity is reported.The ASPS presents excellent sensitivity(0.93 V/104 Pa of voltage and 4.92 nA/104 Pa of current at a pressure of 10-200 kPa)and high durability(over 10,000 cycles).Moreover,the ASPS exhibits a wide measurement range,good linearity,rapid response time,and stable frequency response.All components were fabricated using silicone,affording satisfactory skinconformality for sensing postures.Through cooperation with a homemade circuit and artificial intelligence algorithm,an information processing strategy was proposed to realize intelligent sensing and recognition.The home-made circuit achieves the acquisition and wireless transmission of ASPS signals(transmission distance up to 50 m),and the algorithm realizes the classification and identification of ASPS signals(accuracy up to 99.5%).This study proposes not only a novel fabrication method for developing self-powered sensors,but also a new information processing strategy for intelligent sensing and recognition,which offers significant application potential in human-machine interaction,physiological analysis,and medical research.
基金Supported by the China National Science and Technology Major Project(2016zx05006)Sinopec Program for Science and Technology Development(P15156,P15159)。
摘要The meandering channel deposit of the upper member of Neogene Guantao Formation in Shengli Chengdao extra-shallow sea oilfield is characterized by rapid change in sedimentary facies.In addition,affected by surface tides and sea water reverberation,the double sensor seismic data processed by conventional methods has low signal-to-noise ratio and low resolution,and thus cannot meet the needs of seismic description and oil-bearing fluid identification of thin reservoirs less than 10 meters thick in this area.The two-step high resolution frequency bandwidth expanding processing technology was used to improve the signal-to-noise ratio and resolution of the seismic data,as a result,the dominant frequency of the seismic data was enhanced from 30 Hz to 50 Hz,and the sand body thickness resolution was enhanced from 10 m to 6 m.On the basis of fine layer control by seismic data,three types of seismic facies models,floodplain,natural levee and point bar,were defined,and the intelligent horizon-facies controlled recognition technology was worked out,which had a prediction error of reservoir thickness of less than 1.5 m.Clearly,the description accuracy of meandering channel sand bodies has been improved.The probability semi-quantitative oiliness identification method of fluid by prestack multi-parameters has been worked out by integrating Poisson’s ratio,fluid factor,product of Lame parameter and density,and other prestack elastic parameters,and the method has a coincidence rate of fluid identification of more than 90%,providing solid technical support for the exploration and development of thin reservoirs in Shengli Chengdao extra-shallow sea oilfield,which is expected to provide reference for the exploration and development of similar oilfields in China.
摘要IEEE/CAA JOURNAL OF AUTOMATICA SINICA publishes high-quality papers in English on original theoretical and experimental research and development in all areas of automation.The coverage of this journal includes but is not limited to:1)Automatic control;2)Systems theory and engineering;3)Automation engineering and applications;4)Computer-aided technology for automation systems;5)Robotics;6)Artificial intelligence and intelligent control;7)Pattern recognition and intelligent systems;8)Information processing and information systems;9)Network based automation.
摘要To improve the performance of the forest fire smoke detection model and achieve a better balance between detection accuracy and speed, an improved YOLOv4 detection model (MoAm-YOLOv4) that combines a lightweight network and attention mechanism was proposed. Based on the YOLOv4 algorithm, the backbone network CSPDarknet53 was replaced with a lightweight network MobilenetV1 to reduce the model’s size. An attention mechanism was added to the three channels before the output to increase its ability to extract forest fire smoke effectively. The algorithm used the K-means clustering algorithm to cluster the smoke dataset, and obtained candidate frames that were close to the smoke images;the dataset was expanded to 2000 images by the random flip expansion method to avoid overfitting in training. The experimental results show that the improved YOLOv4 algorithm has excellent detection effect. Its mAP can reach 93.45%, precision can get 93.28%, and the model size is only 45.58 MB. Compared with YOLOv4 algorithm, MoAm-YOLOv4 improves the accuracy by 1.3% and reduces the model size by 80% while sacrificing only 0.27% mAP, showing reasonable practicability.