Impulse components in vibration signals are important fault features of complex machines. Sparse coding (SC) algorithm has been introduced as an impulse feature extraction method, but it could not guarantee a satisf...Impulse components in vibration signals are important fault features of complex machines. Sparse coding (SC) algorithm has been introduced as an impulse feature extraction method, but it could not guarantee a satisfactory performance in processing vibration signals with heavy background noises. In this paper, a method based on fusion sparse coding (FSC) and online dictionary learning is proposed to extract impulses efficiently. Firstly, fusion scheme of different sparse coding algorithms is presented to ensure higher reconstruction accuracy. Then, an improved online dictionary learning method using FSC scheme is established to obtain redundant dictionary and it can capture specific features of training samples and reconstruct the sparse approximation of vibration signals. Simulation shows that this method has a good performance in solving sparse coefficients and training redundant dictionary compared with other methods. Lastly, the proposed method is further applied to processing aircraft engine rotor vibration signals. Compared with other feature extraction approaches, our method can extract impulse features accurately and efficiently from heavy noisy vibration signal, which has significant supports for machinery fault detection and diagnosis.展开更多
A novel hashing method based on multiple heterogeneous features is proposed to improve the accuracy of the image retrieval system. First, it leverages the imbalanced distribution of the similar and dissimilar samples ...A novel hashing method based on multiple heterogeneous features is proposed to improve the accuracy of the image retrieval system. First, it leverages the imbalanced distribution of the similar and dissimilar samples in the feature space to boost the performance of each weak classifier in the asymmetric boosting framework. Then, the weak classifier based on a novel linear discriminate analysis (LDA) algorithm which is learned from the subspace of heterogeneous features is integrated into the framework. Finally, the proposed method deals with each bit of the code sequentially, which utilizes the samples misclassified in each round in order to learn compact and balanced code. The heterogeneous information from different modalities can be effectively complementary to each other, which leads to much higher performance. The experimental results based on the two public benchmarks demonstrate that this method is superior to many of the state- of-the-art methods. In conclusion, the performance of the retrieval system can be improved with the help of multiple heterogeneous features and the compact hash codes which can be learned by the imbalanced learning method.展开更多
In recent years,with the massive growth of image data,how to match the image required by users quickly and efficiently becomes a challenge.Compared with single-view feature,multi-view feature is more accurate to descr...In recent years,with the massive growth of image data,how to match the image required by users quickly and efficiently becomes a challenge.Compared with single-view feature,multi-view feature is more accurate to describe image information.The advantages of hash method in reducing data storage and improving efficiency also make us study how to effectively apply to large-scale image retrieval.In this paper,a hash algorithm of multi-index image retrieval based on multi-view feature coding is proposed.By learning the data correlation between different views,this algorithm uses multi-view data with deeper level image semantics to achieve better retrieval results.This algorithm uses a quantitative hash method to generate binary sequences,and uses the hash code generated by the association features to construct database inverted index files,so as to reduce the memory burden and promote the efficient matching.In order to reduce the matching error of hash code and ensure the retrieval accuracy,this algorithm uses inverted multi-index structure instead of single-index structure.Compared with other advanced image retrieval method,this method has better retrieval performance.展开更多
To solve the problems of the AMR-WB+(Extended Adaptive Multi-Rate-WideBand)semi-open-loop coding mode selection algorithm,features for ACELP(Algebraic Code Excited Linear Prediction)and TCX(Transform Coded eXcitation)...To solve the problems of the AMR-WB+(Extended Adaptive Multi-Rate-WideBand)semi-open-loop coding mode selection algorithm,features for ACELP(Algebraic Code Excited Linear Prediction)and TCX(Transform Coded eXcitation)classification are investigated.11 classifying features in the AMR-WB+codec are selected and 2 novel classifying features,i.e.,EFM(Energy Flatness Measurement)and stdEFM(standard deviation of EFM),are proposed.Consequently,a novel semi-open-loop mode selection algorithm based on EFM and selected AMR-WB+features is proposed.The results of classifying test and listening test show that the performance of the novel algorithm is much better than that of the AMR-WB+semi-open-loop coding mode selection algorithm.展开更多
To solve the problem that using a single feature cannot play the role of multiple features of Android application in malicious code detection, an Android malicious code detection mechanism is proposed based on integra...To solve the problem that using a single feature cannot play the role of multiple features of Android application in malicious code detection, an Android malicious code detection mechanism is proposed based on integrated learning on the basis of dynamic and static detection. Considering three types of Android behavior characteristics, a three-layer hybrid algorithm was proposed. And it combined the malicious code detection based on digital signature to improve the detection efficiency. The digital signature of the known malicious code was extracted to form a malicious sample library. The authority that can reflect Android malicious behavior, API call and the running system call features were also extracted. An expandable hybrid discriminant algorithm was designed for the above three types of features. The algorithm was tested with machine learning method by constructing the optimal classifier suitable for the above features. Finally, the Android malicious code detection system was designed and implemented based on the multi-layer hybrid algorithm. The experimental results show that the system performs Android malicious code detection based on the combination of signature and dynamic and static features. Compared with other related work, the system has better performance in execution efficiency and detection rate.展开更多
In expression recognition, feature representation is critical for successful recognition since it contains distinctive information of expressions. In this paper, a new approach for representing facial expression featu...In expression recognition, feature representation is critical for successful recognition since it contains distinctive information of expressions. In this paper, a new approach for representing facial expression features is proposed with its objective to describe features in an effective and efficient way in order to improve the recognition performance. The method combines the facial action coding system(FACS) and 'uniform' local binary patterns(LBP) to represent facial expression features from coarse to fine. The facial feature regions are extracted by active shape models(ASM) based on FACS to obtain the gray-level texture. Then, LBP is used to represent expression features for enhancing the discriminant. A facial expression recognition system is developed based on this feature extraction method by using K nearest neighborhood(K-NN) classifier to recognize facial expressions. Finally, experiments are carried out to evaluate this feature extraction method. The significance of removing the unrelated facial regions and enhancing the discrimination ability of expression features in the recognition process is indicated by the results, in addition to its convenience.展开更多
Bubble seed image filling is an important prerequisite for the image segmentation of flotation bubble that can be used to improve flotation automatic control.These common image filling algorithms in dealing with compl...Bubble seed image filling is an important prerequisite for the image segmentation of flotation bubble that can be used to improve flotation automatic control.These common image filling algorithms in dealing with complex bubble image exists under-filling and over-filling problems.A new filling algorithm based on boundary point feature and scan lines(PFSL)is proposed in the paper.The filling a|gorithm describes these boundary points of image objects by means of chain codes.The features of each boundary point,including convex points,concave points,left points and right points,are defined by the point's entrancing chain code and leaving chain code.The algorithm firstly finds out all double-matched boundary points based on the features of boundary points,and fill image objects by these double-matched boundary points on scan lines.Experimental results of bubble seed image filling show that under-filling and over-filling problem can be eliminated by the proposed algorithm.展开更多
Two signature systems based on smart cards and fingerprint features are proposed. In one signature system, the cryptographic key is stored in the smart card and is only accessible when the signer's extracted fingerpr...Two signature systems based on smart cards and fingerprint features are proposed. In one signature system, the cryptographic key is stored in the smart card and is only accessible when the signer's extracted fingerprint features match his stored template. To resist being tampered on public channel, the user's message and the signed message are encrypted by the signer's public key and the user's public key, respectively. In the other signature system, the keys are generated by combining the signer's fingerprint features, check bits, and a rememberable key, and there are no matching process and keys stored on the smart card. Additionally, there is generally more than one public key in this system, that is, there exist some pseudo public keys except a real one.展开更多
To extract features of fabric defects effectively and reduce dimension of feature space,a feature extraction method of fabric defects based on complex contourlet transform (CCT) and principal component analysis (PC...To extract features of fabric defects effectively and reduce dimension of feature space,a feature extraction method of fabric defects based on complex contourlet transform (CCT) and principal component analysis (PCA) is proposed.Firstly,training samples of fabric defect images are decomposed by CCT.Secondly,PCA is applied in the obtained low-frequency component and part of highfrequency components to get a lower dimensional feature space.Finally,components of testing samples obtained by CCT are projected onto the feature space where different types of fabric defects are distinguished by the minimum Euclidean distance method.A large number of experimental results show that,compared with PCA,the method combining wavdet low-frequency component with PCA (WLPCA),the method combining contourlet transform with PCA (CPCA),and the method combining wavelet low-frequency and highfrequency components with PCA (WPCA),the proposed method can extract features of common fabric defect types effectively.The recognition rate is greatly improved while the dimension is reduced.展开更多
Stance detection is the task of attitude identification toward a standpoint.Previous work of stance detection has focused on feature extraction but ignored the fact that irrelevant features exist as noise during highe...Stance detection is the task of attitude identification toward a standpoint.Previous work of stance detection has focused on feature extraction but ignored the fact that irrelevant features exist as noise during higher-level abstracting.Moreover,because the target is not always mentioned in the text,most methods have ignored target information.In order to solve these problems,we propose a neural network ensemble method that combines the timing dependence bases on long short-term memory(LSTM)and the excellent extracting performance of convolutional neural networks(CNNs).The method can obtain multi-level features that consider both local and global features.We also introduce attention mechanisms to magnify target information-related features.Furthermore,we employ sparse coding to remove noise to obtain characteristic features.Performance was improved by using sparse coding on the basis of attention employment and feature extraction.We evaluate our approach on the SemEval-2016Task 6-A public dataset,achieving a performance that exceeds the benchmark and those of participating teams.展开更多
In recent years, the accuracy of speech recognition (SR) has been one of the most active areas of research. Despite that SR systems are working reasonably well in quiet conditions, they still suffer severe performance...In recent years, the accuracy of speech recognition (SR) has been one of the most active areas of research. Despite that SR systems are working reasonably well in quiet conditions, they still suffer severe performance degradation in noisy conditions or distorted channels. It is necessary to search for more robust feature extraction methods to gain better performance in adverse conditions. This paper investigates the performance of conventional and new hybrid speech feature extraction algorithms of Mel Frequency Cepstrum Coefficient (MFCC), Linear Prediction Coding Coefficient (LPCC), perceptual linear production (PLP), and RASTA-PLP in noisy conditions through using multivariate Hidden Markov Model (HMM) classifier. The behavior of the proposal system is evaluated using TIDIGIT human voice dataset corpora, recorded from 208 different adult speakers in both training and testing process. The theoretical basis for speech processing and classifier procedures were presented, and the recognition results were obtained based on word recognition rate.展开更多
Wake-Up-Word Speech Recognition task (WUW-SR) is a computationally very demand, particularly the stage of feature extraction which is decoded with corresponding Hidden Markov Models (HMMs) in the back-end stage of the...Wake-Up-Word Speech Recognition task (WUW-SR) is a computationally very demand, particularly the stage of feature extraction which is decoded with corresponding Hidden Markov Models (HMMs) in the back-end stage of the WUW-SR. The state of the art WUW-SR system is based on three different sets of features: Mel-Frequency Cepstral Coefficients (MFCC), Linear Predictive Coding Coefficients (LPC), and Enhanced Mel-Frequency Cepstral Coefficients (ENH_MFCC). In (front-end of Wake-Up-Word Speech Recognition System Design on FPGA) [1], we presented an experimental FPGA design and implementation of a novel architecture of a real-time spectrogram extraction processor that generates MFCC, LPC, and ENH_MFCC spectrograms simultaneously. In this paper, the details of converting the three sets of spectrograms 1) Mel-Frequency Cepstral Coefficients (MFCC), 2) Linear Predictive Coding Coefficients (LPC), and 3) Enhanced Mel-Frequency Cepstral Coefficients (ENH_MFCC) to their equivalent features are presented. In the WUW- SR system, the recognizer’s frontend is located at the terminal which is typically connected over a data network to remote back-end recognition (e.g., server). The WUW-SR is shown in Figure 1. The three sets of speech features are extracted at the front-end. These extracted features are then compressed and transmitted to the server via a dedicated channel, where subsequently they are decoded.展开更多
近年来,随着物联网(Internet of Things,IoT)、语义通信以及智慧城市等经典机器间通信(Machine to Machine,M2M)场景的快速发展,海量视觉数据在设备间的实时传输与高效处理成为了一项关键挑战。在此背景下,传统以人眼感知质量为核心的...近年来,随着物联网(Internet of Things,IoT)、语义通信以及智慧城市等经典机器间通信(Machine to Machine,M2M)场景的快速发展,海量视觉数据在设备间的实时传输与高效处理成为了一项关键挑战。在此背景下,传统以人眼感知质量为核心的图像编码方法,因其优化目标与机器视觉任务需求存在本质差异,往往在面向机器视觉分析时出现分析精度不足的问题。为此,面向机器视觉的图像编码(Image Coding for Machine,ICM)应运而生,其核心目标是在保证下游机器视觉任务(如分类、检测、分割等)分析精度的同时,实现尽可能低的编码码率,从而更好地适配M2M场景中的带宽与存储约束。然而,现有ICM方法仍面临两大瓶颈:其一,在极低码率条件下性能急剧下降。这是由于现有方法多依赖于端到端的非线性变换提取视觉特征,未能充分挖掘和利用图像中高层语义信息的紧凑表示,导致特征编码效率不足;其二,在开放场景下的泛化能力弱。多数方法针对单一任务、单一数据集进行优化,缺乏对未知类别、跨域数据的适应能力,难以在实际动态环境中保持稳定的分析性能。为突破上述限制,本文提出一种文本提示引导的面向机器视觉图像编码框架(Text-prompted Image Coding for Machine,T-ICM)。该框架的核心思想是将图像信息解耦为语义信息与纹理信息两个互补的组成部分,其中,语义信息以结构化文本提示(如对象类别、位置描述)的形式进行表示与编码,纹理信息则通过一种任务无关的通用视觉特征进行提取与压缩。在编码端,文本提示因其高度抽象和语义紧凑的特性,可以显著降低整体码率;通用特征则通过我们提出的分组特征编码模块进行高效压缩。在解码端,文本提示不仅用于直接解析完成分类、检测等任务,更重要的是作为引导信号,通过提示编码器与掩膜解码器,动态调整重建通用特征的语义感知区域,实现特征层面的域自适应与任务适配,从而显著提升模型在开放场景下的鲁棒性。本文在多个标准数据集与任务上对T-ICM进行了全面评估。实验表明,在语义分割和实例分割等密集预测任务上,T-ICM在极低码率下仍能保持接近原始图像输入的分析精度,其性能显著优于H.266/VVC、基于深度学习的图像编码器以及现有的其他ICM方法。本研究通过将语义信息迁移至高度压缩的文本模态进行传输,并利用其引导特征重建,T-ICM在编码效率与任务性能之间实现了更优的权衡,为未来语义通信、边缘智能协同,以及自适应机器视觉系统的发展提供了新的思路与技术支撑。展开更多
基金supported by the National Natural Science Foundation of China (No. 51201182)
摘要Impulse components in vibration signals are important fault features of complex machines. Sparse coding (SC) algorithm has been introduced as an impulse feature extraction method, but it could not guarantee a satisfactory performance in processing vibration signals with heavy background noises. In this paper, a method based on fusion sparse coding (FSC) and online dictionary learning is proposed to extract impulses efficiently. Firstly, fusion scheme of different sparse coding algorithms is presented to ensure higher reconstruction accuracy. Then, an improved online dictionary learning method using FSC scheme is established to obtain redundant dictionary and it can capture specific features of training samples and reconstruct the sparse approximation of vibration signals. Simulation shows that this method has a good performance in solving sparse coefficients and training redundant dictionary compared with other methods. Lastly, the proposed method is further applied to processing aircraft engine rotor vibration signals. Compared with other feature extraction approaches, our method can extract impulse features accurately and efficiently from heavy noisy vibration signal, which has significant supports for machinery fault detection and diagnosis.
基金The National Natural Science Foundation of China(No.61305058)the Natural Science Foundation of Higher Education Institutions of Jiangsu Province(No.12KJB520003)+1 种基金the Natural Science Foundation of Jiangsu Province(No.BK20130471)the Scientific Research Foundation for Advanced Talents by Jiangsu University(No.13JDG093)
摘要A novel hashing method based on multiple heterogeneous features is proposed to improve the accuracy of the image retrieval system. First, it leverages the imbalanced distribution of the similar and dissimilar samples in the feature space to boost the performance of each weak classifier in the asymmetric boosting framework. Then, the weak classifier based on a novel linear discriminate analysis (LDA) algorithm which is learned from the subspace of heterogeneous features is integrated into the framework. Finally, the proposed method deals with each bit of the code sequentially, which utilizes the samples misclassified in each round in order to learn compact and balanced code. The heterogeneous information from different modalities can be effectively complementary to each other, which leads to much higher performance. The experimental results based on the two public benchmarks demonstrate that this method is superior to many of the state- of-the-art methods. In conclusion, the performance of the retrieval system can be improved with the help of multiple heterogeneous features and the compact hash codes which can be learned by the imbalanced learning method.
基金supported in part by the National Natural Science Foundation of China under Grant 61772561,author J.Q,http://gffzzf112c495998e46dehqxf9nb5www5v6bun.ffgz.tsg.suse.edu.cn/in part by the Key Research and Development Plan of Hunan Province under Grant 2018NK2012,author J.Q,http://gffzzc62e3dd5537e48ffhqxf9nb5www5v6bun.ffgz.tsg.suse.edu.cn/+7 种基金in part by the Key Research and Development Plan of Hunan Province under Grant 2019SK2022,author Y.T,http://gffzzc62e3dd5537e48ffhqxf9nb5www5v6bun.ffgz.tsg.suse.edu.cn/in part by the Science Research Projects of Hunan Provincial Education Department under Grant 18A174,author X.X,http://gffzzb782639b0a1c434fhqxf9nb5www5v6bun.ffgz.tsg.suse.edu.cn/in part by the Science Research Projects of Hunan Provincial Education Department under Grant 19B584,author Y.T,http://gffzzb782639b0a1c434fhqxf9nb5www5v6bun.ffgz.tsg.suse.edu.cn/in part by the Degree&Postgraduate Education Reform Project of Hunan Province under Grant 2019JGYB154,author J.Q,http://gffzzb07429f713164f88hqxf9nb5www5v6bun.ffgz.tsg.suse.edu.cn/in part by the Postgraduate Excellent teaching team Project of Hunan Province under Grant[2019]370-133,author J.Q,http://gffzzb07429f713164f88hqxf9nb5www5v6bun.ffgz.tsg.suse.edu.cn/in part by the Postgraduate Education and Teaching Reform Project of Central South University of Forestry&Technology under Grant 2019JG013,author X.X,http://gffzz022d7edae406446chqxf9nb5www5v6bun.ffgz.tsg.suse.edu.cn/in part by the Natural Science Foundation of Hunan Province(No.2020JJ4140),author Y.T,http://gffzzc62e3dd5537e48ffhqxf9nb5www5v6bun.ffgz.tsg.suse.edu.cn/in part by the Natural Science Foundation of Hunan Province(No.2020JJ4141),author X.X,http://gffzzc62e3dd5537e48ffhqxf9nb5www5v6bun.ffgz.tsg.suse.edu.cn/.
摘要In recent years,with the massive growth of image data,how to match the image required by users quickly and efficiently becomes a challenge.Compared with single-view feature,multi-view feature is more accurate to describe image information.The advantages of hash method in reducing data storage and improving efficiency also make us study how to effectively apply to large-scale image retrieval.In this paper,a hash algorithm of multi-index image retrieval based on multi-view feature coding is proposed.By learning the data correlation between different views,this algorithm uses multi-view data with deeper level image semantics to achieve better retrieval results.This algorithm uses a quantitative hash method to generate binary sequences,and uses the hash code generated by the association features to construct database inverted index files,so as to reduce the memory burden and promote the efficient matching.In order to reduce the matching error of hash code and ensure the retrieval accuracy,this algorithm uses inverted multi-index structure instead of single-index structure.Compared with other advanced image retrieval method,this method has better retrieval performance.
摘要To solve the problems of the AMR-WB+(Extended Adaptive Multi-Rate-WideBand)semi-open-loop coding mode selection algorithm,features for ACELP(Algebraic Code Excited Linear Prediction)and TCX(Transform Coded eXcitation)classification are investigated.11 classifying features in the AMR-WB+codec are selected and 2 novel classifying features,i.e.,EFM(Energy Flatness Measurement)and stdEFM(standard deviation of EFM),are proposed.Consequently,a novel semi-open-loop mode selection algorithm based on EFM and selected AMR-WB+features is proposed.The results of classifying test and listening test show that the performance of the novel algorithm is much better than that of the AMR-WB+semi-open-loop coding mode selection algorithm.
摘要To solve the problem that using a single feature cannot play the role of multiple features of Android application in malicious code detection, an Android malicious code detection mechanism is proposed based on integrated learning on the basis of dynamic and static detection. Considering three types of Android behavior characteristics, a three-layer hybrid algorithm was proposed. And it combined the malicious code detection based on digital signature to improve the detection efficiency. The digital signature of the known malicious code was extracted to form a malicious sample library. The authority that can reflect Android malicious behavior, API call and the running system call features were also extracted. An expandable hybrid discriminant algorithm was designed for the above three types of features. The algorithm was tested with machine learning method by constructing the optimal classifier suitable for the above features. Finally, the Android malicious code detection system was designed and implemented based on the multi-layer hybrid algorithm. The experimental results show that the system performs Android malicious code detection based on the combination of signature and dynamic and static features. Compared with other related work, the system has better performance in execution efficiency and detection rate.
基金supported by National Natural Science Foundation of China(No.61273339)
摘要In expression recognition, feature representation is critical for successful recognition since it contains distinctive information of expressions. In this paper, a new approach for representing facial expression features is proposed with its objective to describe features in an effective and efficient way in order to improve the recognition performance. The method combines the facial action coding system(FACS) and 'uniform' local binary patterns(LBP) to represent facial expression features from coarse to fine. The facial feature regions are extracted by active shape models(ASM) based on FACS to obtain the gray-level texture. Then, LBP is used to represent expression features for enhancing the discriminant. A facial expression recognition system is developed based on this feature extraction method by using K nearest neighborhood(K-NN) classifier to recognize facial expressions. Finally, experiments are carried out to evaluate this feature extraction method. The significance of removing the unrelated facial regions and enhancing the discrimination ability of expression features in the recognition process is indicated by the results, in addition to its convenience.
基金supported by the National Science&Technology Support Plan of China(No.2009BAB48B02)the Basic Scientific Research Foundation for Institution of Higher Education(No.2008AA062101)
摘要Bubble seed image filling is an important prerequisite for the image segmentation of flotation bubble that can be used to improve flotation automatic control.These common image filling algorithms in dealing with complex bubble image exists under-filling and over-filling problems.A new filling algorithm based on boundary point feature and scan lines(PFSL)is proposed in the paper.The filling a|gorithm describes these boundary points of image objects by means of chain codes.The features of each boundary point,including convex points,concave points,left points and right points,are defined by the point's entrancing chain code and leaving chain code.The algorithm firstly finds out all double-matched boundary points based on the features of boundary points,and fill image objects by these double-matched boundary points on scan lines.Experimental results of bubble seed image filling show that under-filling and over-filling problem can be eliminated by the proposed algorithm.
基金This project was supported by the National Science Foundation of China (60763009)China Postdoctoral Science Foundation (2005038041)Hainan Natural Science Foundation (80528).
摘要Two signature systems based on smart cards and fingerprint features are proposed. In one signature system, the cryptographic key is stored in the smart card and is only accessible when the signer's extracted fingerprint features match his stored template. To resist being tampered on public channel, the user's message and the signed message are encrypted by the signer's public key and the user's public key, respectively. In the other signature system, the keys are generated by combining the signer's fingerprint features, check bits, and a rememberable key, and there are no matching process and keys stored on the smart card. Additionally, there is generally more than one public key in this system, that is, there exist some pseudo public keys except a real one.
基金National Natural Science Foundation of China(No.60872065)the Key Laboratory of Textile Science&Technology,Ministry of Education,China(No.P1111)+1 种基金the Key Laboratory of Advanced Textile Materials and Manufacturing Technology,Ministry of Education,China(No.2010001)the Priority Academic Program Development of Jiangsu Higher Education Institution,China
摘要To extract features of fabric defects effectively and reduce dimension of feature space,a feature extraction method of fabric defects based on complex contourlet transform (CCT) and principal component analysis (PCA) is proposed.Firstly,training samples of fabric defect images are decomposed by CCT.Secondly,PCA is applied in the obtained low-frequency component and part of highfrequency components to get a lower dimensional feature space.Finally,components of testing samples obtained by CCT are projected onto the feature space where different types of fabric defects are distinguished by the minimum Euclidean distance method.A large number of experimental results show that,compared with PCA,the method combining wavdet low-frequency component with PCA (WLPCA),the method combining contourlet transform with PCA (CPCA),and the method combining wavelet low-frequency and highfrequency components with PCA (WPCA),the proposed method can extract features of common fabric defect types effectively.The recognition rate is greatly improved while the dimension is reduced.
基金This work is supported by the Fundamental Research Funds for the Central Universities(Grant No.2572019BH03).
摘要Stance detection is the task of attitude identification toward a standpoint.Previous work of stance detection has focused on feature extraction but ignored the fact that irrelevant features exist as noise during higher-level abstracting.Moreover,because the target is not always mentioned in the text,most methods have ignored target information.In order to solve these problems,we propose a neural network ensemble method that combines the timing dependence bases on long short-term memory(LSTM)and the excellent extracting performance of convolutional neural networks(CNNs).The method can obtain multi-level features that consider both local and global features.We also introduce attention mechanisms to magnify target information-related features.Furthermore,we employ sparse coding to remove noise to obtain characteristic features.Performance was improved by using sparse coding on the basis of attention employment and feature extraction.We evaluate our approach on the SemEval-2016Task 6-A public dataset,achieving a performance that exceeds the benchmark and those of participating teams.
摘要In recent years, the accuracy of speech recognition (SR) has been one of the most active areas of research. Despite that SR systems are working reasonably well in quiet conditions, they still suffer severe performance degradation in noisy conditions or distorted channels. It is necessary to search for more robust feature extraction methods to gain better performance in adverse conditions. This paper investigates the performance of conventional and new hybrid speech feature extraction algorithms of Mel Frequency Cepstrum Coefficient (MFCC), Linear Prediction Coding Coefficient (LPCC), perceptual linear production (PLP), and RASTA-PLP in noisy conditions through using multivariate Hidden Markov Model (HMM) classifier. The behavior of the proposal system is evaluated using TIDIGIT human voice dataset corpora, recorded from 208 different adult speakers in both training and testing process. The theoretical basis for speech processing and classifier procedures were presented, and the recognition results were obtained based on word recognition rate.
摘要Wake-Up-Word Speech Recognition task (WUW-SR) is a computationally very demand, particularly the stage of feature extraction which is decoded with corresponding Hidden Markov Models (HMMs) in the back-end stage of the WUW-SR. The state of the art WUW-SR system is based on three different sets of features: Mel-Frequency Cepstral Coefficients (MFCC), Linear Predictive Coding Coefficients (LPC), and Enhanced Mel-Frequency Cepstral Coefficients (ENH_MFCC). In (front-end of Wake-Up-Word Speech Recognition System Design on FPGA) [1], we presented an experimental FPGA design and implementation of a novel architecture of a real-time spectrogram extraction processor that generates MFCC, LPC, and ENH_MFCC spectrograms simultaneously. In this paper, the details of converting the three sets of spectrograms 1) Mel-Frequency Cepstral Coefficients (MFCC), 2) Linear Predictive Coding Coefficients (LPC), and 3) Enhanced Mel-Frequency Cepstral Coefficients (ENH_MFCC) to their equivalent features are presented. In the WUW- SR system, the recognizer’s frontend is located at the terminal which is typically connected over a data network to remote back-end recognition (e.g., server). The WUW-SR is shown in Figure 1. The three sets of speech features are extracted at the front-end. These extracted features are then compressed and transmitted to the server via a dedicated channel, where subsequently they are decoded.
摘要近年来,随着物联网(Internet of Things,IoT)、语义通信以及智慧城市等经典机器间通信(Machine to Machine,M2M)场景的快速发展,海量视觉数据在设备间的实时传输与高效处理成为了一项关键挑战。在此背景下,传统以人眼感知质量为核心的图像编码方法,因其优化目标与机器视觉任务需求存在本质差异,往往在面向机器视觉分析时出现分析精度不足的问题。为此,面向机器视觉的图像编码(Image Coding for Machine,ICM)应运而生,其核心目标是在保证下游机器视觉任务(如分类、检测、分割等)分析精度的同时,实现尽可能低的编码码率,从而更好地适配M2M场景中的带宽与存储约束。然而,现有ICM方法仍面临两大瓶颈:其一,在极低码率条件下性能急剧下降。这是由于现有方法多依赖于端到端的非线性变换提取视觉特征,未能充分挖掘和利用图像中高层语义信息的紧凑表示,导致特征编码效率不足;其二,在开放场景下的泛化能力弱。多数方法针对单一任务、单一数据集进行优化,缺乏对未知类别、跨域数据的适应能力,难以在实际动态环境中保持稳定的分析性能。为突破上述限制,本文提出一种文本提示引导的面向机器视觉图像编码框架(Text-prompted Image Coding for Machine,T-ICM)。该框架的核心思想是将图像信息解耦为语义信息与纹理信息两个互补的组成部分,其中,语义信息以结构化文本提示(如对象类别、位置描述)的形式进行表示与编码,纹理信息则通过一种任务无关的通用视觉特征进行提取与压缩。在编码端,文本提示因其高度抽象和语义紧凑的特性,可以显著降低整体码率;通用特征则通过我们提出的分组特征编码模块进行高效压缩。在解码端,文本提示不仅用于直接解析完成分类、检测等任务,更重要的是作为引导信号,通过提示编码器与掩膜解码器,动态调整重建通用特征的语义感知区域,实现特征层面的域自适应与任务适配,从而显著提升模型在开放场景下的鲁棒性。本文在多个标准数据集与任务上对T-ICM进行了全面评估。实验表明,在语义分割和实例分割等密集预测任务上,T-ICM在极低码率下仍能保持接近原始图像输入的分析精度,其性能显著优于H.266/VVC、基于深度学习的图像编码器以及现有的其他ICM方法。本研究通过将语义信息迁移至高度压缩的文本模态进行传输,并利用其引导特征重建,T-ICM在编码效率与任务性能之间实现了更优的权衡,为未来语义通信、边缘智能协同,以及自适应机器视觉系统的发展提供了新的思路与技术支撑。