Non-negative Matrix Factorization(NMF)is a computationally intensive matrix operation that resource-constrained clients struggle to complete locally.Privacy-preserving outsourcing allows clients to offload heavy compu...Non-negative Matrix Factorization(NMF)is a computationally intensive matrix operation that resource-constrained clients struggle to complete locally.Privacy-preserving outsourcing allows clients to offload heavy computing tasks to powerful servers,effectively solving the problem of local computing difficulties.However,the existing privacy-preserving NMF outsourcing schemes only allow one server to perform outsourcing computation,resulting in low efficiency on the server side.In order to improve the efficiency of outsourcing computation,we propose a privacy-preserving parallel NMF outsourcing scheme with multiple edge servers.We adopt the matrix blocking technique to divide the computation task into multiple subtasks,and design the NMF parallel computation algorithm based on the multiplication updating rule.The proposed scheme implements the parallel outsourcing of non-negative matrix factorization based on multiple edge servers.We use random permutation matrices to encrypt original matrix,thereby protecting data privacy.In addition,we utilize the iterative nature of the NMF algorithm for result verification.Theoretical analysis and experimental results prove the advantages of the proposed scheme.展开更多
The increasing connectivity of modern vehicles exposes the in-vehicle controller area network(CAN)bus to various cyberattacks,including denial-of-service,fuzzy injection,and spoofing attacks.Existing machine learning ...The increasing connectivity of modern vehicles exposes the in-vehicle controller area network(CAN)bus to various cyberattacks,including denial-of-service,fuzzy injection,and spoofing attacks.Existing machine learning and deep learning intrusion detection systems(IDS)often rely on labeled data,struggle with class imbalance,lack interpretability,and fail to generalize well across different datasets.This paper proposes a lightweight and interpretable IDS framework based on non-negative matrix factorization(NMF)to address these limitations.Our contributions include:(i)evaluating NMF as both a standalone unsupervised detector and an interpretable feature extractor(NMF-W)for classical,unsupervised,and deep sequence models;(ii)providing comprehensive benchmarking on the car-hacking dataset(CHD),demonstrating improved robustness in mixed-attack and cross-attack scenarios,with class imbalance addressed through oversampling and class weighting;(iii)offering a component-level interpretability analysis that links NMF factors to meaningful CAN traffic patterns;and(iv)validating cross-dataset transferability on the offset-ratio and time interval-based intrusion detection system(OTIDS)dataset.Additional ablation and efficiency studies confirm the practical feasibility of deploying NMF-based IDS on embedded automotive controllers.Overall,this work presents a balanced IDS solution that combines detection accuracy,computational efficiency,and explainability,thereby advancing the security of in-vehicle networks.展开更多
Azimuth ambiguity significantly degrades the quality of synthetic aperture radar images.Sub-look spectral analysis(SSA)is a common ambiguity-detection method,but its performance is limited by threshold sensitivity and...Azimuth ambiguity significantly degrades the quality of synthetic aperture radar images.Sub-look spectral analysis(SSA)is a common ambiguity-detection method,but its performance is limited by threshold sensitivity and the high correlation of specific ambiguities across sub-looks.To overcome these specific limitations,this paper proposes an improved detection method.It first increases the number of sub-looks and constructs a high-dimensional multi-look matrix to enrich the coherence differences between targets and ambiguities.Non-negative matrix factorization is then employed to decompose this matrix,effectively separating the coherent target components from the variably coherent ambiguity components without relying on predefined thresholds.Experimental results on real data demonstrate that the proposed improvements achieve superior azimuth-ambiguity-detection performance compared with conventional SSA methods.展开更多
CircRNAs,widely found throughout the human bodies,play a crucial role in regulating various biological processes and are closely linked to complex human diseases.Investigating potential associations between circRNAs a...CircRNAs,widely found throughout the human bodies,play a crucial role in regulating various biological processes and are closely linked to complex human diseases.Investigating potential associations between circRNAs and diseases can enhance our understanding of diseases and provide new strategies and tools for early diagnosis,treatment,and disease prevention.However,existing models have limitations in accurately capturing similarities,handling the sparse and noise attributes of association networks,and fully leveraging bioinformatical aspects from multiple viewpoints.To address these issues,this study introduces a new non-negative matrix factorization-based framework called NMFMSN.First,we incorporate circRNA sequence data and disease semantic information to compute circRNA and disease similarity,respectively.Given the sparse known associations between circRNAs and diseases,we reconstruct the network to complete more associations by imputing missing links based on neighboring circRNA and disease interactions.Finally,we integrate these two similarity networks into a non-negative matrix factorization framework to identify potential circRNA-disease associations.Upon conducting 5-fold cross-validation and leave-one-out cross-validation,the AUC values for NMFMSN reach 0.9712 and 0.9768,respectively,outperforming the currently most advanced models.Case studies on lung cancer and hepatocellular carcinoma show that NMFMSN is a good way to predict new associations between circRNAs and diseases.展开更多
Fine particulatematter(PM2.5)samples were collected in two neighboring cities,Beijing and Baoding,China.High-concentration events of PM2.5 in which the average mass concentration exceeded 75μg/m3 were freque...Fine particulatematter(PM2.5)samples were collected in two neighboring cities,Beijing and Baoding,China.High-concentration events of PM2.5 in which the average mass concentration exceeded 75μg/m3 were frequently observed during the heating season.Dispersion Normalized Positive Matrix Factorization was applied for the source apportionment of PM2.5 as minimize the dilution effects of meteorology and better reflect the source strengths in these two cities.Secondary nitrate had the highest contribution for Beijing(37.3%),and residential heating/biomass burning was the largest for Baoding(27.1%).Secondary nitrate,mobile,biomass burning,district heating,oil combustion,aged sea salt sources showed significant differences between the heating and non-heating seasons in Beijing for same period(2019.01.10–2019.08.22)(Mann-Whitney Rank Sum Test P<0.05).In case of Baoding,soil,residential heating/biomass burning,incinerator,coal combustion,oil combustion sources showed significant differences.The results of Pearson correlation analysis for the common sources between the two cities showed that long-range transported sources and some sources with seasonal patterns such as oil combustion and soil had high correlation coefficients.Conditional Bivariate Probability Function(CBPF)was used to identify the inflow directions for the sources,and joint-PSCF(Potential Source Contribution Function)was performed to determine the common potential source areas for sources affecting both cities.These models facilitated a more precise verification of city-specific influences on PM2.5 sources.The results of this study will aid in prioritizing air pollution mitigation strategies during the heating season and strengthening air quality management to reduce the impact of downwind neighboring cities.展开更多
Substantial effects of photochemical reaction losses of volatile organic compounds(VOCs)on factor profiles can be investigated by comparing the differences between daytime and nighttime dispersion-normalized VOC data ...Substantial effects of photochemical reaction losses of volatile organic compounds(VOCs)on factor profiles can be investigated by comparing the differences between daytime and nighttime dispersion-normalized VOC data resolved profiles.Hourly speciated VOC data measured in Shijiazhuang,China from May to September 2021 were used to conduct study.The mean VOC concentration in the daytime and at nighttime were 32.8 and 36.0 ppbv,respectively.Alkanes and aromatics concentrations in the daytime(12.9 and 3.08 ppbv)were lower than nighttime(15.5 and 3.63 ppbv),whereas that of alkenes showed the opposite tendency.The concentration differences between daytime and nighttime for alkynes and halogenated hydrocarbonswere uniformly small.The reactivities of the dominant species in factor profiles for gasoline emissions,natural gas and diesel vehicles,and liquefied petroleum gas were relatively low and their profiles were less affected by photochemical losses.Photochemical losses produced a substantial impact on the profiles of solvent use,petrochemical industry emissions,combustion sources,and biogenic emissions where the dominant species in these factor profiles had high reactivities.Although the profile of biogenic emissions was substantially affected by photochemical loss of isoprene,the low emissions at nighttime also had an important impact on its profile.Chemical losses of highly active VOC species substantially reduced their concentrations in apportioned factor profiles.This study results were consistent with the analytical results obtained through initial concentration estimation,suggesting that the initial concentration estimation could be the most effective currently availablemethod for the source analyses of active VOCs although with uncertainty.展开更多
High-dimensional and incomplete(HDI) matrices are commonly encountered in various big data-related applications for illustrating the complex interactions among numerous entities, like the user-item interactions in a c...High-dimensional and incomplete(HDI) matrices are commonly encountered in various big data-related applications for illustrating the complex interactions among numerous entities, like the user-item interactions in a commercial recommender system or the user-user interactions in a social network services system. The factorization of such an HDI matrix can embed the involved entities into the low-dimensional feature space for acquiring their principal representation, which is a vital task in various application scenes and is often established through the Latent Factor Analysis(LFA). Nevertheless, an HDI matrix can be huge when the corresponding application explodes to involve millions of users, items, or other interactive nodes. In this case, a parallel optimization algorithm is desired for raising the scalability and time efficiency of an LFA model. This paper provides a comprehensive review of the existing parallel optimization algorithms for the LFA model. Specifically, it performs: 1) discussion and summary of these algorithms based on computing architecture and mode, 2) empirical studies of representative models, and3) summary of the current challenges and future directions in this domain. This survey aims to offer an exhaustive review of Parallel Optimization Algorithms for High-Dimensional and Incomplete Matrix Factorization, thereby fostering further research in this field.展开更多
Due to the non-stationary characteristics of vibration signals acquired from rolling element bearing fault, thc time-frequency analysis is often applied to describe the local information of these unstable signals smar...Due to the non-stationary characteristics of vibration signals acquired from rolling element bearing fault, thc time-frequency analysis is often applied to describe the local information of these unstable signals smartly. However, it is difficult to classitythe high dimensional feature matrix directly because of too large dimensions for many classifiers. This paper combines the concepts of time-frequency distribution(TFD) with non-negative matrix factorization(NMF), and proposes a novel TFD matrix factorization method to enhance representation and identification of bearing fault. Throughout this method, the TFD of a vibration signal is firstly accomplished to describe the localized faults with short-time Fourier transform(STFT). Then, the supervised NMF mapping is adopted to extract the fault features from TFD. Meanwhile, the fault samples can be clustered and recognized automatically by using the clustering property of NMF. The proposed method takes advantages of the NMF in the parts-based representation and the adaptive clustering. The localized fault features of interest can be extracted as well. To evaluate the performance of the proposed method, the 9 kinds of the bearing fault on a test bench is performed. The proposed method can effectively identify the fault severity and different fault types. Moreover, in comparison with the artificial neural network(ANN), NMF yields 99.3% mean accuracy which is much superior to ANN. This research presents a simple and practical resolution for the fault diagnosis problem of rolling element bearing in high dimensional feature space.展开更多
This study aimed to investigate the pollution characteristics, source apportionment, and health risks associated with trace metal(loid)s(TMs) in the major agricultural producing areas in Chongqing, China. We analyzed ...This study aimed to investigate the pollution characteristics, source apportionment, and health risks associated with trace metal(loid)s(TMs) in the major agricultural producing areas in Chongqing, China. We analyzed the source apportionment and assessed the health risk of TMs in agricultural soils by using positive matrix factorization(PMF) model and health risk assessment(HRA) model based on Monte Carlo simulation. Meanwhile, we combined PMF and HRA models to explore the health risks of TMs in agricultural soils by different pollution sources to determine the priority control factors. Results showed that the average contents of cadmium(Cd), arsenic (As), lead(Pb), chromium(Cr), copper(Cu), nickel(Ni), and zinc(Zn) in the soil were found to be 0.26, 5.93, 27.14, 61.32, 23.81, 32.45, and 78.65 mg/kg, respectively. Spatial analysis and source apportionment analysis revealed that urban and industrial sources, agricultural sources, and natural sources accounted for 33.0%, 27.7%, and 39.3% of TM accumulation in the soil, respectively. In the HRA model based on Monte Carlo simulation, noncarcinogenic risks were deemed negligible(hazard index <1), the carcinogenic risks were at acceptable level(10-6<total carcinogenic risk ≤ 10-4), with higher risks observed for children compared to adults. The relationship between TMs, their sources, and health risks indicated that urban and industrial sources were primarily associated with As, contributing to 75.1% of carcinogenic risks and 55.7% of non-carcinogenic risks, making them the primary control factors. Meanwhile, agricultural sources were primarily linked to Cd and Pb, contributing to 13.1% of carcinogenic risks and 21.8% of non-carcinogenic risks, designating them as secondary control factors.展开更多
The constrained weighted-non-negative matrix factorization(CW-NMF)hybrid receptor model was applied to study the influence of steelmaking activities on PM2.5(particulate matter with equivalent aerodynamic diameter ...The constrained weighted-non-negative matrix factorization(CW-NMF)hybrid receptor model was applied to study the influence of steelmaking activities on PM2.5(particulate matter with equivalent aerodynamic diameter less than 2.5μm)composition in Dunkerque,Northern France.Semi-diurnal PM2.5samples were collected using a high volume sampler in winter 2010 and spring 2011 and were analyzed for trace metals,water-soluble ions,and total carbon using inductively coupled plasma–atomic emission spectrometry(ICP-AES),ICP-mass spectrometry(ICP-MS),ionic chromatography and micro elemental carbon analyzer.The elemental composition shows that NO3-,SO42-,NH_4~+and total carbon are the main PM2.5constituents.Trace metals data were interpreted using concentration roses and both influences of integrated steelworks and electric steel plant were evidenced.The distinction between the two sources is made possible by the use Zn/Fe and Zn/Mn diagnostic ratios.Moreover Rb/Cr,Pb/Cr and Cu/Cd combination ratio are proposed to distinguish the ISW-sintering stack from the ISW-fugitive emissions.The a priori knowledge on the influencing source was introduced in the CW-NMF to guide the calculation.Eleven source profiles with various contributions were identified:8 are characteristics of coastal urban background site profiles and 3 are related to the steelmaking activities.Between them,secondary nitrates,secondary sulfates and combustion profiles give the highest contributions and account for 93%of the PM2.5concentration.The steelwork facilities contribute in about 2%of the total PM2.5concentration and appear to be the main source of Cr,Cu,Fe,Mn,Zn.展开更多
This paper presents a novel medical image registration algorithm named total variation constrained graphregularization for non-negative matrix factorization(TV-GNMF).The method utilizes non-negative matrix factorizati...This paper presents a novel medical image registration algorithm named total variation constrained graphregularization for non-negative matrix factorization(TV-GNMF).The method utilizes non-negative matrix factorization by total variation constraint and graph regularization.The main contributions of our work are the following.First,total variation is incorporated into NMF to control the diffusion speed.The purpose is to denoise in smooth regions and preserve features or details of the data in edge regions by using a diffusion coefficient based on gradient information.Second,we add graph regularization into NMF to reveal intrinsic geometry and structure information of features to enhance the discrimination power.Third,the multiplicative update rules and proof of convergence of the TV-GNMF algorithm are given.Experiments conducted on datasets show that the proposed TV-GNMF method outperforms other state-of-the-art algorithms.展开更多
A current problem in diet recommendation systems is the matching of food preferences with nutritional requirements,taking into account individual characteristics,such as body weight with individual health conditions,s...A current problem in diet recommendation systems is the matching of food preferences with nutritional requirements,taking into account individual characteristics,such as body weight with individual health conditions,such as diabetes.Current dietary recommendations employ association rules,content-based collaborative filtering,and constraint-based methods,which have several limitations.These limitations are due to the existence of a special user group and an imbalance of non-simple attributes.Making use of traditional dietary recommendation algorithm researches,we combine the Adaboost classifier with probabilistic matrix factorization.We present a personalized diet recommendation algorithm by taking advantage of probabilistic matrix factorization via Adaboost.A probabilistic matrix factorization method extracts the implicit factors between individual food preferences and nutritional characteristics.From this,we can make use of those features with strong influence while discarding those with little influence.After incorporating these changes into our approach,we evaluated our algorithm’s performance.Our results show that our method performed better than others at matching preferred foods with dietary requirements,benefiting user health as a result.The algorithm fully considers the constraint relationship between users’attributes and nutritional characteristics of foods.Considering many complex factors in our algorithm,the recommended food result set meets both health standards and users’dietary preferences.A comparison of our algorithm with others demonstrated that our method offers high accuracy and interpretability.展开更多
Nonnegative matrix factorization (NMF) is a method to get parts-based features of information and form the typical profiles. But the basis vectors NMF gets are not orthogonal so that parts-based features of informatio...Nonnegative matrix factorization (NMF) is a method to get parts-based features of information and form the typical profiles. But the basis vectors NMF gets are not orthogonal so that parts-based features of information are usually redundancy. In this paper, we propose two different approaches based on localized non-negative matrix factorization (LNMF) to obtain the typical user session profiles and typical semantic profiles of junk mails. The LNMF get basis vectors as orthogonal as possible so that it can get accurate profiles. The experiments show that the approach based on LNMF can obtain better profiles than the approach based on NMF. Key words localized non-negative matrix factorization - profile - log mining - mail filtering CLC number TP 391 Foundation item: Supported by the National Natural Science Foundation of China (60373066, 60303024), National Grand Fundamental Research 973 Program of China (2002CB312000), National Research Foundation for the Doctoral Program of Higher Education of China (20020286004).Biography: Jiang Ji-xiang (1980-), male, Master candidate, research direction: data mining, knowledge representation on the Web.展开更多
Currently,functional connectomes constructed from neuroimaging data have emerged as a powerful tool in identifying brain disorders.If one brain disease just manifests as some cognitive dysfunction,it means that the di...Currently,functional connectomes constructed from neuroimaging data have emerged as a powerful tool in identifying brain disorders.If one brain disease just manifests as some cognitive dysfunction,it means that the disease may affect some local connectivity in the brain functional network.That is,there are functional abnormalities in the sub-network.Therefore,it is crucial to accurately identify them in pathological diagnosis.To solve these problems,we proposed a sub-network extraction method based on graph regularization nonnegative matrix factorization(GNMF).The dynamic functional networks of normal subjects and early mild cognitive impairment(eMCI)subjects were vectorized and the functional connection vectors(FCV)were assembled to aggregation matrices.Then GNMF was applied to factorize the aggregation matrix to get the base matrix,in which the column vectors were restored to a common sub-network and a distinctive sub-network,and visualization and statistical analysis were conducted on the two sub-networks,respectively.Experimental results demonstrated that,compared with other matrix factorization methods,the proposed method can more obviously reflect the similarity between the common subnetwork of eMCI subjects and normal subjects,as well as the difference between the distinctive sub-network of eMCI subjects and normal subjects,Therefore,the high-dimensional features in brain functional networks can be best represented locally in the lowdimensional space,which provides a new idea for studying brain functional connectomes.展开更多
This paper considers a problem of unsupervised spectral unmixing of hyperspectral data. Based on the Linear Mixing Model ( LMM), a new method under the framework of nonnegative matrix fac- torization (NMF) is prop...This paper considers a problem of unsupervised spectral unmixing of hyperspectral data. Based on the Linear Mixing Model ( LMM), a new method under the framework of nonnegative matrix fac- torization (NMF) is proposed, namely minimum distance constrained nonnegative matrix factoriza- tion (MDC-NMF). In this paper, firstly, a new regularization term, called endmember distance (ED) is considered, which is defined as the sum of the squared Euclidean distances from each end- member to their geometric center. Compared with the simplex volume, ED has better optimization properties and is conceptually intuitive. Secondly, a projected gradient (PG) scheme is adopted, and by the virtue of ED, in this scheme the optimal step size along the feasible descent direction can be calculated easily at each iteration. Thirdly, a finite step ( no more than the number of endmem- bers) terminated algorithm is used to project a point on the canonical simplex, by which the abun- dance nonnegative constraint and abundance sum-to-one constraint can be accurately satisfied in a light amount of computation. The experimental results, based on a set of synthetic data and real da- ta, demonstrate that, in the same running time, MDC-NMF outperforms several other similar meth- ods proposed recently.展开更多
An image fusion method combining complex contourlet transform(CCT) with nonnegative matrix factorization(NMF) is proposed in this paper.After two images are decomposed by CCT,NMF is applied to their highand low-freque...An image fusion method combining complex contourlet transform(CCT) with nonnegative matrix factorization(NMF) is proposed in this paper.After two images are decomposed by CCT,NMF is applied to their highand low-frequency components,respectively,and finally an image is synthesized.Subjective-visual-quality of the image fusion result is compared with those of the image fusion methods based on NMF and the combination of wavelet /contourlet onsubsampled contourlet with NMF.The experimental results are evaluated quantitatively,and the running time is also contrasted.It is shown that the proposed image fusion method can gain larger information entropy,standard deviation and mean gradient,which means that it can better integrate featured information from all source images,avoid background noise and promote space clearness in the fusion image effectively.展开更多
Link prediction has attracted wide attention among interdisciplinaryresearchers as an important issue in complex network. It aims to predict the missing links in current networks and new links that will appear in futu...Link prediction has attracted wide attention among interdisciplinaryresearchers as an important issue in complex network. It aims to predict the missing links in current networks and new links that will appear in future networks.Despite the presence of missing links in the target network of link prediction studies, the network it processes remains macroscopically as a large connectedgraph. However, the complexity of the real world makes the complex networksabstracted from real systems often contain many isolated nodes. This phenomenon leads to existing link prediction methods not to efficiently implement the prediction of missing edges on isolated nodes. Therefore, the cold-start linkprediction is favored as one of the most valuable subproblems of traditional linkprediction. However, due to the loss of many links in the observation network, thetopological information available for completing the link prediction task is extremely scarce. This presents a severe challenge for the study of cold-start link prediction. Therefore, how to mine and fuse more available non-topologicalinformation from observed network becomes the key point to solve the problemof cold-start link prediction. In this paper, we propose a framework for solving thecold-start link prediction problem, a joint-weighted symmetric nonnegative matrixfactorization model fusing graph regularization information, based on low-rankapproximation algorithms in the field of machine learning. First, the nonlinear features in high-dimensional space of node attributes are captured by the designedgraph regularization term. Second, using a weighted matrix, we associate the attribute similarity and first order structure information of nodes and constrain eachother. Finally, a unified framework for implementing cold-start link prediction isconstructed by using a symmetric nonnegative matrix factorization model to integrate the multiple information extracted together. Extensive experimental validationon five real networks with attributes shows that the proposed model has very goodpredictive performance when predicting missing edges of isolated nodes.展开更多
Contrastive learning is a significant research direction in the field of deep learning.However,existing data augmentation methods often lead to issues such as semantic drift in generated views while the complexity of ...Contrastive learning is a significant research direction in the field of deep learning.However,existing data augmentation methods often lead to issues such as semantic drift in generated views while the complexity of model pre-training limits further improvement in the performance of existing methods.To address these challenges,we propose the Efficient Clustering Network based on Matrix Factorization(ECN-MF).Specifically,we design a batched low-rank Singular Value Decomposition(SVD)algorithm for data augmentation to eliminate redundant information and uncover major patterns of variation and key information in the data.Additionally,we design a Mutual Information-Enhanced Clustering Module(MI-ECM)to accelerate the training process by leveraging a simple architecture to bring samples from the same cluster closer while pushing samples from other clusters apart.Extensive experiments on six datasets demonstrate that ECN-MF exhibits more effective performance compared to state-of-the-art algorithms.展开更多
Non-negative matrix factorization (NMF) is a technique for dimensionality reduction by placing non-negativity constraints on the matrix. Based on the PARAFAC model, NMF was extended for three-dimension data decompos...Non-negative matrix factorization (NMF) is a technique for dimensionality reduction by placing non-negativity constraints on the matrix. Based on the PARAFAC model, NMF was extended for three-dimension data decomposition. The three-dimension nonnegative matrix factorization (NMF3) algorithm, which was concise and easy to implement, was given in this paper. The NMF3 algorithm implementation was based on elements but not on vectors. It could decompose a data array directly without unfolding, which was not similar to that the traditional algorithms do, It has been applied to the simulated data array decomposition and obtained reasonable results. It showed that NMF3 could be introduced for curve resolution in chemometrics.展开更多
One of the most important problems in complex networks is to identify the influential vertices for understanding and controlling of information diffusion and disease spreading.Most of the current centrality algorithms...One of the most important problems in complex networks is to identify the influential vertices for understanding and controlling of information diffusion and disease spreading.Most of the current centrality algorithms focus on single feature or manually extract the attributes,which occasionally results in the failure to fully capture the vertex’s importance.A new vertex centrality approach based on symmetric nonnegative matrix factorization(SNMF),called VCSNMF,is proposed in this paper.For highlight the characteristics of a network,the adjacency matrix and the degree matrix are fused to represent original data of the network via a weighted linear combination.First,SNMF automatically extracts the latent characteristics of vertices by factorizing the established original data matrix.Then we prove that each vertex’s composite feature which is constructed with one-dimensional factor matrix can be approximated as the term of eigenvector associated with the spectral radius of the network,otherwise obtained by the factor matrix on the hyperspace.Finally,VCSNMF integrates the composite feature and the topological structure to evaluate the performance of vertices.To verify the effectiveness of the VCSNMF criterion,eight existing centrality approaches are used as comparison measures to rank influential vertices in ten real-world networks.The experimental results assert the superiority of the method.展开更多
基金supported in part by Shandong Provincial Natural Science Foundation under Grant(ZR2024MF038)Qingdao Natural Science Foundation(25-1-1-103-zyyd-jchZ).
摘要Non-negative Matrix Factorization(NMF)is a computationally intensive matrix operation that resource-constrained clients struggle to complete locally.Privacy-preserving outsourcing allows clients to offload heavy computing tasks to powerful servers,effectively solving the problem of local computing difficulties.However,the existing privacy-preserving NMF outsourcing schemes only allow one server to perform outsourcing computation,resulting in low efficiency on the server side.In order to improve the efficiency of outsourcing computation,we propose a privacy-preserving parallel NMF outsourcing scheme with multiple edge servers.We adopt the matrix blocking technique to divide the computation task into multiple subtasks,and design the NMF parallel computation algorithm based on the multiplication updating rule.The proposed scheme implements the parallel outsourcing of non-negative matrix factorization based on multiple edge servers.We use random permutation matrices to encrypt original matrix,thereby protecting data privacy.In addition,we utilize the iterative nature of the NMF algorithm for result verification.Theoretical analysis and experimental results prove the advantages of the proposed scheme.
基金supported in part by the Basic Science Research Program through the NRF funded by the Ministry of Education under Grant 2021R1A6A1A03039493in part by the Regional Innovation System&Education(RISE)program through the Gyeongbuk RISE CENTER,funded by the Ministry of Education(MOE)and the Gyeongsangbuk-do,Republic of Korea(2025-RISE-15-115).
摘要The increasing connectivity of modern vehicles exposes the in-vehicle controller area network(CAN)bus to various cyberattacks,including denial-of-service,fuzzy injection,and spoofing attacks.Existing machine learning and deep learning intrusion detection systems(IDS)often rely on labeled data,struggle with class imbalance,lack interpretability,and fail to generalize well across different datasets.This paper proposes a lightweight and interpretable IDS framework based on non-negative matrix factorization(NMF)to address these limitations.Our contributions include:(i)evaluating NMF as both a standalone unsupervised detector and an interpretable feature extractor(NMF-W)for classical,unsupervised,and deep sequence models;(ii)providing comprehensive benchmarking on the car-hacking dataset(CHD),demonstrating improved robustness in mixed-attack and cross-attack scenarios,with class imbalance addressed through oversampling and class weighting;(iii)offering a component-level interpretability analysis that links NMF factors to meaningful CAN traffic patterns;and(iv)validating cross-dataset transferability on the offset-ratio and time interval-based intrusion detection system(OTIDS)dataset.Additional ablation and efficiency studies confirm the practical feasibility of deploying NMF-based IDS on embedded automotive controllers.Overall,this work presents a balanced IDS solution that combines detection accuracy,computational efficiency,and explainability,thereby advancing the security of in-vehicle networks.
基金supported by the National Natural Science Foundation of China(62271408)Shanghai Aerospace Science Technology Innovation Fund(SAST2024-024)the Innovation Foundation for Doctor Dissertation of Northwestern Polytechnical University(CX2024065)。
摘要Azimuth ambiguity significantly degrades the quality of synthetic aperture radar images.Sub-look spectral analysis(SSA)is a common ambiguity-detection method,but its performance is limited by threshold sensitivity and the high correlation of specific ambiguities across sub-looks.To overcome these specific limitations,this paper proposes an improved detection method.It first increases the number of sub-looks and constructs a high-dimensional multi-look matrix to enrich the coherence differences between targets and ambiguities.Non-negative matrix factorization is then employed to decompose this matrix,effectively separating the coherent target components from the variably coherent ambiguity components without relying on predefined thresholds.Experimental results on real data demonstrate that the proposed improvements achieve superior azimuth-ambiguity-detection performance compared with conventional SSA methods.
基金the Gansu Province Industrial Support Plan(No.2023CYZC-25)Natural Science Foundation of Gansu Province(No.23JRRA770)the National Natural Science Foundation of China(No.62162040)。
摘要CircRNAs,widely found throughout the human bodies,play a crucial role in regulating various biological processes and are closely linked to complex human diseases.Investigating potential associations between circRNAs and diseases can enhance our understanding of diseases and provide new strategies and tools for early diagnosis,treatment,and disease prevention.However,existing models have limitations in accurately capturing similarities,handling the sparse and noise attributes of association networks,and fully leveraging bioinformatical aspects from multiple viewpoints.To address these issues,this study introduces a new non-negative matrix factorization-based framework called NMFMSN.First,we incorporate circRNA sequence data and disease semantic information to compute circRNA and disease similarity,respectively.Given the sparse known associations between circRNAs and diseases,we reconstruct the network to complete more associations by imputing missing links based on neighboring circRNA and disease interactions.Finally,we integrate these two similarity networks into a non-negative matrix factorization framework to identify potential circRNA-disease associations.Upon conducting 5-fold cross-validation and leave-one-out cross-validation,the AUC values for NMFMSN reach 0.9712 and 0.9768,respectively,outperforming the currently most advanced models.Case studies on lung cancer and hepatocellular carcinoma show that NMFMSN is a good way to predict new associations between circRNAs and diseases.
基金supported by the National Institute of Environmental Research(NIER)funded by the Ministry of Environment(No.NIER-2019-04-02-039)supported by Particulate Matter Management Specialized Graduate Program through the Korea Environmental Industry&Technology Institute(KEITI)funded by the Ministry of Environment(MOE).
摘要Fine particulatematter(PM2.5)samples were collected in two neighboring cities,Beijing and Baoding,China.High-concentration events of PM2.5 in which the average mass concentration exceeded 75μg/m3 were frequently observed during the heating season.Dispersion Normalized Positive Matrix Factorization was applied for the source apportionment of PM2.5 as minimize the dilution effects of meteorology and better reflect the source strengths in these two cities.Secondary nitrate had the highest contribution for Beijing(37.3%),and residential heating/biomass burning was the largest for Baoding(27.1%).Secondary nitrate,mobile,biomass burning,district heating,oil combustion,aged sea salt sources showed significant differences between the heating and non-heating seasons in Beijing for same period(2019.01.10–2019.08.22)(Mann-Whitney Rank Sum Test P<0.05).In case of Baoding,soil,residential heating/biomass burning,incinerator,coal combustion,oil combustion sources showed significant differences.The results of Pearson correlation analysis for the common sources between the two cities showed that long-range transported sources and some sources with seasonal patterns such as oil combustion and soil had high correlation coefficients.Conditional Bivariate Probability Function(CBPF)was used to identify the inflow directions for the sources,and joint-PSCF(Potential Source Contribution Function)was performed to determine the common potential source areas for sources affecting both cities.These models facilitated a more precise verification of city-specific influences on PM2.5 sources.The results of this study will aid in prioritizing air pollution mitigation strategies during the heating season and strengthening air quality management to reduce the impact of downwind neighboring cities.
基金supported by the National Key R&D Program of China(No.2023YFC3705801)the National Natural Science Foundation of China(No.42177085).
摘要Substantial effects of photochemical reaction losses of volatile organic compounds(VOCs)on factor profiles can be investigated by comparing the differences between daytime and nighttime dispersion-normalized VOC data resolved profiles.Hourly speciated VOC data measured in Shijiazhuang,China from May to September 2021 were used to conduct study.The mean VOC concentration in the daytime and at nighttime were 32.8 and 36.0 ppbv,respectively.Alkanes and aromatics concentrations in the daytime(12.9 and 3.08 ppbv)were lower than nighttime(15.5 and 3.63 ppbv),whereas that of alkenes showed the opposite tendency.The concentration differences between daytime and nighttime for alkynes and halogenated hydrocarbonswere uniformly small.The reactivities of the dominant species in factor profiles for gasoline emissions,natural gas and diesel vehicles,and liquefied petroleum gas were relatively low and their profiles were less affected by photochemical losses.Photochemical losses produced a substantial impact on the profiles of solvent use,petrochemical industry emissions,combustion sources,and biogenic emissions where the dominant species in these factor profiles had high reactivities.Although the profile of biogenic emissions was substantially affected by photochemical loss of isoprene,the low emissions at nighttime also had an important impact on its profile.Chemical losses of highly active VOC species substantially reduced their concentrations in apportioned factor profiles.This study results were consistent with the analytical results obtained through initial concentration estimation,suggesting that the initial concentration estimation could be the most effective currently availablemethod for the source analyses of active VOCs although with uncertainty.
基金supported in part by the National Key Research and Development Program of China(2024YFF0908200)the National Natural Science Foundation of China(62302402,62272078)+1 种基金the Chongqing Natural Science Foundation(CSTB2024TIAD-KPX0018,CSTB2023NSCO-LZX006)the Southwest University Graduate Research Innovation Project(SWUB24050)
摘要High-dimensional and incomplete(HDI) matrices are commonly encountered in various big data-related applications for illustrating the complex interactions among numerous entities, like the user-item interactions in a commercial recommender system or the user-user interactions in a social network services system. The factorization of such an HDI matrix can embed the involved entities into the low-dimensional feature space for acquiring their principal representation, which is a vital task in various application scenes and is often established through the Latent Factor Analysis(LFA). Nevertheless, an HDI matrix can be huge when the corresponding application explodes to involve millions of users, items, or other interactive nodes. In this case, a parallel optimization algorithm is desired for raising the scalability and time efficiency of an LFA model. This paper provides a comprehensive review of the existing parallel optimization algorithms for the LFA model. Specifically, it performs: 1) discussion and summary of these algorithms based on computing architecture and mode, 2) empirical studies of representative models, and3) summary of the current challenges and future directions in this domain. This survey aims to offer an exhaustive review of Parallel Optimization Algorithms for High-Dimensional and Incomplete Matrix Factorization, thereby fostering further research in this field.
基金Supported by Shaanxi Provincial Overall Innovation Project of Science and Technology,China(Grant No.2013KTCQ01-06)
摘要Due to the non-stationary characteristics of vibration signals acquired from rolling element bearing fault, thc time-frequency analysis is often applied to describe the local information of these unstable signals smartly. However, it is difficult to classitythe high dimensional feature matrix directly because of too large dimensions for many classifiers. This paper combines the concepts of time-frequency distribution(TFD) with non-negative matrix factorization(NMF), and proposes a novel TFD matrix factorization method to enhance representation and identification of bearing fault. Throughout this method, the TFD of a vibration signal is firstly accomplished to describe the localized faults with short-time Fourier transform(STFT). Then, the supervised NMF mapping is adopted to extract the fault features from TFD. Meanwhile, the fault samples can be clustered and recognized automatically by using the clustering property of NMF. The proposed method takes advantages of the NMF in the parts-based representation and the adaptive clustering. The localized fault features of interest can be extracted as well. To evaluate the performance of the proposed method, the 9 kinds of the bearing fault on a test bench is performed. The proposed method can effectively identify the fault severity and different fault types. Moreover, in comparison with the artificial neural network(ANN), NMF yields 99.3% mean accuracy which is much superior to ANN. This research presents a simple and practical resolution for the fault diagnosis problem of rolling element bearing in high dimensional feature space.
基金supported by Project of Chongqing Science and Technology Bureau (cstc2022jxjl0005)。
摘要This study aimed to investigate the pollution characteristics, source apportionment, and health risks associated with trace metal(loid)s(TMs) in the major agricultural producing areas in Chongqing, China. We analyzed the source apportionment and assessed the health risk of TMs in agricultural soils by using positive matrix factorization(PMF) model and health risk assessment(HRA) model based on Monte Carlo simulation. Meanwhile, we combined PMF and HRA models to explore the health risks of TMs in agricultural soils by different pollution sources to determine the priority control factors. Results showed that the average contents of cadmium(Cd), arsenic (As), lead(Pb), chromium(Cr), copper(Cu), nickel(Ni), and zinc(Zn) in the soil were found to be 0.26, 5.93, 27.14, 61.32, 23.81, 32.45, and 78.65 mg/kg, respectively. Spatial analysis and source apportionment analysis revealed that urban and industrial sources, agricultural sources, and natural sources accounted for 33.0%, 27.7%, and 39.3% of TM accumulation in the soil, respectively. In the HRA model based on Monte Carlo simulation, noncarcinogenic risks were deemed negligible(hazard index <1), the carcinogenic risks were at acceptable level(10-6<total carcinogenic risk ≤ 10-4), with higher risks observed for children compared to adults. The relationship between TMs, their sources, and health risks indicated that urban and industrial sources were primarily associated with As, contributing to 75.1% of carcinogenic risks and 55.7% of non-carcinogenic risks, making them the primary control factors. Meanwhile, agricultural sources were primarily linked to Cd and Pb, contributing to 13.1% of carcinogenic risks and 21.8% of non-carcinogenic risks, designating them as secondary control factors.
基金financially supported by the Nord-Pas-de-Calais Region Councilthe Ministry of Higher Education and Research+1 种基金the European Regional Development FundsAdib Kfoury acknowledges the“Pole Metropolitain Cote d'Opale”(PMCO)for its PhD financial support
摘要The constrained weighted-non-negative matrix factorization(CW-NMF)hybrid receptor model was applied to study the influence of steelmaking activities on PM2.5(particulate matter with equivalent aerodynamic diameter less than 2.5μm)composition in Dunkerque,Northern France.Semi-diurnal PM2.5samples were collected using a high volume sampler in winter 2010 and spring 2011 and were analyzed for trace metals,water-soluble ions,and total carbon using inductively coupled plasma–atomic emission spectrometry(ICP-AES),ICP-mass spectrometry(ICP-MS),ionic chromatography and micro elemental carbon analyzer.The elemental composition shows that NO3-,SO42-,NH_4~+and total carbon are the main PM2.5constituents.Trace metals data were interpreted using concentration roses and both influences of integrated steelworks and electric steel plant were evidenced.The distinction between the two sources is made possible by the use Zn/Fe and Zn/Mn diagnostic ratios.Moreover Rb/Cr,Pb/Cr and Cu/Cd combination ratio are proposed to distinguish the ISW-sintering stack from the ISW-fugitive emissions.The a priori knowledge on the influencing source was introduced in the CW-NMF to guide the calculation.Eleven source profiles with various contributions were identified:8 are characteristics of coastal urban background site profiles and 3 are related to the steelmaking activities.Between them,secondary nitrates,secondary sulfates and combustion profiles give the highest contributions and account for 93%of the PM2.5concentration.The steelwork facilities contribute in about 2%of the total PM2.5concentration and appear to be the main source of Cr,Cu,Fe,Mn,Zn.
基金supported by the National Natural Science Foundation of China(61702251,41971424,61701191,U1605254)the Natural Science Basic Research Plan in Shaanxi Province of China(2018JM6030)+4 种基金the Key Technical Project of Fujian Province(2017H6015)the Science and Technology Project of Xiamen(3502Z20183032)the Doctor Scientific Research Starting Foundation of Northwest University(338050050)Youth Academic Talent Support Program of Northwest University(360051900151)the Natural Sciences and Engineering Research Council of Canada,Canada。
摘要This paper presents a novel medical image registration algorithm named total variation constrained graphregularization for non-negative matrix factorization(TV-GNMF).The method utilizes non-negative matrix factorization by total variation constraint and graph regularization.The main contributions of our work are the following.First,total variation is incorporated into NMF to control the diffusion speed.The purpose is to denoise in smooth regions and preserve features or details of the data in edge regions by using a diffusion coefficient based on gradient information.Second,we add graph regularization into NMF to reveal intrinsic geometry and structure information of features to enhance the discrimination power.Third,the multiplicative update rules and proof of convergence of the TV-GNMF algorithm are given.Experiments conducted on datasets show that the proposed TV-GNMF method outperforms other state-of-the-art algorithms.
基金This work was supported in part by the National Natural Science Foundation of China(51679105,51809112,51939003,61872160)“Thirteenth Five Plan”Science and Technology Project of Education Department,Jilin Province(JJKH20200990KJ).
摘要A current problem in diet recommendation systems is the matching of food preferences with nutritional requirements,taking into account individual characteristics,such as body weight with individual health conditions,such as diabetes.Current dietary recommendations employ association rules,content-based collaborative filtering,and constraint-based methods,which have several limitations.These limitations are due to the existence of a special user group and an imbalance of non-simple attributes.Making use of traditional dietary recommendation algorithm researches,we combine the Adaboost classifier with probabilistic matrix factorization.We present a personalized diet recommendation algorithm by taking advantage of probabilistic matrix factorization via Adaboost.A probabilistic matrix factorization method extracts the implicit factors between individual food preferences and nutritional characteristics.From this,we can make use of those features with strong influence while discarding those with little influence.After incorporating these changes into our approach,we evaluated our algorithm’s performance.Our results show that our method performed better than others at matching preferred foods with dietary requirements,benefiting user health as a result.The algorithm fully considers the constraint relationship between users’attributes and nutritional characteristics of foods.Considering many complex factors in our algorithm,the recommended food result set meets both health standards and users’dietary preferences.A comparison of our algorithm with others demonstrated that our method offers high accuracy and interpretability.
摘要Nonnegative matrix factorization (NMF) is a method to get parts-based features of information and form the typical profiles. But the basis vectors NMF gets are not orthogonal so that parts-based features of information are usually redundancy. In this paper, we propose two different approaches based on localized non-negative matrix factorization (LNMF) to obtain the typical user session profiles and typical semantic profiles of junk mails. The LNMF get basis vectors as orthogonal as possible so that it can get accurate profiles. The experiments show that the approach based on LNMF can obtain better profiles than the approach based on NMF. Key words localized non-negative matrix factorization - profile - log mining - mail filtering CLC number TP 391 Foundation item: Supported by the National Natural Science Foundation of China (60373066, 60303024), National Grand Fundamental Research 973 Program of China (2002CB312000), National Research Foundation for the Doctoral Program of Higher Education of China (20020286004).Biography: Jiang Ji-xiang (1980-), male, Master candidate, research direction: data mining, knowledge representation on the Web.
基金supported by the National Natural Science Foundation of China(No.51877013),(ZJ),(http://gffzzf112c495998e46deh5nu0c5bv6nbf6nwf.ffgz.tsg.suse.edu.cn/)the Natural Science Foundation of Jiangsu Province(No.BK20181463),(ZJ),(http://gffzzc03b1aca7f8143e9h5nu0c5bv6nbf6nwf.ffgz.tsg.suse.edu.cn/)sponsored by Qing Lan Project of Jiangsu Province(no specific grant number),(ZJ),(http://gffzzf5c0e997a8b74428h5nu0c5bv6nbf6nwf.ffgz.tsg.suse.edu.cn/).
摘要Currently,functional connectomes constructed from neuroimaging data have emerged as a powerful tool in identifying brain disorders.If one brain disease just manifests as some cognitive dysfunction,it means that the disease may affect some local connectivity in the brain functional network.That is,there are functional abnormalities in the sub-network.Therefore,it is crucial to accurately identify them in pathological diagnosis.To solve these problems,we proposed a sub-network extraction method based on graph regularization nonnegative matrix factorization(GNMF).The dynamic functional networks of normal subjects and early mild cognitive impairment(eMCI)subjects were vectorized and the functional connection vectors(FCV)were assembled to aggregation matrices.Then GNMF was applied to factorize the aggregation matrix to get the base matrix,in which the column vectors were restored to a common sub-network and a distinctive sub-network,and visualization and statistical analysis were conducted on the two sub-networks,respectively.Experimental results demonstrated that,compared with other matrix factorization methods,the proposed method can more obviously reflect the similarity between the common subnetwork of eMCI subjects and normal subjects,as well as the difference between the distinctive sub-network of eMCI subjects and normal subjects,Therefore,the high-dimensional features in brain functional networks can be best represented locally in the lowdimensional space,which provides a new idea for studying brain functional connectomes.
基金Supported by the National Natural Science Foundation of China ( No. 60872083 ) and the National High Technology Research and Development Program of China (No. 2007AA12Z149).
摘要This paper considers a problem of unsupervised spectral unmixing of hyperspectral data. Based on the Linear Mixing Model ( LMM), a new method under the framework of nonnegative matrix fac- torization (NMF) is proposed, namely minimum distance constrained nonnegative matrix factoriza- tion (MDC-NMF). In this paper, firstly, a new regularization term, called endmember distance (ED) is considered, which is defined as the sum of the squared Euclidean distances from each end- member to their geometric center. Compared with the simplex volume, ED has better optimization properties and is conceptually intuitive. Secondly, a projected gradient (PG) scheme is adopted, and by the virtue of ED, in this scheme the optimal step size along the feasible descent direction can be calculated easily at each iteration. Thirdly, a finite step ( no more than the number of endmem- bers) terminated algorithm is used to project a point on the canonical simplex, by which the abun- dance nonnegative constraint and abundance sum-to-one constraint can be accurately satisfied in a light amount of computation. The experimental results, based on a set of synthetic data and real da- ta, demonstrate that, in the same running time, MDC-NMF outperforms several other similar meth- ods proposed recently.
基金Supported by National Natural Science Foundation of China (No. 60872065)
摘要An image fusion method combining complex contourlet transform(CCT) with nonnegative matrix factorization(NMF) is proposed in this paper.After two images are decomposed by CCT,NMF is applied to their highand low-frequency components,respectively,and finally an image is synthesized.Subjective-visual-quality of the image fusion result is compared with those of the image fusion methods based on NMF and the combination of wavelet /contourlet onsubsampled contourlet with NMF.The experimental results are evaluated quantitatively,and the running time is also contrasted.It is shown that the proposed image fusion method can gain larger information entropy,standard deviation and mean gradient,which means that it can better integrate featured information from all source images,avoid background noise and promote space clearness in the fusion image effectively.
基金supported by the Teaching Reform Research Project of Qinghai Minzu University,China(2021-JYYB-009)the“Chunhui Plan”Cooperative Scientific Research Project of the Ministry of Education of China(2018).
摘要Link prediction has attracted wide attention among interdisciplinaryresearchers as an important issue in complex network. It aims to predict the missing links in current networks and new links that will appear in future networks.Despite the presence of missing links in the target network of link prediction studies, the network it processes remains macroscopically as a large connectedgraph. However, the complexity of the real world makes the complex networksabstracted from real systems often contain many isolated nodes. This phenomenon leads to existing link prediction methods not to efficiently implement the prediction of missing edges on isolated nodes. Therefore, the cold-start linkprediction is favored as one of the most valuable subproblems of traditional linkprediction. However, due to the loss of many links in the observation network, thetopological information available for completing the link prediction task is extremely scarce. This presents a severe challenge for the study of cold-start link prediction. Therefore, how to mine and fuse more available non-topologicalinformation from observed network becomes the key point to solve the problemof cold-start link prediction. In this paper, we propose a framework for solving thecold-start link prediction problem, a joint-weighted symmetric nonnegative matrixfactorization model fusing graph regularization information, based on low-rankapproximation algorithms in the field of machine learning. First, the nonlinear features in high-dimensional space of node attributes are captured by the designedgraph regularization term. Second, using a weighted matrix, we associate the attribute similarity and first order structure information of nodes and constrain eachother. Finally, a unified framework for implementing cold-start link prediction isconstructed by using a symmetric nonnegative matrix factorization model to integrate the multiple information extracted together. Extensive experimental validationon five real networks with attributes shows that the proposed model has very goodpredictive performance when predicting missing edges of isolated nodes.
基金supported by the Key Research and Development Program of Hainan Province(Grant Nos.ZDYF2023GXJS163,ZDYF2024GXJS014)National Natural Science Foundation of China(NSFC)(Grant Nos.62162022,62162024)+3 种基金the Major Science and Technology Project of Hainan Province(Grant No.ZDKJ2020012)Hainan Provincial Natural Science Foundation of China(Grant No.620MS021)Youth Foundation Project of Hainan Natural Science Foundation(621QN211)Innovative Research Project for Graduate Students in Hainan Province(Grant Nos.Qhys2023-96,Qhys2023-95).
摘要Contrastive learning is a significant research direction in the field of deep learning.However,existing data augmentation methods often lead to issues such as semantic drift in generated views while the complexity of model pre-training limits further improvement in the performance of existing methods.To address these challenges,we propose the Efficient Clustering Network based on Matrix Factorization(ECN-MF).Specifically,we design a batched low-rank Singular Value Decomposition(SVD)algorithm for data augmentation to eliminate redundant information and uncover major patterns of variation and key information in the data.Additionally,we design a Mutual Information-Enhanced Clustering Module(MI-ECM)to accelerate the training process by leveraging a simple architecture to bring samples from the same cluster closer while pushing samples from other clusters apart.Extensive experiments on six datasets demonstrate that ECN-MF exhibits more effective performance compared to state-of-the-art algorithms.
摘要Non-negative matrix factorization (NMF) is a technique for dimensionality reduction by placing non-negativity constraints on the matrix. Based on the PARAFAC model, NMF was extended for three-dimension data decomposition. The three-dimension nonnegative matrix factorization (NMF3) algorithm, which was concise and easy to implement, was given in this paper. The NMF3 algorithm implementation was based on elements but not on vectors. It could decompose a data array directly without unfolding, which was not similar to that the traditional algorithms do, It has been applied to the simulated data array decomposition and obtained reasonable results. It showed that NMF3 could be introduced for curve resolution in chemometrics.
基金the National Natural Science Foundation of China(Nos.11361033 and 11861045)。
摘要One of the most important problems in complex networks is to identify the influential vertices for understanding and controlling of information diffusion and disease spreading.Most of the current centrality algorithms focus on single feature or manually extract the attributes,which occasionally results in the failure to fully capture the vertex’s importance.A new vertex centrality approach based on symmetric nonnegative matrix factorization(SNMF),called VCSNMF,is proposed in this paper.For highlight the characteristics of a network,the adjacency matrix and the degree matrix are fused to represent original data of the network via a weighted linear combination.First,SNMF automatically extracts the latent characteristics of vertices by factorizing the established original data matrix.Then we prove that each vertex’s composite feature which is constructed with one-dimensional factor matrix can be approximated as the term of eigenvector associated with the spectral radius of the network,otherwise obtained by the factor matrix on the hyperspace.Finally,VCSNMF integrates the composite feature and the topological structure to evaluate the performance of vertices.To verify the effectiveness of the VCSNMF criterion,eight existing centrality approaches are used as comparison measures to rank influential vertices in ten real-world networks.The experimental results assert the superiority of the method.