Characterized by self-monitoring and agile adaptation to fast changing dynamics in complex production environments,smart manufacturing as envisioned under Industry 4.0 aims to improve the throughput and reliability of...Characterized by self-monitoring and agile adaptation to fast changing dynamics in complex production environments,smart manufacturing as envisioned under Industry 4.0 aims to improve the throughput and reliability of production beyond the state-of-the-art.While the widespread application of deep learning(DL)has opened up new opportunities to accomplish the goal,data quality and model interpretability have continued to present a roadblock for the widespread acceptance of DL for real-world applications.This has motivated research on two fronts:data curation,which aims to provide quality data as input for meaningful DL-based analysis,and model interpretation,which intends to reveal the physical reasoning underlying DL model outputs and promote trust from the users.This paper summarizes several key techniques in data curation where breakthroughs in data denoising,outlier detection,imputation,balancing,and semantic annotation have demonstrated the effectiveness in information extraction from noisy,incomplete,insufficient,and/or unannotated data.Also highlighted are model interpretation methods that address the“black-box”nature of DL towards model transparency.展开更多
The extensive accumulation of genetic,genomic,expression,and breeding data on Prunus species often results in valuable information being lost or difficult to access for breeding purposes.We report a recent effort to i...The extensive accumulation of genetic,genomic,expression,and breeding data on Prunus species often results in valuable information being lost or difficult to access for breeding purposes.We report a recent effort to increase curation on Prunus data in the Genome Database for Rosaceae(GDR,rosaceae.org)and a case study that explores 25 years of curated data(from 1998 to 2023)to uncover the genetic architecture of key traits in Prunus species,provide actionable insights for breeding,and encourage the use of shared molecular data across Prunus species.The curated data includes 177 genetic maps,primarily for almond(19),apricot(21),peach(52),and sweet cherry(46).A total of 28971 trait-associated loci were reported,with 72.4% derived from genome-wide association studies,18.7% from quantitative trait loci(QTL),and 8.9% from Mendelian trait loci.Notably,76.4% of these loci are associated with morphological and quality traits,reflecting breeders'focus on consumer preferences.We identified 16 potential QTL hotspots linked to key traits such as morphology,phenology,fruit quality,and disease resistance.Additionally,we identified 17 high-priority syntenic regions among peach,sweet cherry,and almond.The colocalized markers and genes within the QTL hotspots and syntenic regions offer a valuable resource for tool development for Prunus breeding,especially for complex polyploid genomes and lesser studied species with limited genetic and genomic data.展开更多
The China National GeneBank Sequence Archive(CNSA)is an open and freely accessible curated data repository built for archiving,sharing,and reutilizing of multiomics data.The remarkable advancement in sequencing techno...The China National GeneBank Sequence Archive(CNSA)is an open and freely accessible curated data repository built for archiving,sharing,and reutilizing of multiomics data.The remarkable advancement in sequencing technologies has triggered a paradigm shift in life science research.However,it also poses tremendous challenges for the research community in data management and reusability.With the dramatic advance of sequencing technologies like spatial transcriptome sequencing,it brings an unprecedented explosion in sequence data and new requirements for data archiving.CNSA was established in 2017 as one of the fundamental infrastructures to offer multiomics data archiving for the worldwide research community.Here,we present the state-of-the-art enhancements of CNSA encompassing the dramatical increase of varied types of data,the latest features and services implemented in CNSA as well as consistent efforts supporting global cooperation in biodiversity preservation and utilization.CNSA provides public archiving and open-sharing services for sequencing data and relevant metadata including genome,transcriptome,metabolism,and proteome from single-cell(also spatial resolved)level to individual and population level,as well as further analyzed results.As of 2024,CNSA has archived>16.3 petabytes of data and provided the data curation,preservation,and open-share service for>1581 publications from>560 institutions.It plays a pivotal role in supporting global scientific projects such as the 10000 Plant Genomes Project.So far,CNSA has been recommended by various academic publishers such as Cell,Elsevier,and Oxford University Press.CNSA is accessible at http://gffzz984746bd41ee48d7sbxfx6w9cxcoq60op.ffgz.tsg.suse.edu.cn/cnsa/.展开更多
The Protein Data Bank(PDB)is an ever-growing database of three-dimensional macromolecular structures that has become a crucial resource for the drug discovery process.Exploring complexed proteins and accessing their a...The Protein Data Bank(PDB)is an ever-growing database of three-dimensional macromolecular structures that has become a crucial resource for the drug discovery process.Exploring complexed proteins and accessing their associated ligands are essential for researchers to understand biological processes and design new compounds of pharmaceutical interest.However,currently available tools for large-scale ligand identification fail to address many of the more complex ways in which ligands are stored and represented in PDB structures.Therefore,a new tool called LigExtract was specifically developed for the large-scale processing of PDB structures and the identification of their ligands.This is a fully opensource tool available to the scientific community,designed to provide end-to-end processing.Users simply provide a list of UniProt IDs,and LigExtract returns a list of ligands,their individual PDB files,a PDB file of the protein chains interacting with the ligand,and a series of log files.These logs record the decisions made during the ligand extraction process and flag additional scenarios that might have to be considered during any follow-up use of the processed files(e.g.,ligands covalently bound to the protein).LigExtract is freely available on GitHub(http://gffzz188fe103f8f1460asbxfx6w9cxcoq60op.ffgz.tsg.suse.edu.cn/comp-medchem/LigExtract).展开更多
Genome data of severe acute respiratory syndrome coronavirus 2(SARS-CoV-2)is essential for virus diagnosis,vaccine development,and variant surveillance.To archive and integrate worldwide SARS-CoV-2 genome data,a serie...Genome data of severe acute respiratory syndrome coronavirus 2(SARS-CoV-2)is essential for virus diagnosis,vaccine development,and variant surveillance.To archive and integrate worldwide SARS-CoV-2 genome data,a series of resources have been constructed,serving as a fundamental infrastructure for SARS-CoV-2 research,pandemic prevention and control,and coronavirus disease 2019(COVID-19)therapy.Here we present an over-view of extant SARS-CoV-2 resources that are devoted to genome data deposition and integration.We review deposition resources in data accessibility,metadata standardization,data curation and annotation;review integrative resources in data source,de-redundancy processing,data curation and quality assessment,and variant annotation.Moreover,we address issues that impede SARS-CoV-2 genome data integration,including low-complexity,inconsistency and absence of isolate name,sequence inconsistency,asynchronous update of genome data,and mismatched metadata.We finally provide insights into data standardization consensus and data submission guidelines,to promote SARS-CoV-2 genome data sharing and integration.展开更多
The FAIR data guiding principles have been recently developed and widely adopted to improve the Findability,Accessibility,Interoperability,and Reuse of digital assets in the face of an exponential increase of data vol...The FAIR data guiding principles have been recently developed and widely adopted to improve the Findability,Accessibility,Interoperability,and Reuse of digital assets in the face of an exponential increase of data volume and complexity.The FAIR data principles have been formulated on a general level and the technological implementation of these principles remains up to the industries and organizations working on maximizing the value of their data.Here,we describe the data management and curation methodologies and best practices developed for FAIRification of clinical exploratory biomarker data collected from over 250 clinical studies.We discuss the data curation effort involved,the resulting output,and the business and scientific impact of our work.Finally,we propose prospective planning for FAIR data to optimize data management efforts and maximize data value.展开更多
The completion of the Human Genome Project lays a foundation for systematically studying the human genome from evolutionary history to precision medicine against diseases.With the explosive growth of biological data, ...The completion of the Human Genome Project lays a foundation for systematically studying the human genome from evolutionary history to precision medicine against diseases.With the explosive growth of biological data, there is an increasing number of biological databases that have been developed in aid of human-related research. Here we present a collection of humanrelated biological databases and provide a mini-review by classifying them into different categories according to their data types. As human-related databases continue to grow not only in count but also in volume, challenges are ahead in big data storage, processing, exchange and curation.展开更多
The journal Genomics,Proteomics&Bioinformatics(GPB)is interested in submissions across all areas of life science,biology,and biomedicine,focusing on large data acquisition,analysis,and curation.
Data repository infrastructures for academics have appeared in waves since the dawn of Web technology.These waves are driven by changes in societal needs,archiving needs and the development of cloud computing resource...Data repository infrastructures for academics have appeared in waves since the dawn of Web technology.These waves are driven by changes in societal needs,archiving needs and the development of cloud computing resources.As such,the data repository landscape has many flavors when it comes to sustainability models,target audiences and feature sets.One thing that links all data repositories is a desire to make the content they host reusable,building on the core principles of cataloging content for economical and research speed efficiency.The FAIR principles are a common goal for all repository infrastructures to aim for.No matter what discipline or infrastructure,the goal of reusable content,for both humans and machines,is a common one.This is the first time that repositories can work toward a common goal that ultimately lends itself to interoperability.The idea that research can move further and faster as we un-silo these fantastic resources is an achievable one.This paper investigates the steps that existing repositories need to take in order to remain useful and relevant in a FAIR research world.展开更多
摘要Characterized by self-monitoring and agile adaptation to fast changing dynamics in complex production environments,smart manufacturing as envisioned under Industry 4.0 aims to improve the throughput and reliability of production beyond the state-of-the-art.While the widespread application of deep learning(DL)has opened up new opportunities to accomplish the goal,data quality and model interpretability have continued to present a roadblock for the widespread acceptance of DL for real-world applications.This has motivated research on two fronts:data curation,which aims to provide quality data as input for meaningful DL-based analysis,and model interpretation,which intends to reveal the physical reasoning underlying DL model outputs and promote trust from the users.This paper summarizes several key techniques in data curation where breakthroughs in data denoising,outlier detection,imputation,balancing,and semantic annotation have demonstrated the effectiveness in information extraction from noisy,incomplete,insufficient,and/or unannotated data.Also highlighted are model interpretation methods that address the“black-box”nature of DL towards model transparency.
基金supported by the USDA National Research Support Project(NRSP1O)the SCRI-NIFA Award 2022-51181-38449.
摘要The extensive accumulation of genetic,genomic,expression,and breeding data on Prunus species often results in valuable information being lost or difficult to access for breeding purposes.We report a recent effort to increase curation on Prunus data in the Genome Database for Rosaceae(GDR,rosaceae.org)and a case study that explores 25 years of curated data(from 1998 to 2023)to uncover the genetic architecture of key traits in Prunus species,provide actionable insights for breeding,and encourage the use of shared molecular data across Prunus species.The curated data includes 177 genetic maps,primarily for almond(19),apricot(21),peach(52),and sweet cherry(46).A total of 28971 trait-associated loci were reported,with 72.4% derived from genome-wide association studies,18.7% from quantitative trait loci(QTL),and 8.9% from Mendelian trait loci.Notably,76.4% of these loci are associated with morphological and quality traits,reflecting breeders'focus on consumer preferences.We identified 16 potential QTL hotspots linked to key traits such as morphology,phenology,fruit quality,and disease resistance.Additionally,we identified 17 high-priority syntenic regions among peach,sweet cherry,and almond.The colocalized markers and genes within the QTL hotspots and syntenic regions offer a valuable resource for tool development for Prunus breeding,especially for complex polyploid genomes and lesser studied species with limited genetic and genomic data.
基金supported by the Guangdong Genomics Data Center(2021B1212100001)Shenzhen Science and Technology Program(KQTD20230301092839007)+1 种基金Biological Breeding-National Science and Technology Major Project(2023ZD04073)the China National GeneBank。
摘要The China National GeneBank Sequence Archive(CNSA)is an open and freely accessible curated data repository built for archiving,sharing,and reutilizing of multiomics data.The remarkable advancement in sequencing technologies has triggered a paradigm shift in life science research.However,it also poses tremendous challenges for the research community in data management and reusability.With the dramatic advance of sequencing technologies like spatial transcriptome sequencing,it brings an unprecedented explosion in sequence data and new requirements for data archiving.CNSA was established in 2017 as one of the fundamental infrastructures to offer multiomics data archiving for the worldwide research community.Here,we present the state-of-the-art enhancements of CNSA encompassing the dramatical increase of varied types of data,the latest features and services implemented in CNSA as well as consistent efforts supporting global cooperation in biodiversity preservation and utilization.CNSA provides public archiving and open-sharing services for sequencing data and relevant metadata including genome,transcriptome,metabolism,and proteome from single-cell(also spatial resolved)level to individual and population level,as well as further analyzed results.As of 2024,CNSA has archived>16.3 petabytes of data and provided the data curation,preservation,and open-share service for>1581 publications from>560 institutions.It plays a pivotal role in supporting global scientific projects such as the 10000 Plant Genomes Project.So far,CNSA has been recommended by various academic publishers such as Cell,Elsevier,and Oxford University Press.CNSA is accessible at http://gffzz984746bd41ee48d7sbxfx6w9cxcoq60op.ffgz.tsg.suse.edu.cn/cnsa/.
基金supported by the Portuguese Science Foundation FCT(Fundacao para a Ciencia e a Tecnologia,Portugal)(Grant Nos.EXPL/QUI-OUT/1288/2021,2022.03752.PTDC,PTDC/MED-QUI/3542/2020,PhD grant 2024.05709.BDANA,UID/04138/2020,and CPCA/A2/6972/2020)supported by the Marie Skłodowska-Curie Actions(TClock4AD)(Grant No.101072895)the European Regional Development Fund(Grant No.LISBOA-01-0246-FEDER-000017).
摘要The Protein Data Bank(PDB)is an ever-growing database of three-dimensional macromolecular structures that has become a crucial resource for the drug discovery process.Exploring complexed proteins and accessing their associated ligands are essential for researchers to understand biological processes and design new compounds of pharmaceutical interest.However,currently available tools for large-scale ligand identification fail to address many of the more complex ways in which ligands are stored and represented in PDB structures.Therefore,a new tool called LigExtract was specifically developed for the large-scale processing of PDB structures and the identification of their ligands.This is a fully opensource tool available to the scientific community,designed to provide end-to-end processing.Users simply provide a list of UniProt IDs,and LigExtract returns a list of ligands,their individual PDB files,a PDB file of the protein chains interacting with the ligand,and a series of log files.These logs record the decisions made during the ligand extraction process and flag additional scenarios that might have to be considered during any follow-up use of the processed files(e.g.,ligands covalently bound to the protein).LigExtract is freely available on GitHub(http://gffzz188fe103f8f1460asbxfx6w9cxcoq60op.ffgz.tsg.suse.edu.cn/comp-medchem/LigExtract).
基金supported by Strategic Priority Research Program of the Chinese Academy of Sciences[XDB38030201,XDB38030400,XDB38050300]Youth Innovation Promotion Association of Chinese Academy of Sciences[2019104]。
摘要Genome data of severe acute respiratory syndrome coronavirus 2(SARS-CoV-2)is essential for virus diagnosis,vaccine development,and variant surveillance.To archive and integrate worldwide SARS-CoV-2 genome data,a series of resources have been constructed,serving as a fundamental infrastructure for SARS-CoV-2 research,pandemic prevention and control,and coronavirus disease 2019(COVID-19)therapy.Here we present an over-view of extant SARS-CoV-2 resources that are devoted to genome data deposition and integration.We review deposition resources in data accessibility,metadata standardization,data curation and annotation;review integrative resources in data source,de-redundancy processing,data curation and quality assessment,and variant annotation.Moreover,we address issues that impede SARS-CoV-2 genome data integration,including low-complexity,inconsistency and absence of isolate name,sequence inconsistency,asynchronous update of genome data,and mismatched metadata.We finally provide insights into data standardization consensus and data submission guidelines,to promote SARS-CoV-2 genome data sharing and integration.
摘要The FAIR data guiding principles have been recently developed and widely adopted to improve the Findability,Accessibility,Interoperability,and Reuse of digital assets in the face of an exponential increase of data volume and complexity.The FAIR data principles have been formulated on a general level and the technological implementation of these principles remains up to the industries and organizations working on maximizing the value of their data.Here,we describe the data management and curation methodologies and best practices developed for FAIRification of clinical exploratory biomarker data collected from over 250 clinical studies.We discuss the data curation effort involved,the resulting output,and the business and scientific impact of our work.Finally,we propose prospective planning for FAIR data to optimize data management efforts and maximize data value.
基金supported by the‘‘100-Talent Program’’of Chinese Academy of Sciencesthe Strategic Priority Research Program of the Chinese Academy of Sciences(Grant No.XDB13040500)+1 种基金the National High-tech R&D Program(863 ProgramGrant No.2012AA020409)by the Ministry of Science and Technology of China awarded to ZZ
摘要The completion of the Human Genome Project lays a foundation for systematically studying the human genome from evolutionary history to precision medicine against diseases.With the explosive growth of biological data, there is an increasing number of biological databases that have been developed in aid of human-related research. Here we present a collection of humanrelated biological databases and provide a mini-review by classifying them into different categories according to their data types. As human-related databases continue to grow not only in count but also in volume, challenges are ahead in big data storage, processing, exchange and curation.
摘要The journal Genomics,Proteomics&Bioinformatics(GPB)is interested in submissions across all areas of life science,biology,and biomedicine,focusing on large data acquisition,analysis,and curation.
摘要Data repository infrastructures for academics have appeared in waves since the dawn of Web technology.These waves are driven by changes in societal needs,archiving needs and the development of cloud computing resources.As such,the data repository landscape has many flavors when it comes to sustainability models,target audiences and feature sets.One thing that links all data repositories is a desire to make the content they host reusable,building on the core principles of cataloging content for economical and research speed efficiency.The FAIR principles are a common goal for all repository infrastructures to aim for.No matter what discipline or infrastructure,the goal of reusable content,for both humans and machines,is a common one.This is the first time that repositories can work toward a common goal that ultimately lends itself to interoperability.The idea that research can move further and faster as we un-silo these fantastic resources is an achievable one.This paper investigates the steps that existing repositories need to take in order to remain useful and relevant in a FAIR research world.