With the wide application of DNA sequencing technology, DNA sequences are still increasingly generated through the Sanger sequencing platform. SeqMan (in the LaserGene package) is an excellent program with an easy-t...With the wide application of DNA sequencing technology, DNA sequences are still increasingly generated through the Sanger sequencing platform. SeqMan (in the LaserGene package) is an excellent program with an easy-to-use graphical user interface (GUI) employed to assemble Sanger sequences into contigs. However, with increasing data size, larger sample sets and more sequenced loci make contig assemble complicated due to the considerable number of manual operations required to run SeqMan. Here, we present the 'autoSeqMan' software program, which can automatedly assemble contigs using SeqMan scripting language. There are two main modules available, namely, 'Classification' and 'Assembly'. Classification first undertakes preprocessing work, whereas Assembly generates a SeqMan script to consecutively assemble contigs for the classified files. Through comparison with manual operation, we showed that autoSeqMan saved substantial time in the preprocessing and assembly of Sanger sequences. We hope this tool will be useful for those with large sample sets to analyze, but with little programming experience. It is freely available at https://github.com/ Sun-Yanbo/autoSeqMan.展开更多
We described the construction of BAC contigs of the genome of a indica variety of Oryza sativa, Guang Lu Ai 4.An entire representative (sixfold coverage of rice chromosomes) and genetically stable BAC library of rice ...We described the construction of BAC contigs of the genome of a indica variety of Oryza sativa, Guang Lu Ai 4.An entire representative (sixfold coverage of rice chromosomes) and genetically stable BAC library of rice genome constructed in this lab has been systematically analysed by restriction enzyme fragmentation and polyacrylamide gel electrophoresis. And all the images thus obtained were subject to image-processing, which consisted of preliminary location of bands, cooperative tracking of lanesby correlation of adjacent bands, a precise densitometric pass, alignment at the marker bands with the standard,optional interactive editing, and normalization of the accepted bands. The contigs were generated based on the Computer Software specially designed for genome map ping. The number of contigs with 600 kb in length on average was 464; of contigs with 1000 kb in length on average was 107; of contigs with 1500 kb in length on average was 23. Therefore, all the contigs we have obtained amounted up to 420 megabases in length. Considering the size of rice genome (430 megabased), the contigs generated in this lab have covered nearly 98% of the rice genome. We are now in the process of mapping the contigs to chromosomes.展开更多
Centella asiatica is renowned for its medicinal properties,particularly due to its triterpenoid saponins,such as asiaticoside and madecassoside,which are in excess demand for the cosmetic industry.However,comprehensiv...Centella asiatica is renowned for its medicinal properties,particularly due to its triterpenoid saponins,such as asiaticoside and madecassoside,which are in excess demand for the cosmetic industry.However,comprehensive genomic resources for this species are lacking,which impedes the understanding of its biosynthetic pathways.Here,we report a telomere-to-telomere(T2T)C.asiatica genome.The genome size is 438.12 Mb with a contig N50 length of 54.12 Mb.The genome comprises 258.87 Mb of repetitive sequences and 25200 protein-coding genes.Comparative genomic analyses revealed C.asiatica as an early-diverging genus within the Apiaceae family with a single whole-genome duplication(WGD,Apiaceae-ω)event following the ancientγ-triplication,contrasting with Apiaceae species that exhibit two WGD events(Apiaceae-αand Apiaceae-ω).We further constructed 3D chromatin structures,A/B compartments,and topologically associated domains(TADs)in C.asiatica leaves,elucidating the influence of chromatin organization on expression WGD-derived genes.Additionally,gene family and functional characterization analysis highlight the key role of CasiOSC03 inα-amyrin production while also revealing significant expansion and high expression of CYP716,CYP714,and UGT73 families involved in asiaticoside biosynthesis compared to other Apiaceae species.Notably,a unique and large UGT73 gene cluster,located within the same TAD,is potentially pivotal for enhancing triterpenoid saponin.Weighted gene coexpression network analysis(WGCNA)further highlighted the pathways modulated in response to methyl jasmonate(MeJA),offering insights into the regulatory networks governing saponin biosynthesis.This work not only provides a valuable genomic resource for C.asiatica but also sheds light on the molecular mechanisms driving the biosynthesis of pharmacologically important metabolites.展开更多
Pucai(graphic)(Typha angustifolia L.),within the Typha spp.,is a distinctive semiaquatic vegetable.Lignin and chlorophyll are two crucial traits and quality indicators for Pucai.In this study,we assembled a 207.00-Mb ...Pucai(graphic)(Typha angustifolia L.),within the Typha spp.,is a distinctive semiaquatic vegetable.Lignin and chlorophyll are two crucial traits and quality indicators for Pucai.In this study,we assembled a 207.00-Mb high-quality gapless genome of Pucai,telomere-to-telomere(T2T)level with a contig N50 length of 13.73 Mb.The most abundant type of repetitive sequence,comprising 16.98%of the genome,is the long terminal repeat retrotransposons(LTR-RT).A total of 30 telomeres and 15 centromeric regions were predicted.Gene families related to lignin,chlorophyll biosynthesis,and disease resistance were greatly expanded,which played important roles in the adaptation of Pucai to wetlands.The slow evolution of Pucai was indicated by theσwhole-genome duplication(WGD)-associated Ks peaks from different Poales and the low activity of recent LTR-RT in Pucai.Meanwhile,we found a unique WGD event in Typhaceae.A statistical analysis and annotation of genomic variations were conducted in interspecies and intraspecies of Typha.Based on the T2T genome,we constructed lignin and chlorophyll metabolic pathways of Pucai.Subsequently,the candidate structural genes and transcription factors that regulate lignin and chlorophyll biosynthesis were identified.The T2T genomic resources will provide molecular information for lignin and chlorophyll accumulation and help to understand genome evolution in Pucai.展开更多
Merremia boisiana,a captivating species endemic to tropical rainforest habitats,belongs to the esteemed Convolvulaceae family.Renowned for its dazzling golden flowers and exceptional growth rate.This plant rapidly exp...Merremia boisiana,a captivating species endemic to tropical rainforest habitats,belongs to the esteemed Convolvulaceae family.Renowned for its dazzling golden flowers and exceptional growth rate.This plant rapidly expands,covering other vegetation,suppressing the growth of native species,and altering light availability and nutrient distribution within the forest,thereby impacting ecological balance.Here,we report the first high-quality M.boisiana genome assembly,comprising 510 Mb with a contig N50 of 21 Mb and an assembly completeness of 98.7%.This assembly includes the identification of 15 chromosomes and the annotation of 37,389 protein-coding genes,with a high annotation rate of 99.2%.By integrating genomic data from other Convolvulaceae species,we analyzed the karyotype evolution of M.boisiana and uncovered the fundamental ploidy level of the Convolvulaceae family.Based on existing research,we identified 110 highly expressed genes involved in the biosynthesis of SA,IAA,JA,and ABA,all of which play essential roles in plant growth.The EPS1 transcription factor,involved in SA synthesis,along with YUC11 and TIR2,which participate in auxin biosynthesis,and OPR2,associated with ABA biosynthesis,collectively contribute to the enhanced root growth of M.boisiana through mechanisms such as gene expansion,gene dosage,and root-specific expression.This study not only sheds light on the genetic complexity of M.boisiana but also provides a promising direction for improving stress resistance in sweet potatoes and advancing ecological research.Our findings promote the sustainable utilization of this species while broadening our understanding of tropical plant genomics.展开更多
Watercress(Nasturtium officinale R.Br.),an herbaceous plant in the cruciferous family,has a long history of use as a vegetable.In this report,we present a high-quality assembly of the watercress genome,based primarily...Watercress(Nasturtium officinale R.Br.),an herbaceous plant in the cruciferous family,has a long history of use as a vegetable.In this report,we present a high-quality assembly of the watercress genome,based primarily on PacBio and Hi-C sequencing data.The assembled genome of watercress was 337.51 Mb in size,with a contig N50 length of 3.26 Mb and a scaffold N50 length of 5.85 Mb.Approximately 49.15%(165.88 Mb)of the assembled genome was annotated as repetitive sequences,with long terminal repeats(LTRs)being the most abundant,representing 41.96% of the genome.About 96.6% of the assembly was anchored onto 16 pseudo-chromosomes using Hi-C data.Analyses of the syntenic relationships within and between species collectively indicated that watercress underwent an additional whole-genome duplication(WGD)event after divergence from Arabidopsis.The time of the watercressspecific WGD was estimated to be around 4.7 to 12.6 million years ago(Mya),which coincided with the Zanclean flood about 5.33 Mya.After the WGD,a total of 14,353 homologous gene pairs,including 22,476 genes were retained,accounting for 57.71% of all genes in the watercress genome.From the perspective of the watercress genome,we found that watercress might adapt to various abiotic stresses caused by the Zanclean flood through two mechanisms:one is enhancing tolerance to various abiotic stresses,and the other is escaping various abiotic stresses by floating growth.The watercress genome and transcriptome presented here provide useful information for subsequent molecular breeding and understanding how watercress adapted to an aquatic environment.展开更多
Pueraria lobata var.thomsonii(hereinafter abbreviated as Podalirius thomsonii),amember of the legume family,is one of the important traditional Chinese herbal medicines,and its puerarin extract is widely used in the h...Pueraria lobata var.thomsonii(hereinafter abbreviated as Podalirius thomsonii),amember of the legume family,is one of the important traditional Chinese herbal medicines,and its puerarin extract is widely used in the health and pharmaceutical industry.Here,we assembled a high-quality genome of P.thomsonii using long-read single-molecule sequencing and Hi-C technologies.The genome assembly is ~1.37 Gb in size and consists of 5145 contigs with a contig N50 of 593.70 kb,further clustered into 11 pseudochromosomes.Genome structural annotation resulted in∼869.33 Mb(~62.70% of the genome)repeat regions and 45270 protein-coding genes.Genome evolution analysis revealed that P.thomsonii is most closely related to soybean and underwent two ancient whole-genome duplication events;one was in the common ancestor shared by legume species and the other occurred independently at around 7.2million years ago,after its speciation.A total of 2373 gene familieswere found to be unique in P.thomsonii compared with five other legume species.Genes andmetabolites related to puerarin content in tuberous tissueswere characterized.A total of 572 genes that were upregulated in the puerarin biosynthesis pathway were identified,and 235 candidate genes were further enriched by omics data.Furthermore,we identified six 8-C-glucosyltransferase(8-C-GT)candidate genes significantly involved in puerarin metabolism.Our study filled a key genomic gap in the legume family,and provided valuable multi-omic resources for the genetic improvement of P.thomsonii.展开更多
Poplar line NL895 can potentially become a model plant for poplar study as it is a widely cultivated elite line.However,the lack of genome resources hindered the use of NL895 as the major plant material in poplar.In t...Poplar line NL895 can potentially become a model plant for poplar study as it is a widely cultivated elite line.However,the lack of genome resources hindered the use of NL895 as the major plant material in poplar.In this study,we provided a high-quality genome assembly for poplar line NL895 with PacBio single molecule real-time(SMRT)sequencing and High-throughput chromosome conformation capture(Hi-C)technology.The raw assembly of NL895 for the diploid genome included 606 contigs with a total size of~815 Mb,and the monoploid genome included 246 contigs with a total size of~412 Mb.The haplotype-resolved chromosomes in the diploid genomes were also generated.All the monoploid,diploid,and haplotype-resolved genomes showed more than 97%completeness and they can largely improve the mapping efficiency in RNA-Seq analysis.By comprehensively comparing the two haplotype genomes we found the heterozygosity of NL895 is much higher than other poplar lines.We also found that NL895 harbors more genomic variants and more gene diversity.The haplotype-specific genes showed higher variable gene expression patterns.These characters would be attributed to the high heterosis of poplar line NL895.The allele-specific expression(ASE)was also investigated and lots of alleles showed biased expressions in different tissues or environmental conditions.Taken together,the genome sequence for NL895 is a valuable tree genomic resource and it would greatly facilitate studies in poplar.展开更多
Avocado(Persea americana)is a member of themagnoliids,an early branching lineage of angiosperms that has high value globally with the fruit being highly nutritious.Here,we report a chromosome-level genome assembly for...Avocado(Persea americana)is a member of themagnoliids,an early branching lineage of angiosperms that has high value globally with the fruit being highly nutritious.Here,we report a chromosome-level genome assembly for the commercial avocado cultivar Hass,which represents 80% of theworld’s avocado consumption.The DNA contigs produced fromPacific Biosciences HiFi readswere further assembled using a previously published version of the genome supported by a genetic map.The total assembly was 913 Mb with a contig N50 of 84 Mb.Contigs assigned to the 12 chromosomes represented 874 Mb and covered 98.8% of benchmarked single-copy genes from embryophytes.Annotation of protein coding sequences identified 48915 avocado genes of which 39207 could be ascribed functions.The genome contained 62.6% repeat elements.Specific biosynthetic pathways of interest in the genome were investigated.The analysis suggested that the predominant pathway of heptose biosynthesis in avocado may be through sedoheptulose 1,7 bisphosphate rather than via alternative routes.Endoglucanase genes were high in number,consistent with avocado using cellulase for fruit ripening.The avocado genome appeared to have a limited number of translocations between homeologous chromosomes,despite having undergone multiple genome duplication events.Proteome clustering with related species permitted identification of genes unique to avocado and other members of the Lauraceae family,as well as genes unique to species diverged near or prior to the divergence of monocots and eudicots.This genome provides a tool to support future advances in the development of elite avocado varieties with higher yields and fruit quality.展开更多
Broccoli(Brassica oleracea var.italica Plenck)is an important vegetable crop,as it is rich in health-beneficial glucosinolates(GSLs).However,the genetic basis of the GSL diversity in Brassicaceae remains unclear.Here ...Broccoli(Brassica oleracea var.italica Plenck)is an important vegetable crop,as it is rich in health-beneficial glucosinolates(GSLs).However,the genetic basis of the GSL diversity in Brassicaceae remains unclear.Here we report a chromosome-level genome assembly of broccoli generated using PacBio HiFi reads and Hi-C technology.The final genome assembly is 613.79 Mb in size,with a contig N50 of 14.70 Mb.The GSL profile and content analysis of different B.oleracea varieties,combined with a phylogenetic tree analysis,sequence alignment,and the construction of a 3D model of the methylthioalkylmalate synthase 1(MAM1)protein,revealed that the gene copy number and amino acid sequence variation both contributed to the diversity of GSL biosynthesis in B.oleracea.The overexpression of BoMAM1(BolI0108790)in broccoli resulted in high accumulation and a high ratio of C4-GSLs,demonstrating that BoMAM1 is the key enzyme in C4-GSL biosynthesis.These results provide valuable insights for future genetic studies and nutritive component applications of Brassica crops.展开更多
Citrus reticulata‘Chachi’(CRC)has long been recognized for its nutritional benefits,health-promoting properties,and pharmacological potential.Despite its importance,the bioactive components of CRC and their biosynth...Citrus reticulata‘Chachi’(CRC)has long been recognized for its nutritional benefits,health-promoting properties,and pharmacological potential.Despite its importance,the bioactive components of CRC and their biosynthetic pathways have remained largely unexplored.In this study,we introduce a gap-free genome assembly for CRC,which has a size of 312.97 Mb and a contig N50 size of 32.18 Mb.We identified key structural genes,transcription factors,and metabolites crucial to flavonoid biosynthesis through genomic,transcriptomic,and metabolomic analyses.Our analyses reveal that 409 flavonoid metabolites,accounting for 83.30%of the total identified,are highly concentrated in the early stage of fruit development.This concentration decreases as the fruit develops,with a notable decline in compounds such as hesperetin,naringin,and most polymethoxyflavones observed in later fruit development stages.Additionally,we have examined the expression of 21 structural genes within the flavonoid biosynthetic pathway,and found a significant reduction in the expression levels of key genes including 4CL,CHS,CHI,FLS,F3H,and 4OMT during fruit development,aligning with the trend of flavonoid metabolite accumulation.In conclusion,this study offers deep insights into the genomic evolution,biosynthesis processes,and the nutritional and medicinal properties of CRC,which lay a solid foundation for further gene function studies and germplasm improvement in citrus.展开更多
Vernicia montana is a dioecious plant widely cultivated for high-quality tung oil production and ornamental purposes in the Euphor-biaceae family.The lack of genomic information has severely hindered molecular breedin...Vernicia montana is a dioecious plant widely cultivated for high-quality tung oil production and ornamental purposes in the Euphor-biaceae family.The lack of genomic information has severely hindered molecular breeding for genetic improvement and early sex identification in V.montana.Here,we present a chromosome-level reference genome of a male V.montana with a total size of 1.29 Gb and a contig N50 of 3.69 Mb.Genome analysis revealed that different repeat lineages drove the expansion of genome size.The model of chromosome evolution in the Euphorbiaceae family suggests that polyploidization-induced genomic structural variation reshaped the chromosome structure,giving rise to the diverse modern chromosomes.Based on whole-genome resequencing data and analyses of selective sweep and genetic diversity,several genes associated with stress resistance and flavonoid synthesis such as CYP450 genes and members of the LRR–RLK family,were identified and presumed to have been selected during the evolutionary process.Genome-wide association studies were conducted and a putative sex-linked insertion and deletion(InDel)(Chr 2:102799917-102799933 bp)was identified and developed as a polymorphic molecular marker capable of effectively detecting the gender of V.montana.This InDel is located in the second intron of VmBASS4,suggesting a possible role of VmBASS4 in sex determination in V.montana.This study sheds light on the genome evolution and sex identification of V.montana,which will facilitate research on the development of agronomically important traits and genomics-assisted breeding.展开更多
The YAC contig construction has been done for the Human X chromosome short armXp21.3—p11.3, a region which contains several genetic disease gene loci and is of highlybiomedical importance. Using known probes(OTC, DXS...The YAC contig construction has been done for the Human X chromosome short armXp21.3—p11.3, a region which contains several genetic disease gene loci and is of highlybiomedical importance. Using known probes(OTC, DXS166, DMDcDNA) and STS markersof this region, YAC screenings are performed by both YAC colony in situ hybridization andPCR methods. Totally 55 YACs are obtained from the YAC libraries of CEPH, ICRF andthe Institute. The size determination, the analysis of 26 pairs of microsatelite STS, the single copy probe hybridization and the Alu-PCR fingerprinting are performed for these YACs.The mapping of these YACs is performed, and finally, 6 YAC contigs in Xp21.3—11.3 are obtained, covering about 15 Mb. This work will greatly facilitate the positional cloning of disease genes or the genome sequencing in this important region.展开更多
A primary physical map of rice chromosome 12 was constructed using marker-based chromosome landing and chromosome walking. A BAC library from IR64 was screened using 84 RFLP markers, 4 STS markers and 6 microsatellite...A primary physical map of rice chromosome 12 was constructed using marker-based chromosome landing and chromosome walking. A BAC library from IR64 was screened using 84 RFLP markers, 4 STS markers and 6 microsatellite markers on chromosome 12 by colony hybridization and polymerase chain reaction (PCR) amplification. A total of 59 contigs consisting of 419 BAC clones including 5 single-clones were physically aligned on rice chromosome 12 with the largest BAC contig covering 855 kb. The whole physical map had a size of ~16 Mb and covered about 52% of rice chromosome 12. This physical map will be certainly helpful for map-based gene cloning of agronomically and biological important genes and understanding the genome structure of the chromosome.展开更多
As a part of the Multinational Genome Sequencing Project of Brassica rapa, linkage group R9 and R3 were sequenced using a bacterial artificial chromosome (BAC) by BAC strategy. The current physical contigs are expec...As a part of the Multinational Genome Sequencing Project of Brassica rapa, linkage group R9 and R3 were sequenced using a bacterial artificial chromosome (BAC) by BAC strategy. The current physical contigs are expected to cover approximately 90% euchromatins of both chromosomes. As the project progresses, BAC selection for sequence extension becomes more limited because BAC libraries are restriction enzyme-specific. To support the project, a random sheared fosmid library was constructed. The library consists of 97536 clones with average insert size of approximately 40 kb corresponding to seven genome equivalents, assuming a Chinese cabbage genome size of 550 Mb. The library was screened with primers designed at the end of sequences of nine points of scaffold gaps where BAC clones cannot be selected to extend the physical contigs. The selected positive clones were end-sequenced to check the overlap between the fosmid clones and the adjacent BAC clones. Nine fosmid clones were selected and fully sequenced. The sequences revealed two completed gap filling and seven sequence extensions, which can be used for further selection of BAC clones confirming that the fosmid library will facilitate the sequence completion of B. rapa.展开更多
Basal angiosperms contain a wide diversity of floral and growth forms and gave rise to the largest recent angiosperm lineages.As none of the basal angiosperm genomes has been sequenced,examining large bacterial artifi...Basal angiosperms contain a wide diversity of floral and growth forms and gave rise to the largest recent angiosperm lineages.As none of the basal angiosperm genomes has been sequenced,examining large bacterial artificial chromosome(BAC) inserts remains the main approach to providing a first glimpse of the structure and organization of their genomes.In this study,we sequenced a 126.9-kbp BAC contig harboring a cinnamyl alcohol dehydrogenase gene(LtuCAD1) in a basal angiosperm species,Liriodendron tulipifera L.,an important timber tree species with significant ecological and economic values.A key enzyme in lignin biosynthesis,CAD catalyzes the final step in the synthesis of monolignols.We carried out phylogenetic analyses of seven full-length CAD family genes(LtuCAD1-7) obtained from a comprehensive Liriodendron expressed sequence tag dataset.The phylogenetic tree suggests that LtuCAD1 is the primary CAD gene involved in lignifications as it is the only Liriodendron CAD grouped with the bona fide CADs class.As well as the LtuCAD1,the BAC contig contained fragmented sequences for one integrase,eight hypothetical proteins,two gag-pol polyproteins,one RNase H family protein,and one chromatin binding protein.Comparative analysis with other angiosperm species suggests that the genomic segment in this BAC has undergone frequent arrangement.This study is our initial step in identifying and understanding lignin biosynthesis genes from basal angiosperm species.Such knowledge can help bridge the information gap between hardwood(angiosperm) and softwood(gymnosperm) species and benefit potential breeding and biotechnology application for enhanced production of biomass and digestibility in L.tulipifera.展开更多
Pepper (Capsicum annuum. L.) is a widely cultivated vegetable crop worldwide and has the second largest planting area and the first largest vegetable output and value in China. Pepper root-knot nematode (Meloidogyn...Pepper (Capsicum annuum. L.) is a widely cultivated vegetable crop worldwide and has the second largest planting area and the first largest vegetable output and value in China. Pepper root-knot nematode (Meloidogyne spp.) is one of the most serious pests of pepper, which caused huge losses every year. Previous studies showed that the Me3 gene is resistant to a wide range of Meloidogyne species, including M. arenaria, M. javanica, and M. incognita. HDA149, a double haploid pepper genotype, harboring the root-knot nematode resistance gene Me3, was used to construct bacterial artificial chro- mosome library (BAC) via the vector of CopyControFM pCC1 in this study. The library consists of 210 200 BAC clones and is equivalent to 5.3 pepper genomes. The average insert size is 95 kb, and most of them are 90-120 kb; but the empty clones are less than 3%. In order to screen the BAC library easily, 550 super pools with 384 BAC clones of each pool were further developed in this study. Specific primers from Me3 gene locus were used for BAC library screening, and more than 20 positive BAC clones were obtained. Then the selected positive BAC clones were analyzed by restriction enzyme digestion, BAC-end sequencing, marker development, and new positive BAC clones exploration, respectively. Finally, the contig with total length of about 300 kb linked to the Me3 locus was constructed based on chromosome walking strategy, which made a solid foundation for the cloning of the important root-knot nematode resistance gene Me3.展开更多
基金supported by the National Natural Science Foundation of China(31671326)the Youth Innovation Promotion Association,Chinese Academy of Sciences
摘要With the wide application of DNA sequencing technology, DNA sequences are still increasingly generated through the Sanger sequencing platform. SeqMan (in the LaserGene package) is an excellent program with an easy-to-use graphical user interface (GUI) employed to assemble Sanger sequences into contigs. However, with increasing data size, larger sample sets and more sequenced loci make contig assemble complicated due to the considerable number of manual operations required to run SeqMan. Here, we present the 'autoSeqMan' software program, which can automatedly assemble contigs using SeqMan scripting language. There are two main modules available, namely, 'Classification' and 'Assembly'. Classification first undertakes preprocessing work, whereas Assembly generates a SeqMan script to consecutively assemble contigs for the classified files. Through comparison with manual operation, we showed that autoSeqMan saved substantial time in the preprocessing and assembly of Sanger sequences. We hope this tool will be useful for those with large sample sets to analyze, but with little programming experience. It is freely available at https://github.com/ Sun-Yanbo/autoSeqMan.
摘要We described the construction of BAC contigs of the genome of a indica variety of Oryza sativa, Guang Lu Ai 4.An entire representative (sixfold coverage of rice chromosomes) and genetically stable BAC library of rice genome constructed in this lab has been systematically analysed by restriction enzyme fragmentation and polyacrylamide gel electrophoresis. And all the images thus obtained were subject to image-processing, which consisted of preliminary location of bands, cooperative tracking of lanesby correlation of adjacent bands, a precise densitometric pass, alignment at the marker bands with the standard,optional interactive editing, and normalization of the accepted bands. The contigs were generated based on the Computer Software specially designed for genome map ping. The number of contigs with 600 kb in length on average was 464; of contigs with 1000 kb in length on average was 107; of contigs with 1500 kb in length on average was 23. Therefore, all the contigs we have obtained amounted up to 420 megabases in length. Considering the size of rice genome (430 megabased), the contigs generated in this lab have covered nearly 98% of the rice genome. We are now in the process of mapping the contigs to chromosomes.
基金supported by the Yunnan Characteristic Plant Extraction Laboratory(2022YKZY001)Yunnan Seed Laboratory(202305AR340004).
摘要Centella asiatica is renowned for its medicinal properties,particularly due to its triterpenoid saponins,such as asiaticoside and madecassoside,which are in excess demand for the cosmetic industry.However,comprehensive genomic resources for this species are lacking,which impedes the understanding of its biosynthetic pathways.Here,we report a telomere-to-telomere(T2T)C.asiatica genome.The genome size is 438.12 Mb with a contig N50 length of 54.12 Mb.The genome comprises 258.87 Mb of repetitive sequences and 25200 protein-coding genes.Comparative genomic analyses revealed C.asiatica as an early-diverging genus within the Apiaceae family with a single whole-genome duplication(WGD,Apiaceae-ω)event following the ancientγ-triplication,contrasting with Apiaceae species that exhibit two WGD events(Apiaceae-αand Apiaceae-ω).We further constructed 3D chromatin structures,A/B compartments,and topologically associated domains(TADs)in C.asiatica leaves,elucidating the influence of chromatin organization on expression WGD-derived genes.Additionally,gene family and functional characterization analysis highlight the key role of CasiOSC03 inα-amyrin production while also revealing significant expansion and high expression of CYP716,CYP714,and UGT73 families involved in asiaticoside biosynthesis compared to other Apiaceae species.Notably,a unique and large UGT73 gene cluster,located within the same TAD,is potentially pivotal for enhancing triterpenoid saponin.Weighted gene coexpression network analysis(WGCNA)further highlighted the pathways modulated in response to methyl jasmonate(MeJA),offering insights into the regulatory networks governing saponin biosynthesis.This work not only provides a valuable genomic resource for C.asiatica but also sheds light on the molecular mechanisms driving the biosynthesis of pharmacologically important metabolites.
基金supported by the Priority Academic Program Development of Jiangsu Higher Education Institutions Project(PAPD),Coordinated Extension of Major Agricultural Technologies Program of Jiangsu(2022-ZYXT-01-3)The study was supported by the Bioinformatics Center of Nanjing Agricultural University.
摘要Pucai(graphic)(Typha angustifolia L.),within the Typha spp.,is a distinctive semiaquatic vegetable.Lignin and chlorophyll are two crucial traits and quality indicators for Pucai.In this study,we assembled a 207.00-Mb high-quality gapless genome of Pucai,telomere-to-telomere(T2T)level with a contig N50 length of 13.73 Mb.The most abundant type of repetitive sequence,comprising 16.98%of the genome,is the long terminal repeat retrotransposons(LTR-RT).A total of 30 telomeres and 15 centromeric regions were predicted.Gene families related to lignin,chlorophyll biosynthesis,and disease resistance were greatly expanded,which played important roles in the adaptation of Pucai to wetlands.The slow evolution of Pucai was indicated by theσwhole-genome duplication(WGD)-associated Ks peaks from different Poales and the low activity of recent LTR-RT in Pucai.Meanwhile,we found a unique WGD event in Typhaceae.A statistical analysis and annotation of genomic variations were conducted in interspecies and intraspecies of Typha.Based on the T2T genome,we constructed lignin and chlorophyll metabolic pathways of Pucai.Subsequently,the candidate structural genes and transcription factors that regulate lignin and chlorophyll biosynthesis were identified.The T2T genomic resources will provide molecular information for lignin and chlorophyll accumulation and help to understand genome evolution in Pucai.
基金supported by the startup funds for the double firstclass disciplines of crop science in Hainan University(RZ2100003362)the National Natural Science Foundation of China(32172614)+2 种基金Hainan Province Science and Technology Special Fund(ZDYF2023XDNY050)Hainan Provincial Natural Science Foundation of China(324RC452)the Project of National Key Laboratory for Tropical Crop Breeding(No.NKLTCB202337).We thank the editor and anonymous reviewers for their insightful comments and suggestions.
摘要Merremia boisiana,a captivating species endemic to tropical rainforest habitats,belongs to the esteemed Convolvulaceae family.Renowned for its dazzling golden flowers and exceptional growth rate.This plant rapidly expands,covering other vegetation,suppressing the growth of native species,and altering light availability and nutrient distribution within the forest,thereby impacting ecological balance.Here,we report the first high-quality M.boisiana genome assembly,comprising 510 Mb with a contig N50 of 21 Mb and an assembly completeness of 98.7%.This assembly includes the identification of 15 chromosomes and the annotation of 37,389 protein-coding genes,with a high annotation rate of 99.2%.By integrating genomic data from other Convolvulaceae species,we analyzed the karyotype evolution of M.boisiana and uncovered the fundamental ploidy level of the Convolvulaceae family.Based on existing research,we identified 110 highly expressed genes involved in the biosynthesis of SA,IAA,JA,and ABA,all of which play essential roles in plant growth.The EPS1 transcription factor,involved in SA synthesis,along with YUC11 and TIR2,which participate in auxin biosynthesis,and OPR2,associated with ABA biosynthesis,collectively contribute to the enhanced root growth of M.boisiana through mechanisms such as gene expansion,gene dosage,and root-specific expression.This study not only sheds light on the genetic complexity of M.boisiana but also provides a promising direction for improving stress resistance in sweet potatoes and advancing ecological research.Our findings promote the sustainable utilization of this species while broadening our understanding of tropical plant genomics.
基金supported by the Jiangsu Seed Industry Revitalization Project[JBGS(2021)015,JBGS(2021)064]the Fundamental Research Funds for the Central Universities(KYPT2024001)+1 种基金a Project Funded by the Priority Academic Program Development of Jiangsu Higher Education Institutions,the Nanjing Science and technology planning project(202109022)the China Agriculture Research System(CARS-23-A-16).
摘要Watercress(Nasturtium officinale R.Br.),an herbaceous plant in the cruciferous family,has a long history of use as a vegetable.In this report,we present a high-quality assembly of the watercress genome,based primarily on PacBio and Hi-C sequencing data.The assembled genome of watercress was 337.51 Mb in size,with a contig N50 length of 3.26 Mb and a scaffold N50 length of 5.85 Mb.Approximately 49.15%(165.88 Mb)of the assembled genome was annotated as repetitive sequences,with long terminal repeats(LTRs)being the most abundant,representing 41.96% of the genome.About 96.6% of the assembly was anchored onto 16 pseudo-chromosomes using Hi-C data.Analyses of the syntenic relationships within and between species collectively indicated that watercress underwent an additional whole-genome duplication(WGD)event after divergence from Arabidopsis.The time of the watercressspecific WGD was estimated to be around 4.7 to 12.6 million years ago(Mya),which coincided with the Zanclean flood about 5.33 Mya.After the WGD,a total of 14,353 homologous gene pairs,including 22,476 genes were retained,accounting for 57.71% of all genes in the watercress genome.From the perspective of the watercress genome,we found that watercress might adapt to various abiotic stresses caused by the Zanclean flood through two mechanisms:one is enhancing tolerance to various abiotic stresses,and the other is escaping various abiotic stresses by floating growth.The watercress genome and transcriptome presented here provide useful information for subsequent molecular breeding and understanding how watercress adapted to an aquatic environment.
基金supported by the National Natural Science Foundation of China(31960420,31870275)the Guangxi Key R&D Program Project(Guike AB1850028)+1 种基金the Guangxi Natural Science Foundation Project(2018GXNSF BA294001,2019GXNSFBA245093)the Special Project for Basic Scientific Research of Guangxi Academy of Agricultural Sciences(Guinongke 2021YT057).
摘要Pueraria lobata var.thomsonii(hereinafter abbreviated as Podalirius thomsonii),amember of the legume family,is one of the important traditional Chinese herbal medicines,and its puerarin extract is widely used in the health and pharmaceutical industry.Here,we assembled a high-quality genome of P.thomsonii using long-read single-molecule sequencing and Hi-C technologies.The genome assembly is ~1.37 Gb in size and consists of 5145 contigs with a contig N50 of 593.70 kb,further clustered into 11 pseudochromosomes.Genome structural annotation resulted in∼869.33 Mb(~62.70% of the genome)repeat regions and 45270 protein-coding genes.Genome evolution analysis revealed that P.thomsonii is most closely related to soybean and underwent two ancient whole-genome duplication events;one was in the common ancestor shared by legume species and the other occurred independently at around 7.2million years ago,after its speciation.A total of 2373 gene familieswere found to be unique in P.thomsonii compared with five other legume species.Genes andmetabolites related to puerarin content in tuberous tissueswere characterized.A total of 572 genes that were upregulated in the puerarin biosynthesis pathway were identified,and 235 candidate genes were further enriched by omics data.Furthermore,we identified six 8-C-glucosyltransferase(8-C-GT)candidate genes significantly involved in puerarin metabolism.Our study filled a key genomic gap in the legume family,and provided valuable multi-omic resources for the genetic improvement of P.thomsonii.
基金supported by the National Natural Science Foundation of China(NSFC accession Nos 32171768 and 32171745)the Fundamental Research Funds for the Central Universities(Nos 2662022YLYYJ008 and 2262022 YLYJ007).
摘要Poplar line NL895 can potentially become a model plant for poplar study as it is a widely cultivated elite line.However,the lack of genome resources hindered the use of NL895 as the major plant material in poplar.In this study,we provided a high-quality genome assembly for poplar line NL895 with PacBio single molecule real-time(SMRT)sequencing and High-throughput chromosome conformation capture(Hi-C)technology.The raw assembly of NL895 for the diploid genome included 606 contigs with a total size of~815 Mb,and the monoploid genome included 246 contigs with a total size of~412 Mb.The haplotype-resolved chromosomes in the diploid genomes were also generated.All the monoploid,diploid,and haplotype-resolved genomes showed more than 97%completeness and they can largely improve the mapping efficiency in RNA-Seq analysis.By comprehensively comparing the two haplotype genomes we found the heterozygosity of NL895 is much higher than other poplar lines.We also found that NL895 harbors more genomic variants and more gene diversity.The haplotype-specific genes showed higher variable gene expression patterns.These characters would be attributed to the high heterosis of poplar line NL895.The allele-specific expression(ASE)was also investigated and lots of alleles showed biased expressions in different tissues or environmental conditions.Taken together,the genome sequence for NL895 is a valuable tree genomic resource and it would greatly facilitate studies in poplar.
基金funding from Hort Innovation,with The University of Queensland and the Australian Government as part of the National Tree Genomics Program,AS17000 Genomics Toolboxsupported by a graduate scholarship from The University of Queensland.
摘要Avocado(Persea americana)is a member of themagnoliids,an early branching lineage of angiosperms that has high value globally with the fruit being highly nutritious.Here,we report a chromosome-level genome assembly for the commercial avocado cultivar Hass,which represents 80% of theworld’s avocado consumption.The DNA contigs produced fromPacific Biosciences HiFi readswere further assembled using a previously published version of the genome supported by a genetic map.The total assembly was 913 Mb with a contig N50 of 84 Mb.Contigs assigned to the 12 chromosomes represented 874 Mb and covered 98.8% of benchmarked single-copy genes from embryophytes.Annotation of protein coding sequences identified 48915 avocado genes of which 39207 could be ascribed functions.The genome contained 62.6% repeat elements.Specific biosynthetic pathways of interest in the genome were investigated.The analysis suggested that the predominant pathway of heptose biosynthesis in avocado may be through sedoheptulose 1,7 bisphosphate rather than via alternative routes.Endoglucanase genes were high in number,consistent with avocado using cellulase for fruit ripening.The avocado genome appeared to have a limited number of translocations between homeologous chromosomes,despite having undergone multiple genome duplication events.Proteome clustering with related species permitted identification of genes unique to avocado and other members of the Lauraceae family,as well as genes unique to species diverged near or prior to the divergence of monocots and eudicots.This genome provides a tool to support future advances in the development of elite avocado varieties with higher yields and fruit quality.
基金supported by the National Key Research and Development Program of China(2022YFF1003000)the National Natural Science Foundation of China(32372682,32272747,32072585,32072568)+2 种基金the International Cooperation Projects of National Key R&D Program of China(2022YFE0108300)the Graduate Research Innovation Project of Hunan(2023XC103)the innovation and entrepreneurship training program for college students(S202310537006X).
摘要Broccoli(Brassica oleracea var.italica Plenck)is an important vegetable crop,as it is rich in health-beneficial glucosinolates(GSLs).However,the genetic basis of the GSL diversity in Brassicaceae remains unclear.Here we report a chromosome-level genome assembly of broccoli generated using PacBio HiFi reads and Hi-C technology.The final genome assembly is 613.79 Mb in size,with a contig N50 of 14.70 Mb.The GSL profile and content analysis of different B.oleracea varieties,combined with a phylogenetic tree analysis,sequence alignment,and the construction of a 3D model of the methylthioalkylmalate synthase 1(MAM1)protein,revealed that the gene copy number and amino acid sequence variation both contributed to the diversity of GSL biosynthesis in B.oleracea.The overexpression of BoMAM1(BolI0108790)in broccoli resulted in high accumulation and a high ratio of C4-GSLs,demonstrating that BoMAM1 is the key enzyme in C4-GSL biosynthesis.These results provide valuable insights for future genetic studies and nutritive component applications of Brassica crops.
基金supported by grants from Guangdong Province Science and Technology Plan Project(2020B020221001)the National Modern Agricultural(Citrus)Technology Systems of China(CARS-26)+3 种基金the Agricultural Competitive Industry Discipline Team Building Project of Guangdong Academy of Agricultural Sciences(202113TD)the Key Project for New Agricultural Cultivar Breeding in Zhejiang Province(2021C02066-1)the National Natural Science Foundation of China(32102318)supported by the public research platform of the College of Horticulture Science,Zhejiang A&F University.
摘要Citrus reticulata‘Chachi’(CRC)has long been recognized for its nutritional benefits,health-promoting properties,and pharmacological potential.Despite its importance,the bioactive components of CRC and their biosynthetic pathways have remained largely unexplored.In this study,we introduce a gap-free genome assembly for CRC,which has a size of 312.97 Mb and a contig N50 size of 32.18 Mb.We identified key structural genes,transcription factors,and metabolites crucial to flavonoid biosynthesis through genomic,transcriptomic,and metabolomic analyses.Our analyses reveal that 409 flavonoid metabolites,accounting for 83.30%of the total identified,are highly concentrated in the early stage of fruit development.This concentration decreases as the fruit develops,with a notable decline in compounds such as hesperetin,naringin,and most polymethoxyflavones observed in later fruit development stages.Additionally,we have examined the expression of 21 structural genes within the flavonoid biosynthetic pathway,and found a significant reduction in the expression levels of key genes including 4CL,CHS,CHI,FLS,F3H,and 4OMT during fruit development,aligning with the trend of flavonoid metabolite accumulation.In conclusion,this study offers deep insights into the genomic evolution,biosynthesis processes,and the nutritional and medicinal properties of CRC,which lay a solid foundation for further gene function studies and germplasm improvement in citrus.
基金supported by the National Natural Science Foundation of China(grants.32171843,32230073)the Science and Technology Innovation Program of Hunan Province(2022RC3055).
摘要Vernicia montana is a dioecious plant widely cultivated for high-quality tung oil production and ornamental purposes in the Euphor-biaceae family.The lack of genomic information has severely hindered molecular breeding for genetic improvement and early sex identification in V.montana.Here,we present a chromosome-level reference genome of a male V.montana with a total size of 1.29 Gb and a contig N50 of 3.69 Mb.Genome analysis revealed that different repeat lineages drove the expansion of genome size.The model of chromosome evolution in the Euphorbiaceae family suggests that polyploidization-induced genomic structural variation reshaped the chromosome structure,giving rise to the diverse modern chromosomes.Based on whole-genome resequencing data and analyses of selective sweep and genetic diversity,several genes associated with stress resistance and flavonoid synthesis such as CYP450 genes and members of the LRR–RLK family,were identified and presumed to have been selected during the evolutionary process.Genome-wide association studies were conducted and a putative sex-linked insertion and deletion(InDel)(Chr 2:102799917-102799933 bp)was identified and developed as a polymorphic molecular marker capable of effectively detecting the gender of V.montana.This InDel is located in the second intron of VmBASS4,suggesting a possible role of VmBASS4 in sex determination in V.montana.This study sheds light on the genome evolution and sex identification of V.montana,which will facilitate research on the development of agronomically important traits and genomics-assisted breeding.
基金Supported by the High Technology Research and Development Programme of Chinathe National Natural Science Foundation of China
摘要The YAC contig construction has been done for the Human X chromosome short armXp21.3—p11.3, a region which contains several genetic disease gene loci and is of highlybiomedical importance. Using known probes(OTC, DXS166, DMDcDNA) and STS markersof this region, YAC screenings are performed by both YAC colony in situ hybridization andPCR methods. Totally 55 YACs are obtained from the YAC libraries of CEPH, ICRF andthe Institute. The size determination, the analysis of 26 pairs of microsatelite STS, the single copy probe hybridization and the Alu-PCR fingerprinting are performed for these YACs.The mapping of these YACs is performed, and finally, 6 YAC contigs in Xp21.3—11.3 are obtained, covering about 15 Mb. This work will greatly facilitate the positional cloning of disease genes or the genome sequencing in this important region.
摘要A primary physical map of rice chromosome 12 was constructed using marker-based chromosome landing and chromosome walking. A BAC library from IR64 was screened using 84 RFLP markers, 4 STS markers and 6 microsatellite markers on chromosome 12 by colony hybridization and polymerase chain reaction (PCR) amplification. A total of 59 contigs consisting of 419 BAC clones including 5 single-clones were physically aligned on rice chromosome 12 with the largest BAC contig covering 855 kb. The whole physical map had a size of ~16 Mb and covered about 52% of rice chromosome 12. This physical map will be certainly helpful for map-based gene cloning of agronomically and biological important genes and understanding the genome structure of the chromosome.
基金This work was supported by grants from the National Academy of Agricultural Science(Code #200901FHT020508369)the BioGreen21 Program(Code #20050301034438 and Code #20070301034037),Rural Development Administration, Republic of Korea
摘要As a part of the Multinational Genome Sequencing Project of Brassica rapa, linkage group R9 and R3 were sequenced using a bacterial artificial chromosome (BAC) by BAC strategy. The current physical contigs are expected to cover approximately 90% euchromatins of both chromosomes. As the project progresses, BAC selection for sequence extension becomes more limited because BAC libraries are restriction enzyme-specific. To support the project, a random sheared fosmid library was constructed. The library consists of 97536 clones with average insert size of approximately 40 kb corresponding to seven genome equivalents, assuming a Chinese cabbage genome size of 550 Mb. The library was screened with primers designed at the end of sequences of nine points of scaffold gaps where BAC clones cannot be selected to extend the physical contigs. The selected positive clones were end-sequenced to check the overlap between the fosmid clones and the adjacent BAC clones. Nine fosmid clones were selected and fully sequenced. The sequences revealed two completed gap filling and seven sequence extensions, which can be used for further selection of BAC clones confirming that the fosmid library will facilitate the sequence completion of B. rapa.
基金supported by aNational Institute of Food and Agriculture/USDA grant(project number SC-1700324,technical contribution No. 5863 of the Clemson University Experiment Station)an investment award provided by Clemson University
摘要Basal angiosperms contain a wide diversity of floral and growth forms and gave rise to the largest recent angiosperm lineages.As none of the basal angiosperm genomes has been sequenced,examining large bacterial artificial chromosome(BAC) inserts remains the main approach to providing a first glimpse of the structure and organization of their genomes.In this study,we sequenced a 126.9-kbp BAC contig harboring a cinnamyl alcohol dehydrogenase gene(LtuCAD1) in a basal angiosperm species,Liriodendron tulipifera L.,an important timber tree species with significant ecological and economic values.A key enzyme in lignin biosynthesis,CAD catalyzes the final step in the synthesis of monolignols.We carried out phylogenetic analyses of seven full-length CAD family genes(LtuCAD1-7) obtained from a comprehensive Liriodendron expressed sequence tag dataset.The phylogenetic tree suggests that LtuCAD1 is the primary CAD gene involved in lignifications as it is the only Liriodendron CAD grouped with the bona fide CADs class.As well as the LtuCAD1,the BAC contig contained fragmented sequences for one integrase,eight hypothetical proteins,two gag-pol polyproteins,one RNase H family protein,and one chromatin binding protein.Comparative analysis with other angiosperm species suggests that the genomic segment in this BAC has undergone frequent arrangement.This study is our initial step in identifying and understanding lignin biosynthesis genes from basal angiosperm species.Such knowledge can help bridge the information gap between hardwood(angiosperm) and softwood(gymnosperm) species and benefit potential breeding and biotechnology application for enhanced production of biomass and digestibility in L.tulipifera.
基金supported by the National High-Tech R&D Program in China (2013AA102603)the Natural Science Foundation of Shandong Province,China (ZR2014YL014)+3 种基金the Youth Scientific Research Foundation of Shandong Academy of Agricultural Sciences,China (2014QNZ03)the Taishan Scholars Program of Shandong Province,China (2016-2020)the National Natural Science Foundation of China (31101425)Prof. Alain Palloxin,French National Institute for Agricultural Research (INRA),for kindly providing the pepper genotype HDA149
摘要Pepper (Capsicum annuum. L.) is a widely cultivated vegetable crop worldwide and has the second largest planting area and the first largest vegetable output and value in China. Pepper root-knot nematode (Meloidogyne spp.) is one of the most serious pests of pepper, which caused huge losses every year. Previous studies showed that the Me3 gene is resistant to a wide range of Meloidogyne species, including M. arenaria, M. javanica, and M. incognita. HDA149, a double haploid pepper genotype, harboring the root-knot nematode resistance gene Me3, was used to construct bacterial artificial chro- mosome library (BAC) via the vector of CopyControFM pCC1 in this study. The library consists of 210 200 BAC clones and is equivalent to 5.3 pepper genomes. The average insert size is 95 kb, and most of them are 90-120 kb; but the empty clones are less than 3%. In order to screen the BAC library easily, 550 super pools with 384 BAC clones of each pool were further developed in this study. Specific primers from Me3 gene locus were used for BAC library screening, and more than 20 positive BAC clones were obtained. Then the selected positive BAC clones were analyzed by restriction enzyme digestion, BAC-end sequencing, marker development, and new positive BAC clones exploration, respectively. Finally, the contig with total length of about 300 kb linked to the Me3 locus was constructed based on chromosome walking strategy, which made a solid foundation for the cloning of the important root-knot nematode resistance gene Me3.