Enterprise applications utilize relational databases and structured business processes, requiring slow and expensive conversion of inputs and outputs, from business documents such as invoices, purchase orders, and rec...Enterprise applications utilize relational databases and structured business processes, requiring slow and expensive conversion of inputs and outputs, from business documents such as invoices, purchase orders, and receipts, into known templates and schemas before processing. We propose a new LLM Agent-based intelligent data extraction, transformation, and load (IntelligentETL) pipeline that not only ingests PDFs and detects inputs within it but also addresses the extraction of structured and unstructured data by developing tools that most efficiently and securely deal with respective data types. We study the efficiency of our proposed pipeline and compare it with enterprise solutions that also utilize LLMs. We establish the supremacy in timely and accurate data extraction and transformation capabilities of our approach for analyzing the data from varied sources based on nested and/or interlinked input constraints.展开更多
Tree logic, inherited from ambient logic, is introduced as the formal foundation of related programming language and type systems, In this paper, we introduce recursion into such logic system, which can describe the t...Tree logic, inherited from ambient logic, is introduced as the formal foundation of related programming language and type systems, In this paper, we introduce recursion into such logic system, which can describe the tree data more dearly and concisely. By making a distinction between proposition and predicate, a concise semantics interpretation for our modal logic is given. We also develop a model checking algorithm for the logic without △ operator. The correctness of the algorithm is shown. Such work can be seen as the basis of the semi-structured data processing language and more flexible type system.展开更多
Deep learning methods are applied into structured data and in typical methods,low-order features are discarded after combining with high-order featuresfor prediction tasks.However,in structured data,ignorance of low-o...Deep learning methods are applied into structured data and in typical methods,low-order features are discarded after combining with high-order featuresfor prediction tasks.However,in structured data,ignorance of low-order features may cause the low prediction rate.To address this issue,in this paper,deeper attention-based network(DAN)is proposed.With DAN method,to keep both low-and high-order features,attention average pooling layer was utilized to aggregate features of each order.Furthermore,by shortcut connections from each layer to attention average pooling layer,DAN can be built extremely deep to obtain enough capacity.Experimental results show DAN has good performance and works effectively.展开更多
Introduction:Traditional dietary surveys are timeconsuming,and manual recording may lead to omissions.Improvement during data collection is essential to enhance accuracy of nutritional surveys.In recent years,large la...Introduction:Traditional dietary surveys are timeconsuming,and manual recording may lead to omissions.Improvement during data collection is essential to enhance accuracy of nutritional surveys.In recent years,large language models(LLMs)have been rapidly developed,which can provide text-processing functions and assist investigators in conducting dietary surveys.Methods:Thirty-eight participants from 15 families in the Huangpu and Jiading districts of Shanghai were selected.A standardized 24-hour dietary recall protocol was conducted using an intelligent recording pen that simultaneously captured audio data.These recordings were then transcribed into text.After preprocessing,we used GLM-4 for prompt engineering and chain-of-thought for collaborative reasoning,output structured data,and analyzed its integrity and consistency.Model performance was evaluated using precision and F1 scores.Results:The overall integrity rate of the LLMbased structured data reached 92.5%,and the overall consistency rate compared with manual recording was 86%.The LLM can accurately and completely recognize the names of ingredients and dining and production locations during the transcription.The LLM achieved 94%precision and an F1 score of 89.7%for the full dataset.Conclusion:LLM-based text recognition and structured data extraction can serve as effective auxiliary tools to improve efficiency and accuracy in traditional dietary surveys.With the rapid advancement of artificial intelligence,more accurate and efficient auxiliary tools can be developed for more precise and efficient data collection in nutrition research.展开更多
Background:Clinical and biomedical research in low-resource settings often faces substantial challenges due to the need for high-quality data with sufficient sample sizes to construct effective models.These constraint...Background:Clinical and biomedical research in low-resource settings often faces substantial challenges due to the need for high-quality data with sufficient sample sizes to construct effective models.These constraints hinder robust model training and prompt researchers to seek methods for leveraging existing knowledge from related studies to support new research efforts.Transfer learning(TL),a machine learning technique,emerges as a powerful solution by utilizing knowledge from pretrained models to enhance the performance of new models,offering promise across various healthcare domains.Despite its conceptual origins in the 1990s,the application of TL in medical research has remained limited,especially beyond image analysis.This review aims to analyze TL applications,highlight overlooked techniques,and suggest improvements for future healthcare research.Methods:Following the PRISMA-ScR guidelines,we conducted a search for published articles that employed TL with structured clinical or biomedical data by searching the SCOPUS,MEDLINE,Web of Science,Embase,and CINAHL databases.Results:We screened 5,080 papers,with 86 meeting the inclusion criteria.Among these,only 2%(2 of 86)utilized external studies,and 5%(4 of 86)addressed scenarios involving multi-site collaborations with privacy constraints.Conclusions:To achieve actionable TL with structured medical data while addressing regional disparities,inequality,and privacy constraints in healthcare research,we advocate for the careful identification of appropriate source data and models,the selection of suitable TL frameworks,and the validation of TL models with proper baselines.展开更多
Among the “three data rights,” the data utilization right has been persistently overlooked, and is similar to a neglected “middle child” in the context of the data rights family. However, it is precisely during th...Among the “three data rights,” the data utilization right has been persistently overlooked, and is similar to a neglected “middle child” in the context of the data rights family. However, it is precisely during the stages of processing and utilization that data undergoes its transformations and where its economic value is ultimately created. A series of recent policy documents on treating data as a factor of production have emphasized that the building of a scientific data property rights system requires a fair and efficient mechanism for benefit distribution, which provides reasonable preference for creators of data value and use value in terms of the income generated by data elements. Constrained by the inertial thinking of property right logic, the data utilization right is often regarded as a “transitional fulcrum” wherein the holders of data resources have to authorize the operators of data products to realize data value thereby. In the future structural design and implementation of the coordination mechanism for the property right system against the backdrop of the data factor-oriented reform, the establishment of data processing and utilization as an independent right will require the implementation of two core initiatives: first, attaching importance to the independent protection of the benefit distribution;second, implementing risk regulation for data security through optimization of governance. These two initiatives will serve as the key for optimizing the data factor governance system and accelerating the release of data value.展开更多
More web pages are widely applying AJAX (Asynchronous JavaScript XML) due to the rich interactivity and incremental communication. By observing, it is found that the AJAX contents, which could not be seen by traditi...More web pages are widely applying AJAX (Asynchronous JavaScript XML) due to the rich interactivity and incremental communication. By observing, it is found that the AJAX contents, which could not be seen by traditional crawler, are well-structured and belong to one specific domain generally. Extracting the structured data from AJAX contents and annotating its semantic are very significant for further applications. In this paper, a structured AJAX data extraction method for agricultural domain based on agricultural ontology was proposed. Firstly, Crawljax, an open AJAX crawling tool, was overridden to explore and retrieve the AJAX contents; secondly, the retrieved contents were partitioned into items and then classified by combining with agricultural ontology. HTML tags and punctuations were used to segment the retrieved contents into entity items. Finally, the entity items were clustered and the semantic annotation was assigned to clustering results according to agricultural ontology. By experimental evaluation, the proposed approach was proved effectively in resource exploring, entity extraction, and semantic annotation.展开更多
Aiming at the problems in the experimental teaching of data structure such as the replacement of students’learning process,the weakening of algorithm thinking,and the distortion of teaching evaluation results caused ...Aiming at the problems in the experimental teaching of data structure such as the replacement of students’learning process,the weakening of algorithm thinking,and the distortion of teaching evaluation results caused by the overuse of large models,this paper analyzes the influence mechanism of large models on students’ability after their intervention in experimental teaching.On the basis of defining the reasonable usage boundary of large models,a new procedural teaching evaluation system is constructed,and an implementation plan for data structure experimental teaching under the constraint of standardized AI usage is proposed.Teaching practice results show that this evaluation model can effectively standardize students’use of AI and has a significant promoting effect on improving students’algorithm understanding ability,program debugging ability,and engineering practical ability.展开更多
This study proposes an AI-supported project-based learning (PBL) framework for the Data Structures course under the Less Teaching, More Learning philosophy. The traditional chapter-based curriculum was reconstructed i...This study proposes an AI-supported project-based learning (PBL) framework for the Data Structures course under the Less Teaching, More Learning philosophy. The traditional chapter-based curriculum was reconstructed into eight authentic Python projects supported by the learning activity management system and a large language model-based AI teaching assistant. A controlled teaching experiment involving two undergraduate classes was conducted to evaluate the effectiveness of the proposed approach. The results indicate that the AI-supported PBL model significantly improved students’ academic performance, programming competence, and self-directed learning compared with conventional instruction. The proposed framework provides a practical and transferable paradigm for AI-enabled curriculum reform in computer science education.展开更多
Organizing unstructured information from books into a well-defined structure is a significant challenge in digital libraries.Most digital libraries can provide only search services at the granularity of books and few ...Organizing unstructured information from books into a well-defined structure is a significant challenge in digital libraries.Most digital libraries can provide only search services at the granularity of books and few libraries allow books to be accessed at the granularity of chapters,as manually constructing directory information for books is time-consuming.Extracting structured data from scanned books thus remains an urgent and important work.In this paper,we propose a novel structured data organization framework called CMSOF to organize scanned data automatically,and apply it to a Chinese medicine digital library.In the framework,image blocks and text blocks on the scanned page of books are separated based on the gray histogram projection method or a hybrid method of region growth and the Ada-Boosting classifier at first,and then the text structure is obtained from text blocks by text size and font type recognition.Finally,image blocks and structured OCRed text are correlated at the semantic level.By integrating the structured data into a Chinese medicine information system(CMIS),we can organize the Chinese medicine books well and users can access the books with flexibility,which indicates that CMSOF is an efficient framework to organize books mixed with images and text.展开更多
The problem of nonparametric identification of a multivariate nonlinearity in a D-input Hammer- stein system is examined. It is demonstrated that if the input measurements are structured, in the sense that there exist...The problem of nonparametric identification of a multivariate nonlinearity in a D-input Hammer- stein system is examined. It is demonstrated that if the input measurements are structured, in the sense that there exists some hidden relation between them, i.e. if they are distributed on some (unknown) d-dimensional space M in IRD, d 〈 D, then the system nonlinearity can be recovered at points on M with the convergence rate O(n-1/(2+d)) dependent on d. This rate is thus faster than the generic rate O(n-1/(2+D)) achieved by typical nonparametric algorithms and controlled solely by the number of inputs D.展开更多
To extract structured data from a web page with customized requirements,a user labels some DOM elements on the page with attribute names.The common features of the labeled elements are utilized to guide the user throu...To extract structured data from a web page with customized requirements,a user labels some DOM elements on the page with attribute names.The common features of the labeled elements are utilized to guide the user through the labeling process to minimize user efforts,and are also utilized to retrieve attribute values.To turn the attribute values into a structured result,the attribute pattern needs to be induced.For this purpose,a space-optimized suffix tree called attribute tree is built to transform the document object model(DOM) tree into a simpler form while preserving its useful properties such as attribute sequence order.The pattern is induced bottom-up on the attribute tree,and is further used to build the structured result.Experiments are conducted and show high performance of our approach in terms of precision,recall and structural correctness.展开更多
Data warehouse provides storage and management for mass data, but data schema evolves with time on. When data schema is changed, added or deleted, the data in data warehouse must comply with the changed data schema, s...Data warehouse provides storage and management for mass data, but data schema evolves with time on. When data schema is changed, added or deleted, the data in data warehouse must comply with the changed data schema, so data warehouse must be re organized or re constructed, but this process is exhausting and wasteful. In order to cope with these problems, this paper develops an approach to model data cube with XML, which emerges as a universal format for data exchange on the Web and which can make data warehouse flexible and scalable. This paper also extends OLAP algebra for XML based data cube, which is called X OLAP. 展开更多
Seismic data structure characteristics means the waveform character arranged in the time sequence at discrete data points in each 2-D or 3-D seismic trace. Hydrocarbon prediction using seismic data structure character...Seismic data structure characteristics means the waveform character arranged in the time sequence at discrete data points in each 2-D or 3-D seismic trace. Hydrocarbon prediction using seismic data structure characteristics is a new reservoir prediction technique. When the main pay interval is in carbonate fracture and fissure-cavern type reservoirs with very strong inhomogeneity, there are some difficulties with hydrocarbon prediction. Because of the special geological conditions of the eighth zone in the Tahe oil field, we apply seismic data structure characteristics to hydrocarbon prediction for the Ordovician reservoir in this zone. We divide the area oil zone into favorable and unfavorable blocks. Eighteen well locations were proposed in the favorable oil block, drilled, and recovered higher output of oil and gas.展开更多
Taking autonomous driving and driverless as the research object,we discuss and define intelligent high-precision map.Intelligent high-precision map is considered as a key link of future travel,a carrier of real-time p...Taking autonomous driving and driverless as the research object,we discuss and define intelligent high-precision map.Intelligent high-precision map is considered as a key link of future travel,a carrier of real-time perception of traffic resources in the entire space-time range,and the criterion for the operation and control of the whole process of the vehicle.As a new form of map,it has distinctive features in terms of cartography theory and application requirements compared with traditional navigation electronic maps.Thus,it is necessary to analyze and discuss its key features and problems to promote the development of research and application of intelligent high-precision map.Accordingly,we propose an information transmission model based on the cartography theory and combine the wheeled robot’s control flow in practical application.Next,we put forward the data logic structure of intelligent high-precision map,and analyze its application in autonomous driving.Then,we summarize the computing mode of“Crowdsourcing+Edge-Cloud Collaborative Computing”,and carry out key technical analysis on how to improve the quality of crowdsourced data.We also analyze the effective application scenarios of intelligent high-precision map in the future.Finally,we present some thoughts and suggestions for the future development of this field.展开更多
The discovery of catalysts has long been constrained by empirical optimization within specific families of materials[1,2].Although single‐atom catalysts focus on surface metal centers and their coordination environme...The discovery of catalysts has long been constrained by empirical optimization within specific families of materials[1,2].Although single‐atom catalysts focus on surface metal centers and their coordination environments[3],perovskite oxides emphasize bulk composition,lattice structure and the control of transition metal sites[4];as a result,the data structures and descriptors used in these different systems are often incompatible.Recently,Moon et al.reported in Nature Materials a deep learning framework for cross‐material catalyst discovery and proposed a crossbreeding neural network(CBNN).By integrating two experimental datasets,namely,carbon‐supported single‐atom catalysts and bulk perovskite oxides,the CBNN enabled the prediction of oxygen evolution reaction(OER)activity for a previously untrained material class:perovskite‐oxide‐supported single‐atom catalysts[5].This work not only validates the extrapolative capability of machine‐learning models in unexplored materials spaces but also provides an important paradigm for catalyst discovery driven by cross‐material knowledge transfer.展开更多
To make inorganic structure data more useful for further studies a five-point list of simple procedures to be followed by authors of crystal structure papers is proposed. 1. A crystal structure should be described wit...To make inorganic structure data more useful for further studies a five-point list of simple procedures to be followed by authors of crystal structure papers is proposed. 1. A crystal structure should be described with the space group corresponding to its true symmetry. 2. A new structure proposal should be tested, if it is realistic in principle. 3. A structure should be described with a space group in a setting given in the International Tables. 4. For a comparison with other structures the structure data should be standardized with the program STRUCTURE TIDY. 5. 揘ew?structure data should be checked in the databases, Chemical Abstracts or on-line internet resources, if they are really new. The list is supplemented with many explanations, commentaries, examples and references.展开更多
The current storage mechanism considered little in data’s keeping characteristics.These can produce fragments of various sizes among data sets.In some cases,these fragments may be serious and harm system performance....The current storage mechanism considered little in data’s keeping characteristics.These can produce fragments of various sizes among data sets.In some cases,these fragments may be serious and harm system performance.In this paper,we manage to modify the current storage mechanism.We introduce an extra storage unit called data bucket into the classical data manage architecture.Next,we modify the data manage mechanism to improve our designs.By keeping data according to their visited information,both the number of fragments and the fragment size are greatly reduced.Considering different data features and storage device conditions,we also improve the solid state drive(SSD)lifetime by keeping data into different spaces.Experiments show that our designs have a positive influence on the SSD storage density and actual service time.展开更多
With the rapid advancement of cloud computing technology,reversible data hiding algorithms in encrypted images(RDH-EI)have developed into an important field of study concentrated on safeguarding privacy in distributed...With the rapid advancement of cloud computing technology,reversible data hiding algorithms in encrypted images(RDH-EI)have developed into an important field of study concentrated on safeguarding privacy in distributed cloud environments.However,existing algorithms often suffer from low embedding capacities and are inadequate for complex data access scenarios.To address these challenges,this paper proposes a novel reversible data hiding algorithm in encrypted images based on adaptive median edge detection(AMED)and ciphertext-policy attributebased encryption(CP-ABE).This proposed algorithm enhances the conventional median edge detection(MED)by incorporating dynamic variables to improve pixel prediction accuracy.The carrier image is subsequently reconstructed using the Huffman coding technique.Encrypted image generation is then achieved by encrypting the image based on system user attributes and data access rights,with the hierarchical embedding of the group’s secret data seamlessly integrated during the encryption process using the CP-ABE scheme.Ultimately,the encrypted image is transmitted to the data hider,enabling independent embedding of the secret data and resulting in the creation of the marked encrypted image.This approach allows only the receiver to extract the authorized group’s secret data,thereby enabling fine-grained,controlled access.Test results indicate that,in contrast to current algorithms,the method introduced here considerably improves the embedding rate while preserving lossless image recovery.Specifically,the average maximum embedding rates for the(3,4)-threshold and(6,6)-threshold schemes reach 5.7853 bits per pixel(bpp)and 7.7781 bpp,respectively,across the BOSSbase,BOW-2,and USD databases.Furthermore,the algorithm facilitates permission-granting and joint-decryption capabilities.Additionally,this paper conducts a comprehensive examination of the algorithm’s robustness using metrics such as image correlation,information entropy,and number of pixel change rate(NPCR),confirming its high level of security.Overall,the algorithm can be applied in a multi-user and multi-level cloud service environment to realize the secure storage of carrier images and secret data.展开更多
Image-guided computer aided surgery system (ICAS) contributes to safeness and success of surgery operations by means of displaying anatomical structures and showing correlative information to surgeons in the process o...Image-guided computer aided surgery system (ICAS) contributes to safeness and success of surgery operations by means of displaying anatomical structures and showing correlative information to surgeons in the process of operation. Based on analysis of requirements for ICAS, a new concept of clinical knowledge-based ICAS was proposed. Designing a reasonable data structure model is essential for realizing this new concept. The traditional data structure is limited in expressing and reusing the clinical knowledge such as locating an anatomical object, topological relations of anatomical objects and correlative clinical attributes. A data structure model called mixed adjacency lists by octree-path-chain (MALOC) was outlined, which can combine patient's images with clinical knowledge, as well as efficiently locate the instrument and search the objects' information. The efficiency of data structures was analyzed and experimental results were given in comparison to other traditional data structures. The result of the nasal surgery experiment proves that MALOC is a proper model for clinical knowledge-based ICAS that has advantages in not only locating the operative instrument precisely but also proving surgeons with real-time operation-correlative information. It is shown that the clinical knowledge-based ICAS with MALOC model has advantages in terms of safety and success of surgical operations, and help in accurately locating the operative instrument and providing operation-correlative knowledge and information to surgeons in the process of operations.展开更多
摘要Enterprise applications utilize relational databases and structured business processes, requiring slow and expensive conversion of inputs and outputs, from business documents such as invoices, purchase orders, and receipts, into known templates and schemas before processing. We propose a new LLM Agent-based intelligent data extraction, transformation, and load (IntelligentETL) pipeline that not only ingests PDFs and detects inputs within it but also addresses the extraction of structured and unstructured data by developing tools that most efficiently and securely deal with respective data types. We study the efficiency of our proposed pipeline and compare it with enterprise solutions that also utilize LLMs. We establish the supremacy in timely and accurate data extraction and transformation capabilities of our approach for analyzing the data from varied sources based on nested and/or interlinked input constraints.
基金Supported by the National Natural Sciences Foun-dation of China (60233010 ,60273034 ,60403014) ,863 ProgramofChina (2002AA116010) ,973 Programof China (2002CB312002)
摘要Tree logic, inherited from ambient logic, is introduced as the formal foundation of related programming language and type systems, In this paper, we introduce recursion into such logic system, which can describe the tree data more dearly and concisely. By making a distinction between proposition and predicate, a concise semantics interpretation for our modal logic is given. We also develop a model checking algorithm for the logic without △ operator. The correctness of the algorithm is shown. Such work can be seen as the basis of the semi-structured data processing language and more flexible type system.
基金Sichuan Science and Technology Program 2018GZDZX0042,2018HH0061.
摘要Deep learning methods are applied into structured data and in typical methods,low-order features are discarded after combining with high-order featuresfor prediction tasks.However,in structured data,ignorance of low-order features may cause the low prediction rate.To address this issue,in this paper,deeper attention-based network(DAN)is proposed.With DAN method,to keep both low-and high-order features,attention average pooling layer was utilized to aggregate features of each order.Furthermore,by shortcut connections from each layer to attention average pooling layer,DAN can be built extremely deep to obtain enough capacity.Experimental results show DAN has good performance and works effectively.
基金Supported by the Ministry of Finance of the People’s Republic of China from 2022 to 2024(grant number 102393220020070000016).
摘要Introduction:Traditional dietary surveys are timeconsuming,and manual recording may lead to omissions.Improvement during data collection is essential to enhance accuracy of nutritional surveys.In recent years,large language models(LLMs)have been rapidly developed,which can provide text-processing functions and assist investigators in conducting dietary surveys.Methods:Thirty-eight participants from 15 families in the Huangpu and Jiading districts of Shanghai were selected.A standardized 24-hour dietary recall protocol was conducted using an intelligent recording pen that simultaneously captured audio data.These recordings were then transcribed into text.After preprocessing,we used GLM-4 for prompt engineering and chain-of-thought for collaborative reasoning,output structured data,and analyzed its integrity and consistency.Model performance was evaluated using precision and F1 scores.Results:The overall integrity rate of the LLMbased structured data reached 92.5%,and the overall consistency rate compared with manual recording was 86%.The LLM can accurately and completely recognize the names of ingredients and dining and production locations during the transcription.The LLM achieved 94%precision and an F1 score of 89.7%for the full dataset.Conclusion:LLM-based text recognition and structured data extraction can serve as effective auxiliary tools to improve efficiency and accuracy in traditional dietary surveys.With the rapid advancement of artificial intelligence,more accurate and efficient auxiliary tools can be developed for more precise and efficient data collection in nutrition research.
基金supported by the Duke/Duke-NUS Collaboration grant.
摘要Background:Clinical and biomedical research in low-resource settings often faces substantial challenges due to the need for high-quality data with sufficient sample sizes to construct effective models.These constraints hinder robust model training and prompt researchers to seek methods for leveraging existing knowledge from related studies to support new research efforts.Transfer learning(TL),a machine learning technique,emerges as a powerful solution by utilizing knowledge from pretrained models to enhance the performance of new models,offering promise across various healthcare domains.Despite its conceptual origins in the 1990s,the application of TL in medical research has remained limited,especially beyond image analysis.This review aims to analyze TL applications,highlight overlooked techniques,and suggest improvements for future healthcare research.Methods:Following the PRISMA-ScR guidelines,we conducted a search for published articles that employed TL with structured clinical or biomedical data by searching the SCOPUS,MEDLINE,Web of Science,Embase,and CINAHL databases.Results:We screened 5,080 papers,with 86 meeting the inclusion criteria.Among these,only 2%(2 of 86)utilized external studies,and 5%(4 of 86)addressed scenarios involving multi-site collaborations with privacy constraints.Conclusions:To achieve actionable TL with structured medical data while addressing regional disparities,inequality,and privacy constraints in healthcare research,we advocate for the careful identification of appropriate source data and models,the selection of suitable TL frameworks,and the validation of TL models with proper baselines.
摘要Among the “three data rights,” the data utilization right has been persistently overlooked, and is similar to a neglected “middle child” in the context of the data rights family. However, it is precisely during the stages of processing and utilization that data undergoes its transformations and where its economic value is ultimately created. A series of recent policy documents on treating data as a factor of production have emphasized that the building of a scientific data property rights system requires a fair and efficient mechanism for benefit distribution, which provides reasonable preference for creators of data value and use value in terms of the income generated by data elements. Constrained by the inertial thinking of property right logic, the data utilization right is often regarded as a “transitional fulcrum” wherein the holders of data resources have to authorize the operators of data products to realize data value thereby. In the future structural design and implementation of the coordination mechanism for the property right system against the backdrop of the data factor-oriented reform, the establishment of data processing and utilization as an independent right will require the implementation of two core initiatives: first, attaching importance to the independent protection of the benefit distribution;second, implementing risk regulation for data security through optimization of governance. These two initiatives will serve as the key for optimizing the data factor governance system and accelerating the release of data value.
基金supported by the Knowledge Innovation Program of the Chinese Academy of Sciencesthe National High-Tech R&D Program of China(2008BAK49B05)
摘要More web pages are widely applying AJAX (Asynchronous JavaScript XML) due to the rich interactivity and incremental communication. By observing, it is found that the AJAX contents, which could not be seen by traditional crawler, are well-structured and belong to one specific domain generally. Extracting the structured data from AJAX contents and annotating its semantic are very significant for further applications. In this paper, a structured AJAX data extraction method for agricultural domain based on agricultural ontology was proposed. Firstly, Crawljax, an open AJAX crawling tool, was overridden to explore and retrieve the AJAX contents; secondly, the retrieved contents were partitioned into items and then classified by combining with agricultural ontology. HTML tags and punctuations were used to segment the retrieved contents into entity items. Finally, the entity items were clustered and the semantic annotation was assigned to clustering results according to agricultural ontology. By experimental evaluation, the proposed approach was proved effectively in resource exploring, entity extraction, and semantic annotation.
摘要Aiming at the problems in the experimental teaching of data structure such as the replacement of students’learning process,the weakening of algorithm thinking,and the distortion of teaching evaluation results caused by the overuse of large models,this paper analyzes the influence mechanism of large models on students’ability after their intervention in experimental teaching.On the basis of defining the reasonable usage boundary of large models,a new procedural teaching evaluation system is constructed,and an implementation plan for data structure experimental teaching under the constraint of standardized AI usage is proposed.Teaching practice results show that this evaluation model can effectively standardize students’use of AI and has a significant promoting effect on improving students’algorithm understanding ability,program debugging ability,and engineering practical ability.
基金Shaanxi Computer Education Society Teaching Reform Project(2025),Exploration and Practice of Data Structures Teaching Driven by Large Language Models(Grant No.SJJG-25-04)Shaanxi Provincial Education Science 14th Five-Year Plan General Project(2025),Project-Based Teaching Reform of Data Structures Driven by Large Language Models(Grant No.SGH25Y3381)。
摘要This study proposes an AI-supported project-based learning (PBL) framework for the Data Structures course under the Less Teaching, More Learning philosophy. The traditional chapter-based curriculum was reconstructed into eight authentic Python projects supported by the learning activity management system and a large language model-based AI teaching assistant. A controlled teaching experiment involving two undergraduate classes was conducted to evaluate the effectiveness of the proposed approach. The results indicate that the AI-supported PBL model significantly improved students’ academic performance, programming competence, and self-directed learning compared with conventional instruction. The proposed framework provides a practical and transferable paradigm for AI-enabled curriculum reform in computer science education.
基金Project supported by the China Academic Digital Associative Library(CADAL)
摘要Organizing unstructured information from books into a well-defined structure is a significant challenge in digital libraries.Most digital libraries can provide only search services at the granularity of books and few libraries allow books to be accessed at the granularity of chapters,as manually constructing directory information for books is time-consuming.Extracting structured data from scanned books thus remains an urgent and important work.In this paper,we propose a novel structured data organization framework called CMSOF to organize scanned data automatically,and apply it to a Chinese medicine digital library.In the framework,image blocks and text blocks on the scanned page of books are separated based on the gray histogram projection method or a hybrid method of region growth and the Ada-Boosting classifier at first,and then the text structure is obtained from text blocks by text size and font type recognition.Finally,image blocks and structured OCRed text are correlated at the semantic level.By integrating the structured data into a Chinese medicine information system(CMIS),we can organize the Chinese medicine books well and users can access the books with flexibility,which indicates that CMSOF is an efficient framework to organize books mixed with images and text.
摘要The problem of nonparametric identification of a multivariate nonlinearity in a D-input Hammer- stein system is examined. It is demonstrated that if the input measurements are structured, in the sense that there exists some hidden relation between them, i.e. if they are distributed on some (unknown) d-dimensional space M in IRD, d 〈 D, then the system nonlinearity can be recovered at points on M with the convergence rate O(n-1/(2+d)) dependent on d. This rate is thus faster than the generic rate O(n-1/(2+D)) achieved by typical nonparametric algorithms and controlled solely by the number of inputs D.
基金Supported by the National High Technology Research and Development Programme of China(No.2009AA01 Z141)the National Natural Science Foundation of China(No.60573117)Beijing Natural Science Foundation(No.4131001)
摘要To extract structured data from a web page with customized requirements,a user labels some DOM elements on the page with attribute names.The common features of the labeled elements are utilized to guide the user through the labeling process to minimize user efforts,and are also utilized to retrieve attribute values.To turn the attribute values into a structured result,the attribute pattern needs to be induced.For this purpose,a space-optimized suffix tree called attribute tree is built to transform the document object model(DOM) tree into a simpler form while preserving its useful properties such as attribute sequence order.The pattern is induced bottom-up on the attribute tree,and is further used to build the structured result.Experiments are conducted and show high performance of our approach in terms of precision,recall and structural correctness.
摘要Data warehouse provides storage and management for mass data, but data schema evolves with time on. When data schema is changed, added or deleted, the data in data warehouse must comply with the changed data schema, so data warehouse must be re organized or re constructed, but this process is exhausting and wasteful. In order to cope with these problems, this paper develops an approach to model data cube with XML, which emerges as a universal format for data exchange on the Web and which can make data warehouse flexible and scalable. This paper also extends OLAP algebra for XML based data cube, which is called X OLAP.
基金This reservoir research is sponsored by the National 973 Subject Project (No. 2001CB209).
摘要Seismic data structure characteristics means the waveform character arranged in the time sequence at discrete data points in each 2-D or 3-D seismic trace. Hydrocarbon prediction using seismic data structure characteristics is a new reservoir prediction technique. When the main pay interval is in carbonate fracture and fissure-cavern type reservoirs with very strong inhomogeneity, there are some difficulties with hydrocarbon prediction. Because of the special geological conditions of the eighth zone in the Tahe oil field, we apply seismic data structure characteristics to hydrocarbon prediction for the Ordovician reservoir in this zone. We divide the area oil zone into favorable and unfavorable blocks. Eighteen well locations were proposed in the favorable oil block, drilled, and recovered higher output of oil and gas.
基金National Key Research and Development Program(No.2018YFB1305001)Major Consulting and Research Project of Chinese Academy of Engineering(No.2018-ZD-02-07)。
摘要Taking autonomous driving and driverless as the research object,we discuss and define intelligent high-precision map.Intelligent high-precision map is considered as a key link of future travel,a carrier of real-time perception of traffic resources in the entire space-time range,and the criterion for the operation and control of the whole process of the vehicle.As a new form of map,it has distinctive features in terms of cartography theory and application requirements compared with traditional navigation electronic maps.Thus,it is necessary to analyze and discuss its key features and problems to promote the development of research and application of intelligent high-precision map.Accordingly,we propose an information transmission model based on the cartography theory and combine the wheeled robot’s control flow in practical application.Next,we put forward the data logic structure of intelligent high-precision map,and analyze its application in autonomous driving.Then,we summarize the computing mode of“Crowdsourcing+Edge-Cloud Collaborative Computing”,and carry out key technical analysis on how to improve the quality of crowdsourced data.We also analyze the effective application scenarios of intelligent high-precision map in the future.Finally,we present some thoughts and suggestions for the future development of this field.
基金financially supported by the National Natural Science Foundation of China(Grant Nos.22471111,22201115,and 22425105)Gansu Provincial Science and Technology Program(Grant No.24ZD13GA015)the Natural Science Foundation of Gansu Province(Grant No.24JRRA387).
摘要The discovery of catalysts has long been constrained by empirical optimization within specific families of materials[1,2].Although single‐atom catalysts focus on surface metal centers and their coordination environments[3],perovskite oxides emphasize bulk composition,lattice structure and the control of transition metal sites[4];as a result,the data structures and descriptors used in these different systems are often incompatible.Recently,Moon et al.reported in Nature Materials a deep learning framework for cross‐material catalyst discovery and proposed a crossbreeding neural network(CBNN).By integrating two experimental datasets,namely,carbon‐supported single‐atom catalysts and bulk perovskite oxides,the CBNN enabled the prediction of oxygen evolution reaction(OER)activity for a previously untrained material class:perovskite‐oxide‐supported single‐atom catalysts[5].This work not only validates the extrapolative capability of machine‐learning models in unexplored materials spaces but also provides an important paradigm for catalyst discovery driven by cross‐material knowledge transfer.
摘要To make inorganic structure data more useful for further studies a five-point list of simple procedures to be followed by authors of crystal structure papers is proposed. 1. A crystal structure should be described with the space group corresponding to its true symmetry. 2. A new structure proposal should be tested, if it is realistic in principle. 3. A structure should be described with a space group in a setting given in the International Tables. 4. For a comparison with other structures the structure data should be standardized with the program STRUCTURE TIDY. 5. 揘ew?structure data should be checked in the databases, Chemical Abstracts or on-line internet resources, if they are really new. The list is supplemented with many explanations, commentaries, examples and references.
基金partly supported by the National Natural Science Foundation of China under Grant No.62072076the Research Fund of National Key Laboratory of Computer Architecture under Grant No.CARCH201811。
摘要The current storage mechanism considered little in data’s keeping characteristics.These can produce fragments of various sizes among data sets.In some cases,these fragments may be serious and harm system performance.In this paper,we manage to modify the current storage mechanism.We introduce an extra storage unit called data bucket into the classical data manage architecture.Next,we modify the data manage mechanism to improve our designs.By keeping data according to their visited information,both the number of fragments and the fragment size are greatly reduced.Considering different data features and storage device conditions,we also improve the solid state drive(SSD)lifetime by keeping data into different spaces.Experiments show that our designs have a positive influence on the SSD storage density and actual service time.
基金the National Natural Science Foundation of China(Grant Numbers 622724786210245062102451).
摘要With the rapid advancement of cloud computing technology,reversible data hiding algorithms in encrypted images(RDH-EI)have developed into an important field of study concentrated on safeguarding privacy in distributed cloud environments.However,existing algorithms often suffer from low embedding capacities and are inadequate for complex data access scenarios.To address these challenges,this paper proposes a novel reversible data hiding algorithm in encrypted images based on adaptive median edge detection(AMED)and ciphertext-policy attributebased encryption(CP-ABE).This proposed algorithm enhances the conventional median edge detection(MED)by incorporating dynamic variables to improve pixel prediction accuracy.The carrier image is subsequently reconstructed using the Huffman coding technique.Encrypted image generation is then achieved by encrypting the image based on system user attributes and data access rights,with the hierarchical embedding of the group’s secret data seamlessly integrated during the encryption process using the CP-ABE scheme.Ultimately,the encrypted image is transmitted to the data hider,enabling independent embedding of the secret data and resulting in the creation of the marked encrypted image.This approach allows only the receiver to extract the authorized group’s secret data,thereby enabling fine-grained,controlled access.Test results indicate that,in contrast to current algorithms,the method introduced here considerably improves the embedding rate while preserving lossless image recovery.Specifically,the average maximum embedding rates for the(3,4)-threshold and(6,6)-threshold schemes reach 5.7853 bits per pixel(bpp)and 7.7781 bpp,respectively,across the BOSSbase,BOW-2,and USD databases.Furthermore,the algorithm facilitates permission-granting and joint-decryption capabilities.Additionally,this paper conducts a comprehensive examination of the algorithm’s robustness using metrics such as image correlation,information entropy,and number of pixel change rate(NPCR),confirming its high level of security.Overall,the algorithm can be applied in a multi-user and multi-level cloud service environment to realize the secure storage of carrier images and secret data.
基金the Shanghai Municipal Education Commission Fund for Young Scholar (No. 02BQ23)the SEC E-Institute: Shanghai High Institutions Grid Project (No. 200304)
摘要Image-guided computer aided surgery system (ICAS) contributes to safeness and success of surgery operations by means of displaying anatomical structures and showing correlative information to surgeons in the process of operation. Based on analysis of requirements for ICAS, a new concept of clinical knowledge-based ICAS was proposed. Designing a reasonable data structure model is essential for realizing this new concept. The traditional data structure is limited in expressing and reusing the clinical knowledge such as locating an anatomical object, topological relations of anatomical objects and correlative clinical attributes. A data structure model called mixed adjacency lists by octree-path-chain (MALOC) was outlined, which can combine patient's images with clinical knowledge, as well as efficiently locate the instrument and search the objects' information. The efficiency of data structures was analyzed and experimental results were given in comparison to other traditional data structures. The result of the nasal surgery experiment proves that MALOC is a proper model for clinical knowledge-based ICAS that has advantages in not only locating the operative instrument precisely but also proving surgeons with real-time operation-correlative information. It is shown that the clinical knowledge-based ICAS with MALOC model has advantages in terms of safety and success of surgical operations, and help in accurately locating the operative instrument and providing operation-correlative knowledge and information to surgeons in the process of operations.