Automated annotation of scientific texts for ML-based keyphrase extraction and validation.

Amusat, Oluwamayowa O; Hegde, Harshad; Mungall, Christopher J; Giannakou, Anna; Byers, Neil P; Gunter, Dan; Fagnan, Kjiersten; Ramakrishnan, Lavanya

Amusat, Oluwamayowa O; Hegde, Harshad; Mungall, Christopher J; Giannakou, Anna; Byers, Neil P; Gunter, Dan; Fagnan, Kjiersten; Ramakrishnan, Lavanya.

Affiliation

Amusat OO; Scientific Data Division, Lawrence Berkeley National Laboratory, 1 Cyclotron road, Berkeley, CA 94720, United States.
Hegde H; Division of Environmental Genomics and Systems Biology, Lawrence Berkeley National Laboratory, 1 Cyclotron road, Berkeley, CA 94720, United States.
Mungall CJ; Division of Environmental Genomics and Systems Biology, Lawrence Berkeley National Laboratory, 1 Cyclotron road, Berkeley, CA 94720, United States.
Giannakou A; Scientific Data Division, Lawrence Berkeley National Laboratory, 1 Cyclotron road, Berkeley, CA 94720, United States.
Byers NP; DOE Joint Genome Institute, Lawrence Berkeley National Laboratory, 1 Cyclotron road, Berkeley, CA 94720, United States.
Gunter D; Scientific Data Division, Lawrence Berkeley National Laboratory, 1 Cyclotron road, Berkeley, CA 94720, United States.
Fagnan K; DOE Joint Genome Institute, Lawrence Berkeley National Laboratory, 1 Cyclotron road, Berkeley, CA 94720, United States.
Ramakrishnan L; Scientific Data Division, Lawrence Berkeley National Laboratory, 1 Cyclotron road, Berkeley, CA 94720, United States.

Database (Oxford) ; 20242024 Sep 27.

Article in En | MEDLINE | ID: mdl-39331731

ABSTRACT

ABSTRACT

Advanced omics technologies and facilities generate a wealth of valuable data daily; however, the data often lack the essential metadata required for researchers to find, curate, and search them effectively. The lack of metadata poses a significant challenge in the utilization of these data sets. Machine learning (ML)-based metadata extraction techniques have emerged as a potentially viable approach to automatically annotating scientific data sets with the metadata necessary for enabling effective search. Text labeling, usually performed manually, plays a crucial role in validating machine-extracted metadata. However, manual labeling is time-consuming and not always feasible; thus, there is a need to develop automated text labeling techniques in order to accelerate the process of scientific innovation. This need is particularly urgent in fields such as environmental genomics and microbiome science, which have historically received less attention in terms of metadata curation and creation of gold-standard text mining data sets. In this paper, we present two novel automated text labeling approaches for the validation of ML-generated metadata for unlabeled texts, with specific applications in environmental genomics. Our techniques show the potential of two new ways to leverage existing information that is only available for select documents within a corpus to validate ML models, which can then be used to describe the remaining documents in the corpus. The first technique exploits relationships between different types of data sources related to the same research study, such as publications and proposals. The second technique takes advantage of domain-specific controlled vocabularies or ontologies. In this paper, we detail applying these approaches in the context of environmental genomics research for ML-generated metadata validation. Our results show that the proposed label assignment approaches can generate both generic and highly specific text labels for the unlabeled texts, with up to 44% of the labels matching with those suggested by a ML keyword extraction algorithm.

Subject(s)

Data Curation; Data Mining; Machine Learning; Data Curation/methods; Data Mining/methods; Metadata

Fulltext

Add to My VHL

XML

PubMed Links

Search on Google

Full text: 1 Collection: 01-internacional Database: MEDLINE Main subject: Data Mining / Data Curation / Machine Learning Language: En Journal: Database (Oxford) Year: 2024 Document type: Article Affiliation country: United States Country of publication: United kingdom

Fulltext

Add to My VHL

XML

PubMed Links

Search on Google