Pretraining model for biological sequence data.

Song, Bosheng; Li, Zimeng; Lin, Xuan; Wang, Jianmin; Wang, Tian; Fu, Xiangzheng

Song, Bosheng; Li, Zimeng; Lin, Xuan; Wang, Jianmin; Wang, Tian; Fu, Xiangzheng.

Brief Funct Genomics ; 20(3): 181-195, 2021 06 09.

Article in English | MEDLINE | ID: covidwho-1246686

ABSTRACT

ABSTRACT

With the development of high-throughput sequencing technology, biological sequence data reflecting life information becomes increasingly accessible. Particularly on the background of the COVID-19 pandemic, biological sequence data play an important role in detecting diseases, analyzing the mechanism and discovering specific drugs. In recent years, pretraining models that have emerged in natural language processing have attracted widespread attention in many research fields not only to decrease training cost but also to improve performance on downstream tasks. Pretraining models are used for embedding biological sequence and extracting feature from large biological sequence corpus to comprehensively understand the biological sequence data. In this survey, we provide a broad review on pretraining models for biological sequence data. Moreover, we first introduce biological sequences and corresponding datasets, including brief description and accessible link. Subsequently, we systematically summarize popular pretraining models for biological sequences based on four categories CNN, word2vec, LSTM and Transformer. Then, we present some applications with proposed pretraining models on downstream tasks to explain the role of pretraining models. Next, we provide a novel pretraining scheme for protein sequences and a multitask benchmark for protein pretraining models. Finally, we discuss the challenges and future directions in pretraining models for biological sequences.

Subject(s)

Algorithms; Computational Biology/methods; Data Mining/methods; High-Throughput Nucleotide Sequencing/methods; Natural Language Processing; Software; Datasets as Topic; Deep Learning; Humans; Models, Theoretical

Keywords

biological sequence; deep learning; pretraining model

Fulltext

XML

PubMed Links

Search on Google

Full text: Available Collection: International databases Database: MEDLINE Main subject: Algorithms / Natural Language Processing / Software / Computational Biology / Data Mining / High-Throughput Nucleotide Sequencing Type of study: Observational study / Reviews Limits: Humans Language: English Journal: Brief Funct Genomics Year: 2021 Document Type: Article

Similar

MEDLINE

LILACS

LIS

Fulltext

XML

PubMed Links

Search on Google