Your browser doesn't support javascript.
A multi-task CNN learning model for taxonomic assignment of human viruses.
Ma, Haoran; Tan, Tin Wee; Ban, Kenneth Hon Kim.
  • Ma H; Department of Biochemistry, Yong Loo Lin School of Medicine, National University of Singapore, 117592, Singapore, Singapore.
  • Tan TW; Department of Biochemistry, Yong Loo Lin School of Medicine, National University of Singapore, 117592, Singapore, Singapore.
  • Ban KHK; National Supercomputing Centre (NSCC), 138632, Singapore, Singapore.
BMC Bioinformatics ; 22(Suppl 6): 194, 2021 Jun 02.
Article in English | MEDLINE | ID: covidwho-1388728
ABSTRACT

BACKGROUND:

Taxonomic assignment is a key step in the identification of human viral pathogens. Current tools for taxonomic assignment from sequencing reads based on alignment or alignment-free k-mer approaches may not perform optimally in cases where the sequences diverge significantly from the reference sequences. Furthermore, many tools may not incorporate the genomic coverage of assigned reads as part of overall likelihood of a correct taxonomic assignment for a sample.

RESULTS:

In this paper, we describe the development of a pipeline that incorporates a multi-task learning model based on convolutional neural network (MT-CNN) and a Bayesian ranking approach to identify and rank the most likely human virus from sequence reads. For taxonomic assignment of reads, the MT-CNN model outperformed Kraken 2, Centrifuge, and Bowtie 2 on reads generated from simulated divergent HIV-1 genomes and was more sensitive in identifying SARS as the closest relation in four RNA sequencing datasets for SARS-CoV-2 virus. For genomic region assignment of assigned reads, the MT-CNN model performed competitively compared with Bowtie 2 and the region assignments were used for estimation of genomic coverage that was incorporated into a naïve Bayesian network together with the proportion of taxonomic assignments to rank the likelihood of candidate human viruses from sequence data.

CONCLUSIONS:

We have developed a pipeline that combines a novel MT-CNN model that is able to identify viruses with divergent sequences together with assignment of the genomic region, with a Bayesian approach to ranking of taxonomic assignments by taking into account both the number of assigned reads and genomic coverage. The pipeline is available at GitHub via https//github.com/MaHaoran627/CNN_Virus .
Subject(s)
Keywords

Full text: Available Collection: International databases Database: MEDLINE Main subject: Viruses / COVID-19 Type of study: Prognostic study Limits: Humans Language: English Journal: BMC Bioinformatics Journal subject: Medical Informatics Year: 2021 Document Type: Article Affiliation country: S12859-021-04084-W

Similar

MEDLINE

...
LILACS

LIS


Full text: Available Collection: International databases Database: MEDLINE Main subject: Viruses / COVID-19 Type of study: Prognostic study Limits: Humans Language: English Journal: BMC Bioinformatics Journal subject: Medical Informatics Year: 2021 Document Type: Article Affiliation country: S12859-021-04084-W