Search | VHL Regional Portal

LEON-BIS: multiple alignment evaluation of sequence neighbours using a Bayesian inference system.

Vanhoutreve, Renaud; Kress, Arnaud; Legrand, Baptiste; Gass, Hélène; Poch, Olivier; Thompson, Julie D.

BMC Bioinformatics ; 17(1): 271, 2016 Jul 07.

Article in English | MEDLINE | ID: mdl-27387560

ABSTRACT

BACKGROUND: A standard procedure in many areas of bioinformatics is to use a multiple sequence alignment (MSA) as the basis for various types of homology-based inference. Applications include 3D structure modelling, protein functional annotation, prediction of molecular interactions, etc. These applications, however sophisticated, are generally highly sensitive to the alignment used, and neglecting non-homologous or uncertain regions in the alignment can lead to significant bias in the subsequent inferences. RESULTS: Here, we present a new method, LEON-BIS, which uses a robust Bayesian framework to estimate the homologous relations between sequences in a protein multiple alignment. Sequences are clustered into sub-families and relations are predicted at different levels, including 'core blocks', 'regions' and full-length proteins. The accuracy and reliability of the predictions are demonstrated in large-scale comparisons using well annotated alignment databases, where the homologous sequence segments are detected with very high sensitivity and specificity. CONCLUSIONS: LEON-BIS uses robust Bayesian statistics to distinguish the portions of multiple sequence alignments that are conserved either across the whole family or within subfamilies. LEON-BIS should thus be useful for automatic, high-throughput genome annotations, 2D/3D structure predictions, protein-protein interaction predictions etc.

Subject(s)

Bayes Theorem , Computational Biology/methods , Proteins/chemistry , Sequence Alignment/methods , Amino Acid Sequence , Humans , Proteins/genetics , Sequence Homology, Amino Acid

SIBIS: a Bayesian model for inconsistent protein sequence estimation.

Khenoussi, Walyd; Vanhoutrève, Renaud; Poch, Olivier; Thompson, Julie D.

Bioinformatics ; 30(17): 2432-9, 2014 Sep 01.

Article in English | MEDLINE | ID: mdl-24825613

ABSTRACT

MOTIVATION: The prediction of protein coding genes is a major challenge that depends on the quality of genome sequencing, the accuracy of the model used to elucidate the exonic structure of the genes and the complexity of the gene splicing process leading to different protein variants. As a consequence, today's protein databases contain a huge amount of inconsistency, due to both natural variants and sequence prediction errors. RESULTS: We have developed a new method, called SIBIS, to detect such inconsistencies based on the evolutionary information in multiple sequence alignments. A Bayesian framework, combined with Dirichlet mixture models, is used to estimate the probability of observing specific amino acids and to detect inconsistent or erroneous sequence segments. We evaluated the performance of SIBIS on a reference set of protein sequences with experimentally validated errors and showed that the sensitivity is significantly higher than previous methods, with only a small loss of specificity. We also assessed a large set of human sequences from the UniProt database and found evidence of inconsistency in 48% of the previously uncharacterized sequences. We conclude that the integration of quality control methods like SIBIS in automatic analysis pipelines will be critical for the robust inference of structural, functional and phylogenetic information from these sequences. AVAILABILITY AND IMPLEMENTATION: Source code, implemented in C on a linux system, and the datasets of protein sequences are freely available for download at http://www.lbgi.fr/â¼julie/SIBIS.

Subject(s)

Sequence Analysis, Protein/methods , Algorithms , Animals , Bayes Theorem , Databases, Protein , Humans , Macaca mulatta , Phylogeny , Sequence Alignment , Software

ABSTRACT

Subject(s)

ABSTRACT

Subject(s)

SEND TO:

SELECTION OF CITATIONS

SEARCH DETAIL