Search | VHL Regional Portal

AGRIS: the Arabidopsis Gene Regulatory Information Server, an update.

Yilmaz, Alper; Mejia-Guerra, Maria Katherine; Kurz, Kyle; Liang, Xiaoyu; Welch, Lonnie; Grotewold, Erich.

Nucleic Acids Res ; 39(Database issue): D1118-22, 2011 Jan.

Article in English | MEDLINE | ID: mdl-21059685

ABSTRACT

The Arabidopsis Gene Regulatory Information Server (AGRIS; http://arabidopsis.med.ohio-state.edu/) provides a comprehensive resource for gene regulatory studies in the model plant Arabidopsis thaliana. Three interlinked databases, AtTFDB, AtcisDB and AtRegNet, furnish comprehensive and updated information on transcription factors (TFs), predicted and experimentally verified cis-regulatory elements (CREs) and their interactions, respectively. In addition to significant contributions in the identification of the entire set of TF-DNA interactions, which are the key to understand the gene regulatory networks that govern Arabidopsis gene expression, tools recently incorporated into AGRIS include the complete set of words length 5-15 present in the Arabidopsis genome and the integration of AtRegNet with visualization tools, such as the recently developed ReIN application. All the information in AGRIS is publicly available and downloadable upon registration.

Subject(s)

Arabidopsis/genetics , Databases, Genetic , Gene Expression Regulation, Plant , Gene Regulatory Networks , Promoter Regions, Genetic , Transcription Factors/metabolism

WordSeeker: concurrent bioinformatics software for discovering genome-wide patterns and word-based genomic signatures.

Lichtenberg, Jens; Kurz, Kyle; Liang, Xiaoyu; Al-ouran, Rami; Neiman, Lev; Nau, Lee J; Welch, Joshua D; Jacox, Edwin; Bitterman, Thomas; Ecker, Klaus; Elnitski, Laura; Drews, Frank; Lee, Stephen Sauchi; Welch, Lonnie R.

BMC Bioinformatics ; 11 Suppl 12: S6, 2010 Dec 21.

Article in English | MEDLINE | ID: mdl-21210985

ABSTRACT

BACKGROUND: An important focus of genomic science is the discovery and characterization of all functional elements within genomes. In silico methods are used in genome studies to discover putative regulatory genomic elements (called words or motifs). Although a number of methods have been developed for motif discovery, most of them lack the scalability needed to analyze large genomic data sets. METHODS: This manuscript presents WordSeeker, an enumerative motif discovery toolkit that utilizes multi-core and distributed computational platforms to enable scalable analysis of genomic data. A controller task coordinates activities of worker nodes, each of which (1) enumerates a subset of the DNA word space and (2) scores words with a distributed Markov chain model. RESULTS: A comprehensive suite of performance tests was conducted to demonstrate the performance, speedup and efficiency of WordSeeker. The scalability of the toolkit enabled the analysis of the entire genome of Arabidopsis thaliana; the results of the analysis were integrated into The Arabidopsis Gene Regulatory Information Server (AGRIS). A public version of WordSeeker was deployed on the Glenn cluster at the Ohio Supercomputer Center. CONCLUSION: WordSeeker effectively utilizes concurrent computing platforms to enable the identification of putative functional elements in genomic data sets. This capability facilitates the analysis of the large quantity of sequenced genomic data.

Subject(s)

DNA/chemistry , Genomics/methods , Regulatory Sequences, Nucleic Acid , Software , Algorithms , Arabidopsis/genetics , Genome, Plant , Markov Chains , Sequence Analysis, DNA

The word landscape of the non-coding segments of the Arabidopsis thaliana genome.

Lichtenberg, Jens; Yilmaz, Alper; Welch, Joshua D; Kurz, Kyle; Liang, Xiaoyu; Drews, Frank; Ecker, Klaus; Lee, Stephen S; Geisler, Matt; Grotewold, Erich; Welch, Lonnie R.

BMC Genomics ; 10: 463, 2009 Oct 08.

Article in English | MEDLINE | ID: mdl-19814816

ABSTRACT

BACKGROUND: Genome sequences can be conceptualized as arrangements of motifs or words. The frequencies and positional distributions of these words within particular non-coding genomic segments provide important insights into how the words function in processes such as mRNA stability and regulation of gene expression. RESULTS: Using an enumerative word discovery approach, we investigated the frequencies and positional distributions of all 65,536 different 8-letter words in the genome of Arabidopsis thaliana. Focusing on promoter regions, introns, and 3' and 5' untranslated regions (3'UTRs and 5'UTRs), we compared word frequencies in these segments to genome-wide frequencies. The statistically interesting words in each segment were clustered with similar words to generate motif logos. We investigated whether words were clustered at particular locations or were distributed randomly within each genomic segment, and we classified the words using gene expression information from public repositories. Finally, we investigated whether particular sets of words appeared together more frequently than others. CONCLUSION: Our studies provide a detailed view of the word composition of several segments of the non-coding portion of the Arabidopsis genome. Each segment contains a unique word-based signature. The respective signatures consist of the sets of enriched words, 'unwords', and word pairs within a segment, as well as the preferential locations and functional classifications for the signature words. Additionally, the positional distributions of enriched words within the segments highlight possible functional elements, and the co-associations of words in promoter regions likely represent the formation of higher order regulatory modules. This work is an important step toward fully cataloguing the functional elements of the Arabidopsis genome.

Subject(s)

Arabidopsis/genetics , Computational Biology/methods , Genome, Plant , Models, Statistical , 3' Untranslated Regions , 5' Untranslated Regions , DNA, Plant/genetics , Gene Expression Regulation, Plant , Introns , Markov Chains , Promoter Regions, Genetic , Sequence Analysis, DNA

Word-based characterization of promoters involved in human DNA repair pathways.

Lichtenberg, Jens; Jacox, Edwin; Welch, Joshua D; Kurz, Kyle; Liang, Xiaoyu; Yang, Mary Qu; Drews, Frank; Ecker, Klaus; Lee, Stephen S; Elnitski, Laura; Welch, Lonnie R.

BMC Genomics ; 10 Suppl 1: S18, 2009 Jul 07.

Article in English | MEDLINE | ID: mdl-19594877

ABSTRACT

BACKGROUND: DNA repair genes provide an important contribution towards the surveillance and repair of DNA damage. These genes produce a large network of interacting proteins whose mRNA expression is likely to be regulated by similar regulatory factors. Full characterization of promoters of DNA repair genes and the similarities among them will more fully elucidate the regulatory networks that activate or inhibit their expression. To address this goal, the authors introduce a technique to find regulatory genomic signatures, which represents a specific application of the genomic signature methodology to classify DNA sequences as putative functional elements within a single organism. RESULTS: The effectiveness of the regulatory genomic signatures is demonstrated via analysis of promoter sequences for genes in DNA repair pathways of humans. The promoters are divided into two classes, the bidirectional promoters and the unidirectional promoters, and distinct genomic signatures are calculated for each class. The genomic signatures include statistically overrepresented words, word clusters, and co-occurring words. The robustness of this method is confirmed by the ability to identify sequences that exist as motifs in TRANSFAC and JASPAR databases, and in overlap with verified binding sites in this set of promoter regions. CONCLUSION: The word-based signatures are shown to be effective by finding occurrences of known regulatory sites. Moreover, the signatures of the bidirectional and unidirectional promoters of human DNA repair pathways are clearly distinct, exhibiting virtually no overlap. In addition to providing an effective characterization method for related DNA sequences, the signatures elucidate putative regulatory aspects of DNA repair pathways, which are notably under-characterized.

Subject(s)

Computational Biology/methods , DNA Repair , Promoter Regions, Genetic , Base Composition , Cluster Analysis , Databases, Genetic , Humans , Models, Statistical

ABSTRACT

Subject(s)

ABSTRACT

Subject(s)

ABSTRACT

Subject(s)

ABSTRACT

Subject(s)

SEND TO:

SELECTION OF CITATIONS

SEARCH DETAIL