Search | VHL Regional Portal

The whole genome sequences and experimentally phased haplotypes of over 100 personal genomes.

Mao, Qing; Ciotlos, Serban; Zhang, Rebecca Yu; Ball, Madeleine P; Chin, Robert; Carnevali, Paolo; Barua, Nina; Nguyen, Staci; Agarwal, Misha R; Clegg, Tom; Connelly, Abram; Vandewege, Ward; Zaranek, Alexander Wait; Estep, Preston W; Church, George M; Drmanac, Radoje; Peters, Brock A.

Gigascience ; 5(1): 42, 2016 10 11.

Article in English | MEDLINE | ID: mdl-27724973

ABSTRACT

BACKGROUND: Since the completion of the Human Genome Project in 2003, it is estimated that more than 200,000 individual whole human genomes have been sequenced. A stunning accomplishment in such a short period of time. However, most of these were sequenced without experimental haplotype data and are therefore missing an important aspect of genome biology. In addition, much of the genomic data is not available to the public and lacks phenotypic information. FINDINGS: As part of the Personal Genome Project, blood samples from 184 participants were collected and processed using Complete Genomics' Long Fragment Read technology. Here, we present the experimental whole genome haplotyping and sequencing of these samples to an average read coverage depth of 100X. This is approximately three-fold higher than the read coverage applied to most whole human genome assemblies and ensures the highest quality results. Currently, 114 genomes from this dataset are freely available in the GigaDB repository and are associated with rich phenotypic data; the remaining 70 should be added in the near future as they are approved through the PGP data release process. For reproducibility analyses, 20 genomes were sequenced at least twice using independent LFR barcoded libraries. Seven genomes were also sequenced using Complete Genomics' standard non-barcoded library process. In addition, we report 2.6 million high-quality, rare variants not previously identified in the Single Nucleotide Polymorphisms database or the 1000 Genomes Project Phase 3 data. CONCLUSIONS: These genomes represent a unique source of haplotype and phenotype data for the scientific community and should help to expand our understanding of human genome evolution and function.

Subject(s)

Genome, Human , High-Throughput Nucleotide Sequencing/methods , Sequence Analysis, DNA/methods , DNA/blood , Haplotypes , Humans , Reproducibility of Results

Harvard Personal Genome Project: lessons from participatory public research.

Ball, Madeleine P; Bobe, Jason R; Chou, Michael F; Clegg, Tom; Estep, Preston W; Lunshof, Jeantine E; Vandewege, Ward; Zaranek, Alexander; Church, George M.

Genome Med ; 6(2): 10, 2014 Feb 28.

Article in English | MEDLINE | ID: mdl-24713084

ABSTRACT

BACKGROUND: Since its initiation in 2005, the Harvard Personal Genome Project has enrolled thousands of volunteers interested in publicly sharing their genome, health and trait data. Because these data are highly identifiable, we use an 'open consent' framework that purposefully excludes promises about privacy and requires participants to demonstrate comprehension prior to enrollment. DISCUSSION: Our model of non-anonymous, public genomes has led us to a highly participatory model of researcher-participant communication and interaction. The participants, who are highly committed volunteers, self-pursue and donate research-relevant datasets, and are actively engaged in conversations with both our staff and other Personal Genome Project participants. We have quantitatively assessed these communications and donations, and report our experiences with returning research-grade whole genome data to participants. We also observe some of the community growth and discussion that has occurred related to our project. SUMMARY: We find that public non-anonymous data is valuable and leads to a participatory research model, which we encourage others to consider. The implementation of this model is greatly facilitated by web-based tools and methods and participant education. Project results are long-term proactive participant involvement and the growth of a community that benefits both researchers and participants.

A public resource facilitating clinical use of genomes.

Ball, Madeleine P; Thakuria, Joseph V; Zaranek, Alexander Wait; Clegg, Tom; Rosenbaum, Abraham M; Wu, Xiaodi; Angrist, Misha; Bhak, Jong; Bobe, Jason; Callow, Matthew J; Cano, Carlos; Chou, Michael F; Chung, Wendy K; Douglas, Shawn M; Estep, Preston W; Gore, Athurva; Hulick, Peter; Labarga, Alberto; Lee, Je-Hyuk; Lunshof, Jeantine E; Kim, Byung Chul; Kim, Jong-Il; Li, Zhe; Murray, Michael F; Nilsen, Geoffrey B; Peters, Brock A; Raman, Anugraha M; Rienhoff, Hugh Y; Robasky, Kimberly; Wheeler, Matthew T; Vandewege, Ward; Vorhaus, Daniel B; Yang, Joyce L; Yang, Luhan; Aach, John; Ashley, Euan A; Drmanac, Radoje; Kim, Seong-Jin; Li, Jin Billy; Peshkin, Leonid; Seidman, Christine E; Seo, Jeong-Sun; Zhang, Kun; Rehm, Heidi L; Church, George M.

Proc Natl Acad Sci U S A ; 109(30): 11920-7, 2012 Jul 24.

Article in English | MEDLINE | ID: mdl-22797899

ABSTRACT

Rapid advances in DNA sequencing promise to enable new diagnostics and individualized therapies. Achieving personalized medicine, however, will require extensive research on highly reidentifiable, integrated datasets of genomic and health information. To assist with this, participants in the Personal Genome Project choose to forgo privacy via our institutional review board- approved "open consent" process. The contribution of public data and samples facilitates both scientific discovery and standardization of methods. We present our findings after enrollment of more than 1,800 participants, including whole-genome sequencing of 10 pilot participant genomes (the PGP-10). We introduce the Genome-Environment-Trait Evidence (GET-Evidence) system. This tool automatically processes genomes and prioritizes both published and novel variants for interpretation. In the process of reviewing the presumed healthy PGP-10 genomes, we find numerous literature references implying serious disease. Although it is sometimes impossible to rule out a late-onset effect, stringent evidence requirements can address the high rate of incidental findings. To that end we develop a peer production system for recording and organizing variant evaluations according to standard evidence guidelines, creating a public forum for reaching consensus on interpretation of clinically relevant variants. Genome analysis becomes a two-step process: using a prioritized list to record variant evaluations, then automatically sorting reviewed variants using these annotations. Genome data, health and trait information, participant samples, and variant interpretations are all shared in the public domain-we invite others to review our results using our participant samples and contribute to our interpretations. We offer our public resource and methods to further personalized medical research.

Subject(s)

Databases, Genetic , Genetic Variation , Genome, Human/genetics , Phenotype , Precision Medicine/methods , Software , Cell Line , Data Collection , Humans , Precision Medicine/trends , Sequence Analysis, DNA

Free Factories: Unified Infrastructure for Data Intensive Web Services.

Zaranek, Alexander Wait; Clegg, Tom; Vandewege, Ward; Church, George M.

Proc USENIX Annu Tech Conf ; 2008: 391-404, 2008 May 01.

Article in English | MEDLINE | ID: mdl-20514356

ABSTRACT

We introduce the Free Factory, a platform for deploying data-intensive web services using small clusters of commodity hardware and free software. Independently administered virtual machines called Freegols give application developers the flexibility of a general purpose web server, along with access to distributed batch processing, cache and storage services. Each cluster exploits idle RAM and disk space for cache, and reserves disks in each node for high bandwidth storage. The batch processing service uses a variation of the MapReduce model. Virtualization allows every CPU in the cluster to participate in batch jobs. Each 48-node cluster can achieve 4-8 gigabytes per second of disk I/O. Our intent is to use multiple clusters to process hundreds of simultaneous requests on multi-hundred terabyte data sets. Currently, our applications achieve 1 gigabyte per second of I/O with 123 disks by scheduling batch jobs on two clusters, one of which is located in a remote data center.

ABSTRACT

Subject(s)

ABSTRACT

ABSTRACT

Subject(s)

ABSTRACT

SEND TO:

SELECTION OF CITATIONS

SEARCH DETAIL