ABSTRACT
BACKGROUND: Full chloroplast genomes provide high resolution taxonomic discrimination between closely related plant species and are quickly replacing single and multi-locus barcoding regions as reference materials of choice for DNA based taxonomic annotation of plants. Bixa orellana, commonly known as "achiote" and "annatto" is a plant used for both human and animal foods and was thus identified for full chloroplast sequencing for the Center for Veterinary Medicine (CVM) Complete Chloroplast Animal Feed database. This work was conducted in collaboration with the Instituto de Medicina Tradicional (IMET) in Iquitos, Peru. There is a wide range of color variation in pods of Bixa orellana for which genetic loci that distinguish phenotypes have not yet been identified. Here we apply whole chloroplast genome sequencing of "red" and "yellow" individuals of Bixa orellana to provide high quality reference genomes to support kmer database development for use identifying this plant from complex mixtures using shotgun data. Additionally, we describe chloroplast gene content, synteny and phylogeny, and identify an indel and snp that may be associated with seed pod color. RESULTS: Fully assembled chloroplast genomes were produced for both red and yellow Bixa orellana accessions (158,918 and 158,823 bp respectively). Synteny and gene content was identical to the only other previously reported full chloroplast genome of Bixa orellana (NC_041550). We observed a 17 base pair deletion at position 58,399-58,415 in both accessions, relative to NC_041550 and a 6 bp deletion at position 75,531-75,526 and a snp at position 86,493 in red Bixa orellana. CONCLUSIONS: Our data provide high quality reference genomes of individuals of red and yellow Bixa orellana to support kmer based identity markers for use with shotgun sequencing approaches for rapid, precise identification of Bixa orellana from complex mixtures. Kmer based phylogeny of full chloroplast genomes supports monophylly of Bixaceae consistent with alignment based approaches. A potentially discriminatory indel and snp were identified that may be correlated with the red phenotype.