I recently received an email notifying me that the journal The Lichenologist had just released a compilation of papers entitled 'Editor's Pick of Papers: from Genome to Ecosystems.' This group consists of the editor's top 15 articles published in the journal over the past two years (volumes 42 and 43), during which well over 100 works have been published in The Lichenologist. I was excited to see that one of my papers was selected! To read my previous blog post outlining the content of the paper, please click here. Many thanks to Peter Crittenden for his work as editor-in-chief and to the British Lichen Society for publishing the journal!
- Brendan
Thursday, December 8, 2011
Tuesday, December 6, 2011
Two new orders of fungi
This week I had a manuscript published in which two entirely new orders of fungi were established! They are Sarrameanales and Trapeliales, and they both belong to the subclass Ostropomycetidae in the class Lecanoromycetes (the main class of lichen-forming fungi). To read more about the taxonomy of these two orders, see the publication in the December 2011 issue of Phytologia. This research can partly be seen as a taxonomic extension of the NSF-funded Assembling the Fungal Tree of Life project, since the distinctness of these orders was only recognized as a result of multi-gene phylogenetic analyses conducted as part of this effort.
We are making progress taxonomically, but even the higher-level work is not done yet; there are still many groups of 'incertae sedis' fungi out there and probably many more legitimate orders to be described... which means that there's plenty more exciting work to be done in the future!
- Brendan
- Brendan
------------------------------------
Reference
Hodkinson, B. P., and J. C. Lendemer. 2011. The orders of Ostropomycetidae (Lecanoromycetes, Ascomycota): Recognition of Sarrameanales and Trapeliales with a request to retain Pertusariales over Agyriales. Phytologia 93(3): 407-412.
Download publication (PDF file)
Hodkinson, B. P., and J. C. Lendemer. 2011. The orders of Ostropomycetidae (Lecanoromycetes, Ascomycota): Recognition of Sarrameanales and Trapeliales with a request to retain Pertusariales over Agyriales. Phytologia 93(3): 407-412.
Download publication (PDF file)
Sunday, November 20, 2011
My New Supercomputing Center
Recently I applied for and received an XSEDE (Extreme Science and Engineering Discovery Environment) Startup Allocation Award (DEB-110024), allowing me to use the San Diego Supercomputing Center (SDSC) Trestles cluster. The cluster itself contains over 10,000 processor cores; you can read more about it here. While there are many different types of research being conducted using this cluster, my work focuses on ecological and evolutionary analyses of molecular sequence data. Previously I conducted my computationally-intensive analyses using the Duke Shared Cluster Resource (DSCR). However, since leaving Duke University a few months ago, I have been without a cluster. I am happy to say that I now have access to a cluster once more! Let the supercomputing resume!
- Brendan
- Brendan
Thursday, November 10, 2011
Outreach and training through Youtube videos
Over the past few months I have acted as the New York Botanical Garden's Workflow Coordinator for the NSF-funded "Digitization TCN Collaborative Research: North American Lichens and Bryophytes: Sensitive Indicators of Environmental Quality and Change" (EF-1115086). This project is a collaboration between multiple institutions across North America and is aimed at cataloging label data from the vast majority of North American lichen and bryophyte specimens. Recently, as part of this project, we at NYBG released two videos on Youtube. The first acts as an introduction to the project for the general public and gives some of the rationale behind it:
The second is a training video that can be used by members of the partner institutions or others who are thinking about taking on similar projects:
As you will probably notice from the videos themselves, Charlie Zimmerman, the imager that I have supervised for this project here at the garden, was the one who did most of the writing, filming and editing... and we greatly appreciate all of his work!
Please enjoy the videos!
- Brendan
The second is a training video that can be used by members of the partner institutions or others who are thinking about taking on similar projects:
As you will probably notice from the videos themselves, Charlie Zimmerman, the imager that I have supervised for this project here at the garden, was the one who did most of the writing, filming and editing... and we greatly appreciate all of his work!
Please enjoy the videos!
- Brendan
Sunday, November 6, 2011
Building linkage-probability-based RNA secondary structure models for phylogenetic inference
RNA secondary structure models are increasingly being integrated into likelihood-based phylogenetic inferences, but the dynamic structure of functional RNA molecules makes any single structural inference necessarily inaccurate. In this post I present an objective method for determining which elements of secondary structure are most stable based on the statistical significance of linkage probabilities between sites on a given RNA molecule. I briefly outline how this information can be integrated into a phylogenetic analysis by creating an input file that contains these statistically significant structural elements.
For some additional background on RNA secondary structure, see this previous post:
http://squamules.blogspot.com/2011/08/its-rna-secondary-structure.html
Functional RNA molecules include pairs of nucleotide sites that are linked to one another physically, resulting in specific secondary structures that define the shape of each molecule. This linkage causes certain sites to evolve in tandem with their counterparts. As such, the secondary structure of RNA molecules has been recognized for some time as a significant consideration in the inference of phylogenies from functional RNA-encoding genes (Kimura 1985, Tillier and Collins 1995). Typically, RNA secondary structure is used to optimize multiple-sequence-alignment accuracy for functional RNA-encoding genes (Gutell et al. 1992, Kjer 1995, Lendemer and Hodkinson 2009). However, in recent years, the use of secondary structure in modeling evolution for likelihood-based phylogenetic inferences has begun to gain popularity (Hodkinson and Lendemer in review, Savill et al. 2001, Telford et al. 2005). This approach requires defining the pairs of linked sites on the encoded RNA molecule and treating these pairs as states that are separate from the standard independent nucleotide states (A, C, T, G) in a phylogenetic inference. This method allows one to properly account for the interdependency of interacting nucleotides, since paired RNA nucleotides are no longer required to be treated as independent sites, leading to a more accurate approach for modeling sequence evolution.
Current protocols for integrating RNA secondary structure data into phylogenetic analyses require a single hypothetical structure to be used as an input. Structures are typically inferred using algorithms that minimize free energy or use other thermodynamic considerations to produce the best single structural inference (Mathews & Turner 2006). However, RNA secondary structure is dynamic, frequently changing in the cell as RNA catalyzes reactions and performs various cellular functions. When RNA molecules encounter certain enzymes and cellular components, the thermodynamic rules that previously favored one structure might strongly favor another. Additionally, different methods of structural inference are not always comparable, and small differences in algorithms can favor significantly different structures; the use of differing structural models in a phylogenetic context can have consequences in terms of both topology inference and the calculation of support (Ullrich et al. 2010).
These problems can largely be solved by removing statistically non-significant linkages from phylogenetic analyses, leaving only the most probable structural elements to be incorporated into downstream inferences. The determination of which RNA secondary structural elements are supported with statistical significance is often overlooked and is certainly not a standard part of the current work flow for scientists integrating secondary structural data into phylogenetic analyses.
Since RNA secondary structure can serve as such a useful tool for revealing the evolutionary history of certain groups, it is essential that objective criteria be established for incorporating structural elements into phylogenetic inferences. The simple method outlined here allows one (a) to evaluate the probability that each site on an RNA-encoding gene is linked to each other site and (b) to produce an 'elemental' secondary structure model for phylogenetic inference containing only the statistically-supported elements of the structure.
The UNAFold package provides a particularly useful set of tools for exploring various aspects of RNA secondary structure (Markham and Zuker 2008). UNAFold's 'hybrid-ss.exe' yields a set of '.plot' files that give the probability of each base binding to each other base for all reasonable pairings. After installing UNAFold and running 'hybrid-ss.exe' on a FASTA-formatted sequence, one can choose the '.plot' file with the number that most closely approximates the typical cellular temperature (in degrees Celsius) of the organism from which the sequence is derived. This '.plot' file can be modified in Excel by sorting according to 'P(i,j)' values (the probability of pairing) and isolating only the rows for which 'P(i,j)' is above 0.95. This stringent 95% pairing probability cut-off seems most easily justifiable; however, other cut-off values could potentially be used in the context of this method.
For integrating this type of data into a phylogenetic analysis (e.g., using RAxML 7.2.8; http://wwwkramer.in.tum.de/exelixis/software.html; Stamatakis 2006), the standard 'Vienna' dot-bracket notation is used (Hofacker et al. 1994). Any standard secondary structure inference program can be used to create an initial structure that may serve as a template; parentheses can be converted to periods using a standard text-editor or secondary structure editing program (e.g., 4SALE; http://4sale.bioapps.biozentrum.uni-wuerzburg.de/; Seibel et al. 2006) for sites whose linkage is statistically non-significant. These procedures will produce a secondary structure model that includes only the statistically supported elements of structure. When this 'elemental' secondary structure model is incorporated into phylogenetic analyses, it could serve to decrease the degree of uncertainty inserted into the standard secondary structure-based inferences.
Future advances may allow the integration of various intermediate linkage probabilities to be considered in the calculation of tree likelihoods. However, it seems that certain theoretical hurdles remain to be overcome before this type of analysis can be possible. Meanwhile, a methodology like the one outlined here could be beneficial if one wishes to reduce the amount of chance introduced into phylogenetic analyses while still accounting for the fact that certain sites are inextricably linked.
- Brendan
---------------------------------------------------
References
Gutell, R. R., A. Power, G. Z. Hertz, E. J. Putz, and G. D. Stormo. 1992. Identifying constraints on the higher-order structure of RNA: continued development and application of comparative sequence analysis methods. Nucleic Acids Research 20(21): 5785–5795.
Hodkinson, B. P., and J. C. Lendemer. In review. Systematics of a enigmatic sterile crustose lichen.
Hofacker, I. L., W. Fontana, P. F. Stadler, S. Bonhoeffer, M. Tacker, and P. Schuster. 1994. Fast folding and comparison of RNA secondary structures. Monatshefte für Chemie / Chemical Monthly 125: 167-188.
Kimura, M. 1985. The role of compensatory neutral mutations in molecular evolution. Journal of Genetics 64(1):7-19.
Kjer, K. M. 1995. Use of rRNA secondary structure in phylogenetic studies to identify homologous positions: an example of alignment and data presentation from frogs. Molecular Phylogenetics and Evolution 4: 314-330.
Lendemer, J. C., and B. P. Hodkinson. 2009. The wisdom of fools: new molecular and morphological insights into the North American apodetiate species of Cladonia. Opuscula Philolichenum 7: 79-100.
Markham, N., and M. Zuker. 2008. UNAFold: software for nucleic acid folding and hybridization. Methods in Molecular Biology 453: 3-31.
Mathews, D. H., and D. H. Turner. 2006. Prediction of RNA secondary structure by free energy minimization. Journal of Molecular Biology 16(3): 270-278.
Savill N. J., D. C. Hoyle, and P. G. Higgs. 2001. RNA sequence evolution with secondary structure constraints: comparison of substitution rate models using maximum likelihood methods. Genetics 157: 399-411.
Seibel P. N., T. Müller, T. Dandekar, J. Schultz, and M. Wolf. 2006. 4SALE - A tool for synchronous RNA sequence and secondary structure alignment and editing. BMC Bioinformatics 7: 498.
Stamatakis, A. 2006. RAxML-VI-HPC: maximum likelihood- based phylogenetic analyses with thousands of taxa and mixed models. Bioinformatics 22: 2688-2690.
Telford, M., M. Wise, and V. Gowri-Shankar. 2005. Consideration of RNA secondary structure significantly improves likelihood-based estimates of phylogeny: examples from the Bilateria. Molecular Biology and Evolution 22: 1129-1136.
Tillier, E. R. M., and R. A. Collins. 1995. Neighbor-joining and maximum likelihood with RNA sequences: addressing interdependence of sites. Molecular Biology and Evolution 12: 7-15.
Ullrich, B., K. Reinhold, O. Niehuis, and B. Misof. 2010. Secondary structure and phylogenetic analysis of the internal transcribed spacers 1 and 2 of bush crickets (Orthoptera: Tettigoniidae: Barbitistini). Journal of Zoological Systematics and Evolutionary Research 48(3): 219-228.
---------------------------------------------------
This article can be cited as:
Hodkinson, B. P. 2011. Building linkage-probability-based RNA secondary structure models for phylogenetic inference. Squamules Unlimited, New York. [Available at: http://squamules.blogspot.com/2011/11/building-linkage-probability-based-rna.html]
For some additional background on RNA secondary structure, see this previous post:
http://squamules.blogspot.com/2011/08/its-rna-secondary-structure.html
Functional RNA molecules include pairs of nucleotide sites that are linked to one another physically, resulting in specific secondary structures that define the shape of each molecule. This linkage causes certain sites to evolve in tandem with their counterparts. As such, the secondary structure of RNA molecules has been recognized for some time as a significant consideration in the inference of phylogenies from functional RNA-encoding genes (Kimura 1985, Tillier and Collins 1995). Typically, RNA secondary structure is used to optimize multiple-sequence-alignment accuracy for functional RNA-encoding genes (Gutell et al. 1992, Kjer 1995, Lendemer and Hodkinson 2009). However, in recent years, the use of secondary structure in modeling evolution for likelihood-based phylogenetic inferences has begun to gain popularity (Hodkinson and Lendemer in review, Savill et al. 2001, Telford et al. 2005). This approach requires defining the pairs of linked sites on the encoded RNA molecule and treating these pairs as states that are separate from the standard independent nucleotide states (A, C, T, G) in a phylogenetic inference. This method allows one to properly account for the interdependency of interacting nucleotides, since paired RNA nucleotides are no longer required to be treated as independent sites, leading to a more accurate approach for modeling sequence evolution.
Current protocols for integrating RNA secondary structure data into phylogenetic analyses require a single hypothetical structure to be used as an input. Structures are typically inferred using algorithms that minimize free energy or use other thermodynamic considerations to produce the best single structural inference (Mathews & Turner 2006). However, RNA secondary structure is dynamic, frequently changing in the cell as RNA catalyzes reactions and performs various cellular functions. When RNA molecules encounter certain enzymes and cellular components, the thermodynamic rules that previously favored one structure might strongly favor another. Additionally, different methods of structural inference are not always comparable, and small differences in algorithms can favor significantly different structures; the use of differing structural models in a phylogenetic context can have consequences in terms of both topology inference and the calculation of support (Ullrich et al. 2010).
These problems can largely be solved by removing statistically non-significant linkages from phylogenetic analyses, leaving only the most probable structural elements to be incorporated into downstream inferences. The determination of which RNA secondary structural elements are supported with statistical significance is often overlooked and is certainly not a standard part of the current work flow for scientists integrating secondary structural data into phylogenetic analyses.
Since RNA secondary structure can serve as such a useful tool for revealing the evolutionary history of certain groups, it is essential that objective criteria be established for incorporating structural elements into phylogenetic inferences. The simple method outlined here allows one (a) to evaluate the probability that each site on an RNA-encoding gene is linked to each other site and (b) to produce an 'elemental' secondary structure model for phylogenetic inference containing only the statistically-supported elements of the structure.
The UNAFold package provides a particularly useful set of tools for exploring various aspects of RNA secondary structure (Markham and Zuker 2008). UNAFold's 'hybrid-ss.exe' yields a set of '.plot' files that give the probability of each base binding to each other base for all reasonable pairings. After installing UNAFold and running 'hybrid-ss.exe' on a FASTA-formatted sequence, one can choose the '.plot' file with the number that most closely approximates the typical cellular temperature (in degrees Celsius) of the organism from which the sequence is derived. This '.plot' file can be modified in Excel by sorting according to 'P(i,j)' values (the probability of pairing) and isolating only the rows for which 'P(i,j)' is above 0.95. This stringent 95% pairing probability cut-off seems most easily justifiable; however, other cut-off values could potentially be used in the context of this method.
For integrating this type of data into a phylogenetic analysis (e.g., using RAxML 7.2.8; http://wwwkramer.in.tum.de/exelixis/software.html; Stamatakis 2006), the standard 'Vienna' dot-bracket notation is used (Hofacker et al. 1994). Any standard secondary structure inference program can be used to create an initial structure that may serve as a template; parentheses can be converted to periods using a standard text-editor or secondary structure editing program (e.g., 4SALE; http://4sale.bioapps.biozentrum.uni-wuerzburg.de/; Seibel et al. 2006) for sites whose linkage is statistically non-significant. These procedures will produce a secondary structure model that includes only the statistically supported elements of structure. When this 'elemental' secondary structure model is incorporated into phylogenetic analyses, it could serve to decrease the degree of uncertainty inserted into the standard secondary structure-based inferences.
Future advances may allow the integration of various intermediate linkage probabilities to be considered in the calculation of tree likelihoods. However, it seems that certain theoretical hurdles remain to be overcome before this type of analysis can be possible. Meanwhile, a methodology like the one outlined here could be beneficial if one wishes to reduce the amount of chance introduced into phylogenetic analyses while still accounting for the fact that certain sites are inextricably linked.
- Brendan
---------------------------------------------------
References
Gutell, R. R., A. Power, G. Z. Hertz, E. J. Putz, and G. D. Stormo. 1992. Identifying constraints on the higher-order structure of RNA: continued development and application of comparative sequence analysis methods. Nucleic Acids Research 20(21): 5785–5795.
Hodkinson, B. P., and J. C. Lendemer. In review. Systematics of a enigmatic sterile crustose lichen.
Hofacker, I. L., W. Fontana, P. F. Stadler, S. Bonhoeffer, M. Tacker, and P. Schuster. 1994. Fast folding and comparison of RNA secondary structures. Monatshefte für Chemie / Chemical Monthly 125: 167-188.
Kimura, M. 1985. The role of compensatory neutral mutations in molecular evolution. Journal of Genetics 64(1):7-19.
Kjer, K. M. 1995. Use of rRNA secondary structure in phylogenetic studies to identify homologous positions: an example of alignment and data presentation from frogs. Molecular Phylogenetics and Evolution 4: 314-330.
Lendemer, J. C., and B. P. Hodkinson. 2009. The wisdom of fools: new molecular and morphological insights into the North American apodetiate species of Cladonia. Opuscula Philolichenum 7: 79-100.
Markham, N., and M. Zuker. 2008. UNAFold: software for nucleic acid folding and hybridization. Methods in Molecular Biology 453: 3-31.
Mathews, D. H., and D. H. Turner. 2006. Prediction of RNA secondary structure by free energy minimization. Journal of Molecular Biology 16(3): 270-278.
Savill N. J., D. C. Hoyle, and P. G. Higgs. 2001. RNA sequence evolution with secondary structure constraints: comparison of substitution rate models using maximum likelihood methods. Genetics 157: 399-411.
Seibel P. N., T. Müller, T. Dandekar, J. Schultz, and M. Wolf. 2006. 4SALE - A tool for synchronous RNA sequence and secondary structure alignment and editing. BMC Bioinformatics 7: 498.
Stamatakis, A. 2006. RAxML-VI-HPC: maximum likelihood- based phylogenetic analyses with thousands of taxa and mixed models. Bioinformatics 22: 2688-2690.
Telford, M., M. Wise, and V. Gowri-Shankar. 2005. Consideration of RNA secondary structure significantly improves likelihood-based estimates of phylogeny: examples from the Bilateria. Molecular Biology and Evolution 22: 1129-1136.
Tillier, E. R. M., and R. A. Collins. 1995. Neighbor-joining and maximum likelihood with RNA sequences: addressing interdependence of sites. Molecular Biology and Evolution 12: 7-15.
Ullrich, B., K. Reinhold, O. Niehuis, and B. Misof. 2010. Secondary structure and phylogenetic analysis of the internal transcribed spacers 1 and 2 of bush crickets (Orthoptera: Tettigoniidae: Barbitistini). Journal of Zoological Systematics and Evolutionary Research 48(3): 219-228.
---------------------------------------------------
This article can be cited as:
Hodkinson, B. P. 2011. Building linkage-probability-based RNA secondary structure models for phylogenetic inference. Squamules Unlimited, New York. [Available at: http://squamules.blogspot.com/2011/11/building-linkage-probability-based-rna.html]
Monday, October 31, 2011
Visiting Harvard
Last week I visited Harvard University and gave a talk on my research into the consortia of bacteria associated with lichens. I was invited by Anne Pringle and was really excited to see some of the great research going on in her lab. One highlight of the visit was a tour of the Farlow Herbarium with Don Pfister. It was interesting to discuss how the herbarium would implement procedures for the NSF-funded TCN North American Lichens and Bryophytes project, for which I am the coodinator at the New York Botanical Garden (NSF Awards - NYBG: EF-1115086; Harvard: EF-1114957). Soon we will be putting online YouTube videos detailing procedures and workflows for this project, which will hopefully help other institutions such as the Farlow Herbarium to get the project up and running quickly once they are ready to begin. I hope that this visit will represent the first step in establishing lasting collaborations with some of the biologists at Harvard!
- Brendan
- Brendan
Thursday, October 13, 2011
Lichen Conservation
I saw this story in the news and thought that it was interesting:
http://www.bbc.co.uk/news/uk-wales-south-west-wales-15180179
It talks about cutting down trees at a location in Wales in order to engineer a park so that it will remain an optimal location for lichen growth. I am very interested in lichen conservation, but I typically take the stance that we should just "leave it alone" (i.e., let the preserved places evolve naturally) rather than actively attempting to create, maintain, or otherwise engineer a suitable habitat for lichens. But perhaps in parts of the world where there are only a few small, fragmented pieces of land on which lichens can thrive, a different approach must be taken (so that the preserved places remain more or less static instead of evolving naturally into old-growth habitats). It's not something that I would jump into without reservations, but perhaps this instance will be a good test case. I still instinctively remain wary of anything that could set off a chain of interventions that may lead us where we never intended to go! What are your thoughts?
- Brendan
http://www.bbc.co.uk/news/uk-wales-south-west-wales-15180179
It talks about cutting down trees at a location in Wales in order to engineer a park so that it will remain an optimal location for lichen growth. I am very interested in lichen conservation, but I typically take the stance that we should just "leave it alone" (i.e., let the preserved places evolve naturally) rather than actively attempting to create, maintain, or otherwise engineer a suitable habitat for lichens. But perhaps in parts of the world where there are only a few small, fragmented pieces of land on which lichens can thrive, a different approach must be taken (so that the preserved places remain more or less static instead of evolving naturally into old-growth habitats). It's not something that I would jump into without reservations, but perhaps this instance will be a good test case. I still instinctively remain wary of anything that could set off a chain of interventions that may lead us where we never intended to go! What are your thoughts?
- Brendan
Tuesday, September 27, 2011
Unravelling Lecidea
Recently, I co-authored a paper (Schmull et al. 2011) in which we presented the results of analyses aimed at determining the phylogenetic placement of numerous lineages of lichen-forming fungi that were previously placed in the genus Lecidea based solely on morphology. For a long time, it has been known that the assemblage of species placed in Lecidea by Zahlbruckner did not form a single evolutionary lineage. However, placing all of the species in known families has been problematic. For our paper, we conducted two separate 6-gene analyses of lichen-forming fungi in the class Lecanoromycetes in order to infer the placement of twenty-five Lecidea taxa. Most species fell within three families: Lecanoraceae, Pilocarpaceae, and Lecideaceae (the familiy of the 'real' Lecidea). Those within the first two families will unquestionably need to be given new generic names in the near future. The main story that I hope will come out of this paper is that there is much more work to be done! We have used molecular data to demonstrate the scope of the problem with the genus Lecidea, but the definitive placement of all described species will require a great deal of additional study. I'm looking forward to continuing work on this group in the future!
- Brendan
-----------------------------------------------
Reference
Schmull, M., J. Miadlikowska, M. Pelzer, E. Stocker-Wörgötter, V. Hofstetter, E. Fraker, B. P. Hodkinson, V. Reeb, M. Kukwa, H. T. Lumbsch, F. Kauff, and F. Lutzoni. 2011. Phylogenetic affiliations of members of the heterogeneous lichen-forming genus Lecidea sensu Zahlbruckner (Lecanoromycetes, Ascomycota). Mycologia 103(5): 983-1003.
Download publication (PDF file)
Download nucleotide alignment (NEXUS file)
Download supplementary data table 1 (PDF file)
Download supplementary data table 2 (PDF file)
- Brendan
-----------------------------------------------
Reference
Schmull, M., J. Miadlikowska, M. Pelzer, E. Stocker-Wörgötter, V. Hofstetter, E. Fraker, B. P. Hodkinson, V. Reeb, M. Kukwa, H. T. Lumbsch, F. Kauff, and F. Lutzoni. 2011. Phylogenetic affiliations of members of the heterogeneous lichen-forming genus Lecidea sensu Zahlbruckner (Lecanoromycetes, Ascomycota). Mycologia 103(5): 983-1003.
Download publication (PDF file)
Download nucleotide alignment (NEXUS file)
Download supplementary data table 1 (PDF file)
Download supplementary data table 2 (PDF file)
Wednesday, September 14, 2011
Musical Lichenology
I recently received an email from Sean Beeching, famous to readers of this blog for his poetry (click here and here for samples). This is what he wrote:
"Here is a video that Nancy Lowe of discoverlife shot during our lichen workshop [http://www.youtube.com/watch?v=rvxnv-6Z6rg]. I am sawing up branches to show the students the lichens that were growing on them in time to Tommy [Jordan]'s banjo playing. He and I play together in the evenings during the workshop. You might also have a look at the lichen key I made for our students at discoverlife, go to the nature guides at the discoverlife.org website and then page down to 'Lichens, Georgia.' The site just passed a billion hits."
For anyone in the southeastern US interested in learning a thing or two about lichens, I would highly recommend any workshop with Sean!
- Brendan
"Here is a video that Nancy Lowe of discoverlife shot during our lichen workshop [http://www.youtube.com/watch?v=rvxnv-6Z6rg]. I am sawing up branches to show the students the lichens that were growing on them in time to Tommy [Jordan]'s banjo playing. He and I play together in the evenings during the workshop. You might also have a look at the lichen key I made for our students at discoverlife, go to the nature guides at the discoverlife.org website and then page down to 'Lichens, Georgia.' The site just passed a billion hits."
For anyone in the southeastern US interested in learning a thing or two about lichens, I would highly recommend any workshop with Sean!
- Brendan
Wednesday, September 7, 2011
Diversity of Lichenology
Not too long ago I reviewed the 100th anniversary issue of the journal/book series Bibliotheca Lichenologica for The Bryologist. Here is what I wrote:
The 100th volume of Bibliotheca Lichenologica (‘‘Diversity of Lichenology — Anniversary Volume’’) provides important contributions to the field and gives us further insights into the biology of lichens while connecting us to our historical roots. While the volumes of this series have taken many different forms, this edition appears as a standard journal volume, with 18 scientific and historical articles from a total of 37 authors representing a diverse array of lichenologists. It should be noted that the emphasis is on lichens of Eurasia and/or the Southern Hemisphere; however, many of the articles will appeal to a general, worldwide audience.
In terms of taxonomy, this volume will be of particular interest to those following the changes in the family Teloschistaceae. Kondratyuk et al. describe 35 new species in the family. Many of these are known only from the type locality, or a small handful of specimens, making further evaluation of some of the taxa difficult, although the authors are to be commended for providing excellent color photographs of the thalli. The work by Fedorenko et al. focusing on the phylogeny of ‘xanthorioid’ lichens represents a significant contribution in terms of both the data generated and the provisional generic concepts articulated. However, both of the aforementioned works seem to highlight the fact that the largest task still remains: an integrated systematic revision of the family Teloschistaceae that includes crustose, foliose, and fruticose forms. In different ways, these works both improve our understanding of this family for which the taxonomy continues to evolve.
This volume also includes important insights into the taxonomy of the oft-neglected lichenicolous fungi. Hafellner presents an excellent ‘traditional’ treatment of the lichenicolous genera Phacothecium and Phacographa. The work is notably thorough, articulating precisely what is known about the genera, while highlighting areas where additional data and analyses are needed. The author also provides a useful key to opegraphoid lichenicolous fungi with widely exposed hymenia, along with a summary table of the phenotypic characters separating the five opegraphoid arthonialean genera with lichenicolous members discussed in the text (i.e., Opegrapha, Lecanographa, Phacothecium, Phacographa, and Plectocarpon).
Among the works that will appeal to a broader audience is Kärnefelt’s contribution entitled ‘‘Fifty influential lichenologists.’’ This portion of the volume provides a veritable ‘‘Who’s Who’’ of lichenology. The careers of some of the world’s major players in our field, both modern and historical, are briefly summarized, starting with ‘‘the father of lichenology,’’ Erik Acharius. Another work with broad appeal is by Lücking et al., entitled ‘‘How many tropical lichens are there… really?’’ This piece discusses the various factors involved in calculating a ‘ballpark’ estimate of the overall diversity of lichens in the neotropics, the tropics in general, and the globe as a whole (the estimate for the latter comes in at 28,000 species!). Although any estimate is subjective in nature, various points that have not been explicitly integrated into previous estimates are considered (e.g., taxonomic ‘orphans,’ species pairs, photosymbiodemes, chemotypes, and cryptic species). Another portion of the volume that will appeal to amateurs and professionals alike is the section by Randlane et al., which provides what is undoubtedly the best synthetic work on the species of Usnea for the continent of Europe. Range maps and photographs of small-scale features make this section both informative and interesting.
Reading this volume also makes it apparent that the changing landscape of lichenological research has led to certain problems that require special attention. Many of the problematic issues plaguing our field seem to be associated with the process of adjusting to the molecular age. A number of the studies published herein leave the reader wanting to know more, especially in terms of molecular data and how they were analyzed (or how they could be analyzed differently). Beyond the simple deposition of sequences in a public repository, alignments must be reviewed and made available if alignment-based phylogenetic analyses are to be reproducible. Authors will find that making their assembled molecular datasets freely available to readers (on their own personal websites if necessary) increases the impact and relevance of their work by permitting others to easily build on their studies. Ultimately, this practice will allow our field to advance more quickly and will raise the bar in terms of research quality. In summary, this volume of Bibliotheca Lichenologica, as its title suggest, provides an excellent picture of the diversity of lichenology and represents quite well the overall state of the field as we enter this next decade. Many of the works contained in this anniversary edition provide important contributions, and any lichenologist’s collection would be enriched by the addition of this volume.
- Brendan
-----------------------------
Citation:
Hodkinson, B. P. 2010. Lichenological Diversity. The Bryologist 113(4): 828-829.
The 100th volume of Bibliotheca Lichenologica (‘‘Diversity of Lichenology — Anniversary Volume’’) provides important contributions to the field and gives us further insights into the biology of lichens while connecting us to our historical roots. While the volumes of this series have taken many different forms, this edition appears as a standard journal volume, with 18 scientific and historical articles from a total of 37 authors representing a diverse array of lichenologists. It should be noted that the emphasis is on lichens of Eurasia and/or the Southern Hemisphere; however, many of the articles will appeal to a general, worldwide audience.
In terms of taxonomy, this volume will be of particular interest to those following the changes in the family Teloschistaceae. Kondratyuk et al. describe 35 new species in the family. Many of these are known only from the type locality, or a small handful of specimens, making further evaluation of some of the taxa difficult, although the authors are to be commended for providing excellent color photographs of the thalli. The work by Fedorenko et al. focusing on the phylogeny of ‘xanthorioid’ lichens represents a significant contribution in terms of both the data generated and the provisional generic concepts articulated. However, both of the aforementioned works seem to highlight the fact that the largest task still remains: an integrated systematic revision of the family Teloschistaceae that includes crustose, foliose, and fruticose forms. In different ways, these works both improve our understanding of this family for which the taxonomy continues to evolve.
This volume also includes important insights into the taxonomy of the oft-neglected lichenicolous fungi. Hafellner presents an excellent ‘traditional’ treatment of the lichenicolous genera Phacothecium and Phacographa. The work is notably thorough, articulating precisely what is known about the genera, while highlighting areas where additional data and analyses are needed. The author also provides a useful key to opegraphoid lichenicolous fungi with widely exposed hymenia, along with a summary table of the phenotypic characters separating the five opegraphoid arthonialean genera with lichenicolous members discussed in the text (i.e., Opegrapha, Lecanographa, Phacothecium, Phacographa, and Plectocarpon).
Among the works that will appeal to a broader audience is Kärnefelt’s contribution entitled ‘‘Fifty influential lichenologists.’’ This portion of the volume provides a veritable ‘‘Who’s Who’’ of lichenology. The careers of some of the world’s major players in our field, both modern and historical, are briefly summarized, starting with ‘‘the father of lichenology,’’ Erik Acharius. Another work with broad appeal is by Lücking et al., entitled ‘‘How many tropical lichens are there… really?’’ This piece discusses the various factors involved in calculating a ‘ballpark’ estimate of the overall diversity of lichens in the neotropics, the tropics in general, and the globe as a whole (the estimate for the latter comes in at 28,000 species!). Although any estimate is subjective in nature, various points that have not been explicitly integrated into previous estimates are considered (e.g., taxonomic ‘orphans,’ species pairs, photosymbiodemes, chemotypes, and cryptic species). Another portion of the volume that will appeal to amateurs and professionals alike is the section by Randlane et al., which provides what is undoubtedly the best synthetic work on the species of Usnea for the continent of Europe. Range maps and photographs of small-scale features make this section both informative and interesting.
Reading this volume also makes it apparent that the changing landscape of lichenological research has led to certain problems that require special attention. Many of the problematic issues plaguing our field seem to be associated with the process of adjusting to the molecular age. A number of the studies published herein leave the reader wanting to know more, especially in terms of molecular data and how they were analyzed (or how they could be analyzed differently). Beyond the simple deposition of sequences in a public repository, alignments must be reviewed and made available if alignment-based phylogenetic analyses are to be reproducible. Authors will find that making their assembled molecular datasets freely available to readers (on their own personal websites if necessary) increases the impact and relevance of their work by permitting others to easily build on their studies. Ultimately, this practice will allow our field to advance more quickly and will raise the bar in terms of research quality. In summary, this volume of Bibliotheca Lichenologica, as its title suggest, provides an excellent picture of the diversity of lichenology and represents quite well the overall state of the field as we enter this next decade. Many of the works contained in this anniversary edition provide important contributions, and any lichenologist’s collection would be enriched by the addition of this volume.
- Brendan
-----------------------------
Citation:
Hodkinson, B. P. 2010. Lichenological Diversity. The Bryologist 113(4): 828-829.
Wednesday, August 31, 2011
ITS RNA secondary structure
I have recently been conducting phylogenetic and taxonomic studies of selected groups of lichen-forming fungi using sequences from the quickly evolving nuclear ribosomal ITS (internal transcribed spacer) region to examine relationships within and between species (e.g., Hodkinson & Lendemer 2011, Hodkinson et al. 2010, Lendemer & Hodkinson 2009, 2010, in prep). In order to properly analyze the evolutionary relationships between the organisms from which these molecules were derived, I built secondary structure models for the RNA molecules encoded by ITS1 and ITS2 (the two rapidly evolving sections of the ITS region) for some of the groups.
The ITS1 and ITS2 spacer regions encode stretches of RNA that fold up in specific conformations and help to assemble the ribosomes (the pieces of cellular machinery that build protein molecules based on specific messenger RNA sequences transcribed from DNA). The particular folding pattern is referred to as the molecule's "secondary structure." Here is an example of a secondary structure model that I put together for ITS2 of Parmotrema perforatum:
Notice the A(adenine)-U(uracil) pairings and the G(guanine)-C(cytosine) pairings, just like the complementary strands of DNA (except that with DNA you have T for thymine instead of U for uracil).
There are two main reasons that one might want to have a secondary structure model when inferring phylogeny:
[1] Nucleotide Alignment - An understanding of the overall structure of the molecule can aid in discerning which sets of sites in different organisms actually represent the same character when they have different states and there are adjacent nucleotides that have been inserted or deleted in some taxa (Kjer 1995). Many studies use principles of secondary structure to aid in alignment.
[2] Phylogenetic Inference - Since paired sites in some sense evolve in tandem (if one nucleotide changes, the linked nucleotide will often change to compensate over evolutionary time), it is most appropriate within a likelihood framework to apply a different model of evolution to the paired nucleotides so that this can be taken into consideration. This type of inference can be done with RAxML (Stamatakis 2006) and I have recently integrated this into my workflow (Hodkinson & Lendemer in prep).
The really interesting thing to think about is the fact that this type of macromolecule needs to be able to move in order to function, which means that the structure is not actually static, but dynamic. While we usually use the 'best' structure for phylogenetic inference, there are actually many structures that are nearly equally good, and the molecule actually changes its conformation through space and time, flipping between these different conformations in order to perform its functions in the cell. To drive the point home, here is a quick video I made of the ITS2 molecule of Cladonia stipitata Lendemer & Hodkinson (2009) shifting between different likely conformations:
------------------------------
Sources cited:
Download publication (PDF file)
Download nucleotide alignment (NEXUS file)
Hodkinson, B. P., and J. C. Lendemer. In prep. Systematics of a enigmatic sterile crustose lichen.
Hodkinson, B. P., J. C. Lendemer, and T. L. Esslinger. 2010. Parmelia barrenoae, a macrolichen new to North America and Africa. North American Fungi 5(3): 1-5.
Download publication (PDF file)
Kjer, K. M. 1995. Use of rRNA secondary structure in phylogenetic studies to identify homologous positions: an example of alignment and data presentation from the frogs. Molecular Phylogenetics and Evolution 4: 314–330.
Lendemer, J. C., and B. P. Hodkinson. 2009. The Wisdom of Fools: new molecular and morphological insights into the North American apodetiate species of Cladonia. Opuscula Philolichenum 7: 79-100.
Download publication (PDF file)
Download nucleotide alignment (NEXUS file)
Lendemer, J. C., and B. P. Hodkinson. 2010. A new perspective on Punctelia subrudecta in North America: previously-rejected morphological characters corroborate molecular phylogenetic evidence and provide insight into an old problem. The Lichenologist 42(4): 405-421.
Download publication (PDF file)
Download nucleotide alignment (NEXUS file)
Stamatakis, A. 2006. RAxML-VI-HPC: maximum likelihood-based phylogenetic analyses with thousands of taxa and mixed models. Bioinformatics 22: 2688–2690.
Wednesday, August 17, 2011
Taxonomy: Art or Science?
When Googling "science definition," the first thing that came up was "The intellectual and practical activity encompassing the systematic study of the structure and behavior of the physical and natural world through observation and experiment." After a little more research, I was surprised to see that this seems to be one of the stricter definitions of science (others may be as broad as "the state of knowing" or some such...), but it is one with which I can get on board. I tend to think of science itself in a very strict sense, as the process of developing and testing hypotheses. However, my big caveat is that there are many activities that are involved in (and are absolutely essential to) the practice of science that are not science per se according to that definition. This does not diminish their value to science. Some of this has to do with the acquisition of background knowledge that informs the hypotheses to be tested, while some of it is associated with making the results of inquiry available and comprehensible to the scientific community and the public.
So then is taxonomy art or science? With taxonomy, there is not a "right" answer, although there are plenty of wrong answers if one wishes to have a system that is informed by the results of scientific inquiry. Taxonomic units are all in some sense arbitrary. Although a group of organisms may form a "clade," whether we recognize that clade with a certain name is somewhat arbitrary. I personally like to think of taxonomic units being defined by specific innovations (morphological, molecular, ecological, etc.) that have changed the evolutionary trajectory of a group, but that rule is certainly not universally applied, and there could certainly be many alternative taxonomies even if such standards were applied.
For me, the argument for taxonomy as an art does not actually diminish taxonomy in any way as part of what we must do in order to be effective and responsible scientists. In fact, having this perspective on taxonomy can help to enhance the understanding of the significance of taxonomy for science. As scientists, we must use what we discover through the scientific process to help facilitate communication about natural phenomena. Taxonomy is a tool that we use to communicate ideas about organisms, so taxonomy is an absolutely necessary part of the pursuit of scientific truth, even if it is not "science" itself.
One test for me of whether taxonomy is itself a science in the very strictest sense of the word is whether it is directly involved in the process of hypothesis testing. One can use principles of phylogenetics, ecology, or molecular biology to test hypotheses, but taxonomic principles would not be used. When we begin to dissect some of the scientific questions that are often deemed "taxonomic questions," it can be argued that they are not actually taxonomic in nature, and that the taxonomic repercussions would really only be a byproduct of obtaining results through scientific inquiry. For instance, a question like "Is this a good genus?" is really asking something like "Do the species form a distinct clade?", which is a question that is evolutionary in nature. Likewise, the question "Do these individuals make up one species?" is perhaps just a way of saying "How can we properly apply a biological, morphological, chemical, ecological, and/or phylogenetic species concept to this group of individuals?", a question that draws on different fields of biology.
I can see that many systematists would hesitate to state that taxonomy is an art, because of what it implies. If it is an art, then it opens the door for people to say that people who do taxonomy are not really scientists at all. But a consummate scientist is not just someone who constantly tests hypotheses one after another without consideration for anything else. To be a scientist, one must also lay the groundwork for scientific pursuits, and defining the terms used to communicate ideas about specific units of the tree of life (whether or not it is itself an artistic pursuit) is crucial to the advancement of science.
- Brendan
So then is taxonomy art or science? With taxonomy, there is not a "right" answer, although there are plenty of wrong answers if one wishes to have a system that is informed by the results of scientific inquiry. Taxonomic units are all in some sense arbitrary. Although a group of organisms may form a "clade," whether we recognize that clade with a certain name is somewhat arbitrary. I personally like to think of taxonomic units being defined by specific innovations (morphological, molecular, ecological, etc.) that have changed the evolutionary trajectory of a group, but that rule is certainly not universally applied, and there could certainly be many alternative taxonomies even if such standards were applied.
For me, the argument for taxonomy as an art does not actually diminish taxonomy in any way as part of what we must do in order to be effective and responsible scientists. In fact, having this perspective on taxonomy can help to enhance the understanding of the significance of taxonomy for science. As scientists, we must use what we discover through the scientific process to help facilitate communication about natural phenomena. Taxonomy is a tool that we use to communicate ideas about organisms, so taxonomy is an absolutely necessary part of the pursuit of scientific truth, even if it is not "science" itself.
One test for me of whether taxonomy is itself a science in the very strictest sense of the word is whether it is directly involved in the process of hypothesis testing. One can use principles of phylogenetics, ecology, or molecular biology to test hypotheses, but taxonomic principles would not be used. When we begin to dissect some of the scientific questions that are often deemed "taxonomic questions," it can be argued that they are not actually taxonomic in nature, and that the taxonomic repercussions would really only be a byproduct of obtaining results through scientific inquiry. For instance, a question like "Is this a good genus?" is really asking something like "Do the species form a distinct clade?", which is a question that is evolutionary in nature. Likewise, the question "Do these individuals make up one species?" is perhaps just a way of saying "How can we properly apply a biological, morphological, chemical, ecological, and/or phylogenetic species concept to this group of individuals?", a question that draws on different fields of biology.
I can see that many systematists would hesitate to state that taxonomy is an art, because of what it implies. If it is an art, then it opens the door for people to say that people who do taxonomy are not really scientists at all. But a consummate scientist is not just someone who constantly tests hypotheses one after another without consideration for anything else. To be a scientist, one must also lay the groundwork for scientific pursuits, and defining the terms used to communicate ideas about specific units of the tree of life (whether or not it is itself an artistic pursuit) is crucial to the advancement of science.
- Brendan
Thursday, August 4, 2011
Using Sequencher for Multiple Sequence Alignments
Much of the molecular research that I have done over the years has involved working with DNA sequences generated through Sanger sequencing. These sequences are never perfect, and always require manual correction. It is especially helpful to correct sequences and align them to other similar sequences simultaneously. In this way, alignment and structural data can be taken into consideration when interpreting the chromatograms for the DNA sequences.
So I wrote a couple of simple Perl scripts that would allow me to make my alignments in Sequencher (the standard program for editing raw sequence reads) and easily move it over to Mesquite or MacClade (standard programs for assembling data matrices for downstream phylogenetic analyses) so that it could be joined with a reference alignment that I had made previously. In this way, I could avoid completely realigning all sequences to one another through an automatic alignment program, thereby preserving certain sequence alignment patterns (note that I often deal with over 1000 sequences at a time). If you use Linux or Macintosh, running a Perl script is generally a pretty simple matter (since Perl interpreters are typically built into the operating system). If you use Windows, you will probably need to download an interpreter like Strawberry Perl or ActivePerl.
The type of data that I was dealing with was a set of bidirectional Sanger sequences (one forward, one reverse primer for each sequence) of fragments ~650 bp in length. These sequences were cloned and therefore had vector overhang on both ends of both strands, which had to be deleted. If you have data that are similar, here is a procedure that can be used to preserve the Sequencher alignment pattern and bring it into MacClade/Mesquite (potentially for merging with a curated reference alignment, if you have one of these):
[a] In the Sequencher alignment, make sure at least one sequencing strand of each pair of strands (from the bidirectionally-sequenced pool of DNA fragments) has all of the corrected bases, and delete the second strand for each pair. This gives an alignment with one strand for each sequence. [This Sequencher alignment can be tweaked visually to align with a reference set that is already pre-aligned by introducing gaps into the Sequencher alignment to accommodate the gaps in the reference alignment.]
[b] The Sequencher alignment can then be exported as a contig in aligned fasta format and subsequently opened in MacClade/Mesquite. [Note: If you have exported the sequences from Sequencher as a concatenated set of sequence fragments, it might use ':' instead of '-' to represent the gaps; make sure all of the gaps are changed to '-' for integration into MacClade or Mesquite (this can be done as a simple search and replace with any text editor).]
For my particular sequences, I had to deal with the issue of all of the sequence names being proceeded by my initials and having strand-specific information tacked on to the end (both standard pieces of information added by the sequencing facility). Here is another blog post with the Perl script that I wrote for editing the fasta file to extract the 10-digit alpha-numeric code used to identify my sequences. Also, I had to line my sequence block up with the portion of my reference alignment with which it correlated. In my particular situation, the block of sequences that I had aligned began 488 bases into the reference alignment. Here is the script that I used to add 488 bases to the front of each sequence in the fasta file (this script relies on having a 10-digit code name for each sequence):
#!/usr/bin/perl
print "\nPlease type the name of your input file: ";
my $filename = <STDIN>;
chomp $filename;
open (FASTA, $filename);
{
if ($filename =~ /(.*)\.[^.]*/)
{
open OUT, ">$1.ed.fasta";
}
}
while (<FASTA>)
{
if ($_ =~ /^>(..........)/)
{
print OUT "\r>$1\r\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\n";
}
else
{
print OUT $_;
}
}
The final step was to simply open up my reference alignment in MacClade and import the newly-generated fasta file of aligned cloned sequences... and they lined up perfectly! I then tweaked exclusion sets, saved the full alignment, and was ready for downstream phylogenetic analyses.
Even though MacClade and Mesquite are very good programs overall for alignment, aligning a set of 1000+ sequences is extremely cumbersome, and Sequencher can be much faster and easier as long as the sequences are relatively conserved. With this set of Perl scripts discussed above, hopefully researchers will no longer perceive impediments or inefficiency in a process that includes aligning and correcting relatively conserved sequences in Sequencher (with all of the raw sequence data) before moving them over to MacClade/Mesquite for final data set assembly and formatting.
- Brendan
----------------------------------------------
References
The above protocols are published in the following sources:
Hodkinson, B. P. 2011. A Phylogenetic, Ecological, and Functional Characterization of Non-Photoautotrophic Bacteria in the Lichen Microbiome. Doctoral Dissertation, Duke University, Durham, NC.
Download Dissertation (PDF file)
Hodkinson, B. P., N. R. Gottel, C. W. Schadt, and F. Lutzoni. 2011. Data from: Photoautotrophic symbiont and geography are major factors affecting highly structured and diverse bacterial communities in the lichen microbiome. Dryad Digital Repository doi:10.5061/dryad.t99b1.
Hodkinson, B. P., N. R. Gottel, C. W. Schadt, and F. Lutzoni. In press. Photoautotrophic symbiont and geography are major factors affecting highly structured and diverse bacterial communities in the lichen microbiome. Environmental Microbiology.
----------------------------------------------
This work was funded in part by NSF DEB-1011504.
So I wrote a couple of simple Perl scripts that would allow me to make my alignments in Sequencher (the standard program for editing raw sequence reads) and easily move it over to Mesquite or MacClade (standard programs for assembling data matrices for downstream phylogenetic analyses) so that it could be joined with a reference alignment that I had made previously. In this way, I could avoid completely realigning all sequences to one another through an automatic alignment program, thereby preserving certain sequence alignment patterns (note that I often deal with over 1000 sequences at a time). If you use Linux or Macintosh, running a Perl script is generally a pretty simple matter (since Perl interpreters are typically built into the operating system). If you use Windows, you will probably need to download an interpreter like Strawberry Perl or ActivePerl.
The type of data that I was dealing with was a set of bidirectional Sanger sequences (one forward, one reverse primer for each sequence) of fragments ~650 bp in length. These sequences were cloned and therefore had vector overhang on both ends of both strands, which had to be deleted. If you have data that are similar, here is a procedure that can be used to preserve the Sequencher alignment pattern and bring it into MacClade/Mesquite (potentially for merging with a curated reference alignment, if you have one of these):
[a] In the Sequencher alignment, make sure at least one sequencing strand of each pair of strands (from the bidirectionally-sequenced pool of DNA fragments) has all of the corrected bases, and delete the second strand for each pair. This gives an alignment with one strand for each sequence. [This Sequencher alignment can be tweaked visually to align with a reference set that is already pre-aligned by introducing gaps into the Sequencher alignment to accommodate the gaps in the reference alignment.]
[b] The Sequencher alignment can then be exported as a contig in aligned fasta format and subsequently opened in MacClade/Mesquite. [Note: If you have exported the sequences from Sequencher as a concatenated set of sequence fragments, it might use ':' instead of '-' to represent the gaps; make sure all of the gaps are changed to '-' for integration into MacClade or Mesquite (this can be done as a simple search and replace with any text editor).]
For my particular sequences, I had to deal with the issue of all of the sequence names being proceeded by my initials and having strand-specific information tacked on to the end (both standard pieces of information added by the sequencing facility). Here is another blog post with the Perl script that I wrote for editing the fasta file to extract the 10-digit alpha-numeric code used to identify my sequences. Also, I had to line my sequence block up with the portion of my reference alignment with which it correlated. In my particular situation, the block of sequences that I had aligned began 488 bases into the reference alignment. Here is the script that I used to add 488 bases to the front of each sequence in the fasta file (this script relies on having a 10-digit code name for each sequence):
#!/usr/bin/perl
print "\nPlease type the name of your input file: ";
my $filename = <STDIN>
chomp $filename;
open (FASTA, $filename);
{
if ($filename =~ /(.*)\.[^.]*/)
{
open OUT, ">$1.ed.fasta";
}
}
while (<FASTA>
{
if ($_ =~ /^>(..........)/)
{
print OUT "\r>$1\r\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\-\n";
}
else
{
print OUT $_;
}
}
The final step was to simply open up my reference alignment in MacClade and import the newly-generated fasta file of aligned cloned sequences... and they lined up perfectly! I then tweaked exclusion sets, saved the full alignment, and was ready for downstream phylogenetic analyses.
Even though MacClade and Mesquite are very good programs overall for alignment, aligning a set of 1000+ sequences is extremely cumbersome, and Sequencher can be much faster and easier as long as the sequences are relatively conserved. With this set of Perl scripts discussed above, hopefully researchers will no longer perceive impediments or inefficiency in a process that includes aligning and correcting relatively conserved sequences in Sequencher (with all of the raw sequence data) before moving them over to MacClade/Mesquite for final data set assembly and formatting.
- Brendan
----------------------------------------------
References
The above protocols are published in the following sources:
Hodkinson, B. P. 2011. A Phylogenetic, Ecological, and Functional Characterization of Non-Photoautotrophic Bacteria in the Lichen Microbiome. Doctoral Dissertation, Duke University, Durham, NC.
Download Dissertation (PDF file)
Hodkinson, B. P., N. R. Gottel, C. W. Schadt, and F. Lutzoni. 2011. Data from: Photoautotrophic symbiont and geography are major factors affecting highly structured and diverse bacterial communities in the lichen microbiome. Dryad Digital Repository doi:10.5061/dryad.t99b1.
Hodkinson, B. P., N. R. Gottel, C. W. Schadt, and F. Lutzoni. In press. Photoautotrophic symbiont and geography are major factors affecting highly structured and diverse bacterial communities in the lichen microbiome. Environmental Microbiology.
----------------------------------------------
This work was funded in part by NSF DEB-1011504.
Monday, July 25, 2011
Taxonomy as Wastewater Treatment
I was thinking about taxonomic treatments of specific groups of organisms; in some cases a 'treatment' is essentially a rehash and synthesis of what has been published previously (but just for a specific subregion, etc). While focusing on the mechanistic details of how my colleagues and I put together treatments (not through rehashing), I thought of the process by which sewage is treated. So I did a Google image search for "taxonomic treatment," and basically got a bunch of photos of journal pages/covers, some phylogenetic trees, and some photos of organisms. I think that's pretty much how most people see biological taxonomy, which explains a lot about why taxonomy is often seen as dull and sometimes (even worse) unscientific. But then I did a Google image search for "wastewater treatment" and I was amazed to see that the images there generally matched my concept of how to do a taxonomic treatment much better than what I saw in the previous search! What I saw were mainly flowcharts... they showed processes like screening, pre-treatment, cleaning, clarification, digestion, storage, disposal... and the processes all flowed into one another and ended up making products for public consumption! Yes! This is it! This is how we really need to be doing taxonomy! Instead of perpetuating the problems that exist, take them head on... get the junk out of the way, and make something that people can use. Many researchers seem to be of the mind that if it's mostly right, it's good enough; but as any wastewater treatment plant manager will tell you, even if it's only 10-20% sewage, it's not fit for public consumption. Let us view our taxonomy in the same manner!
Lichenologist James Lendemer is famously quoted as saying "I think of myself as a bounty hunter." Perhaps I should think of myself as a manager of a wastewater treatment plant. Maybe that's not as glorious, but it certainly is important. So much work remains to be done before we get close to having a reasonable set of names for the organisms on the Earth. As long as humans are involved, our nomenclatural system will be imperfect and will require constant cleaning, management, and enforcement of standards... and there I will stand, ready to take on the nastiest and dirtiest of the problems!
- Brendan
P.S. Some recent big news in the NYC area has been the big fire at a wastewater treatment plant that sent sewage spewing into the Hudson River. Thought exercise for taxonomists: Can you think of any events like this one (speaking metaphorically) that affected your particular group of organisms?
Lichenologist James Lendemer is famously quoted as saying "I think of myself as a bounty hunter." Perhaps I should think of myself as a manager of a wastewater treatment plant. Maybe that's not as glorious, but it certainly is important. So much work remains to be done before we get close to having a reasonable set of names for the organisms on the Earth. As long as humans are involved, our nomenclatural system will be imperfect and will require constant cleaning, management, and enforcement of standards... and there I will stand, ready to take on the nastiest and dirtiest of the problems!
- Brendan
P.S. Some recent big news in the NYC area has been the big fire at a wastewater treatment plant that sent sewage spewing into the Hudson River. Thought exercise for taxonomists: Can you think of any events like this one (speaking metaphorically) that affected your particular group of organisms?
Tuesday, July 12, 2011
Man It Feels Good 2 B A Lichen
Recently I was thinking about the plight of the sterile crustose lichens, specifically those in Eastern North America. One could feel sorry for them, being so taxonomically neglected and underrepresented in all major surveys of biodiversity. But in a certain kind of way, I think that they must be very proud. It's certainly amazing how they've been able to get by on so little (so little sex, so little attention from humans). Inspired by the story of the sterile crusts, I decided to write some lyrics, which I entitled "Man, It Feels Good 2 B A Lichen," and set the words to the music of the similarly-titled song (made famous by the movie Office Space) by the Geto Boys, about being a gangsta, not a lichen. Some people have said that it's groundbreaking... that it's a whole new genre (known as "Lichen Rap" or, more commonly, "Lich-Hop")... I just like to think of it as one of the products resulting from the inspiration that I get from the amazing organisms that I study!
Man, It Feels Good 2 B A Lichen (a sterile crust’s song)
Man, it feels good to be a lichen
Live sterile crust lichens ain’t indoors
Real sterile crust lichens are hard to identify
‘cause sterile crust lichens don’t make spores
Man, it feels good to be a lichen
I mean one that you don’t really know
Livin’ as a sterile crust, drivin’ people crazy
‘cause I can’t be identified for sho’
Now sterile crust lichens come in all shapes and colors
Some got killed in the past
But if NSF could just get their back with some fundin’
We could study them and hopefully they’ll last
Now all I gotta say to you
Sexually-reproducin’ lichen-formin’ fungi recombinin’
When your spore can’t find no algae what you think you gonna do?
Man, it feels good to be a lichen
Man, it feels good to be a lichen
Flying round on the currents all day
‘cause when a sterile crust lichen goes and tries to reproduce
It bundles up the partners and it blows away
Now when a lichen like this one is livin’ in your ‘hood
It’s most likely that you ain’t gonna know
‘cause lots of sterile crust lichens are small and inconspicuous
But under UV they might glow
Now all I gotta say to you
Sexually-reproducin’ lichen-formin’ fungi recombinin’
When your spore can’t find no algae what you think you gonna do?
Man, it feels good to be a lichen
Stay tuned for news about the record release party in the Bronx, to take place later this year!
- Brendan
Tuesday, July 5, 2011
Perl: Renaming DNA Sequences
When I began my dissertation studies, I did not know the wonders of the Perl programming language. However, within the past year, it has proven to be an invaluable tool for manipulating DNA sequence data sets and helping me to tackle projects that once seemed too large in scope. In this post I will just give one example of a Perl script that I wrote after getting some training from Dr. Bob Thomson, who I met at the NSF-Sponsored "Fast, Free Phylogenies" workshop at NIMBioS (Knoxville, TN).
When processing large numbers of DNA sequences, it always helps to have a standardized naming system so that the sequences can be handled in an automated way during downstream analyses. For a recent large-scale cloning experiment that involved picking 2880 clonal bacterial colonies (to amplify and sequence a vector-inserted 16S gene fragment from each), I developed a 10-digit alpha-numeric code that allowed me to encode all of the necessary data about my sequences into each specific sequence identifier. However, the sequencing facility also needed to use its own codes to keep track of my sequences, so I ended up with long names that had my own codes in the middle with information identifying them as my sequences in front and information about the individual sequence reads themselves tacked on at the end. Therefore, to recover the names (without retaining the sequencing facility's additions) in an automated way, I wrote a simple Perl script to edit a fasta file containing the sequences (this was run after the process of manual sequence correction had been finished).
The following script (‘Clon_16S_fasta_renamer.pl’) allowed me to extract the 10-digit alpha-numeric codes that I used in my dissertation studies (Hodkinson 2011) from the long names (with extraneous information) that come from the sequencing facility. It creates a new fasta file with these modified identifiers. Specifically, it takes sequences that have "BH_" (my initials), followed by a 10-digit code, followed by additional characters, and simply renames each sequence using just the 10-digit code (effectively stripping out "BH_" at the beginning and and extra characters at the end). The new file will have the same name, but the extension will be replaced by ".ed.fasta". This can be easily modified for any set of sequences that are identified using a standardized naming scheme.
If you need to know how to run a Perl script, you can look it up on Google, but here is one example of how to run a Perl script using Windows (it's actually easier on almost any other type of operating system). Since I was performing a simple task with a very specific data set, it was easy for me to use basic Perl commands. However, for more complex sequence manipulations, BioPerl provides an excellent collection of Perl modules for biological applications.
- Brendan
----------------------------------------------
References
The above script is published in the following sources:
Hodkinson, B. P. 2011. A Phylogenetic, Ecological, and Functional Characterization of Non-Photoautotrophic Bacteria in the Lichen Microbiome. Doctoral Dissertation, Duke University, Durham, NC.
Download Dissertation (PDF file)
Hodkinson, B. P., N. R. Gottel, C. W. Schadt, and F. Lutzoni. 2011. Data from: Photoautotrophic symbiont and geography are major factors affecting highly structured and diverse bacterial communities in the lichen microbiome. Dryad Digital Repository doi:10.5061/dryad.t99b1.
Hodkinson, B. P., N. R. Gottel, C. W. Schadt, and F. Lutzoni. In press. Photoautotrophic symbiont and geography are major factors affecting highly structured and diverse bacterial communities in the lichen microbiome. Environmental Microbiology.
----------------------------------------------
This work was funded in part by NSF DEB-1011504 and EF-0832858.
When processing large numbers of DNA sequences, it always helps to have a standardized naming system so that the sequences can be handled in an automated way during downstream analyses. For a recent large-scale cloning experiment that involved picking 2880 clonal bacterial colonies (to amplify and sequence a vector-inserted 16S gene fragment from each), I developed a 10-digit alpha-numeric code that allowed me to encode all of the necessary data about my sequences into each specific sequence identifier. However, the sequencing facility also needed to use its own codes to keep track of my sequences, so I ended up with long names that had my own codes in the middle with information identifying them as my sequences in front and information about the individual sequence reads themselves tacked on at the end. Therefore, to recover the names (without retaining the sequencing facility's additions) in an automated way, I wrote a simple Perl script to edit a fasta file containing the sequences (this was run after the process of manual sequence correction had been finished).
The following script (‘Clon_16S_fasta_renamer.pl’) allowed me to extract the 10-digit alpha-numeric codes that I used in my dissertation studies (Hodkinson 2011) from the long names (with extraneous information) that come from the sequencing facility. It creates a new fasta file with these modified identifiers. Specifically, it takes sequences that have "BH_" (my initials), followed by a 10-digit code, followed by additional characters, and simply renames each sequence using just the 10-digit code (effectively stripping out "BH_" at the beginning and and extra characters at the end). The new file will have the same name, but the extension will be replaced by ".ed.fasta". This can be easily modified for any set of sequences that are identified using a standardized naming scheme.
#!/usr/bin/perl
print "\nPlease type the name of your input file: ";
my $filename = <STDIN>;
chomp $filename;
open (FASTA, $filename);
{
if ($filename =~ /(.*)\.[^.]*/)
{
open OUT, ">$1.ed.fasta";
}
}
while ( <FASTA>)
{
if ($_ =~ /^>BH\_(..........)/)
{
print OUT ">$1\n";
}
if ($_ =~ /^[A,C,G,T,R,Y,K,M,S,W,B,D,H,V,N,:,-][A,C,G,T,R,Y,K,M,S,W,B,D,H,V,N,:,-][A,C,G,T,R,Y,K,M,S,W,B,D,H,V,N,:,-][A,C,G,T,R,Y,K,M,S,W,B,D,H,V,N,:,-]*/)
{
print OUT $_;
}
if ($_ =~ /^[A,C,G,T,R,Y,K,M,S,W,B,D,H,V,N,:,-][A,C,G,T,R,Y,K,M,S,W,B,D,H,V,N,:,-]$/)
{
print OUT $_;
}
if ($_ =~ /^[A,C,G,T,R,Y,K,M,S,W,B,D,H,V,N,:,-]$/)
{
print OUT $_;
}
}
If you need to know how to run a Perl script, you can look it up on Google, but here is one example of how to run a Perl script using Windows (it's actually easier on almost any other type of operating system). Since I was performing a simple task with a very specific data set, it was easy for me to use basic Perl commands. However, for more complex sequence manipulations, BioPerl provides an excellent collection of Perl modules for biological applications.
- Brendan
----------------------------------------------
References
The above script is published in the following sources:
Hodkinson, B. P. 2011. A Phylogenetic, Ecological, and Functional Characterization of Non-Photoautotrophic Bacteria in the Lichen Microbiome. Doctoral Dissertation, Duke University, Durham, NC.
Download Dissertation (PDF file)
Hodkinson, B. P., N. R. Gottel, C. W. Schadt, and F. Lutzoni. 2011. Data from: Photoautotrophic symbiont and geography are major factors affecting highly structured and diverse bacterial communities in the lichen microbiome. Dryad Digital Repository doi:10.5061/dryad.t99b1.
Hodkinson, B. P., N. R. Gottel, C. W. Schadt, and F. Lutzoni. In press. Photoautotrophic symbiont and geography are major factors affecting highly structured and diverse bacterial communities in the lichen microbiome. Environmental Microbiology.
----------------------------------------------
This work was funded in part by NSF DEB-1011504 and EF-0832858.
Saturday, June 18, 2011
Molecular Phylogenetics Workshop
Next week I will be running a short Molecular Phylogenetics Workshop at Roan Mountain State Park in Tennessee (June 22, 10:30-2:00). The workshop coincides with the meeting of the American Bryological and Lichenological Society, but I will be presenting general principles of molecular evolution and phylogenetic inference that are applicable to any set of organisms.
Here is the abstract:
"Cryptogams are notorious for their paucity of morphological characters when compared with higher plants and animals. As a result, an understanding of molecular data and what they can reveal in terms of evolution is perhaps more crucial in these organisms than in many others. Workshop participants will explore principles of molecular phylogenetics and learn basic protocols for running phylogenetic analyses. The main objectives will be (1) to promote an understanding of how events in the course of molecular sequence evolution affect phylogenetic inference, (2) to explore the advantages and disadvantages of different phylogenetic methods, and (3) to facilitate sound research into the phylogenetic history of life. The workshop will include both lecture and discussion. Participants are invited to bring their own data sets for more detailed evaluation at the conclusion of the workshop."
For those scheduled to attend, I look forward to seeing you there! For those not attending, I hope to see you at a future workshop!
- Brendan
----------------------------------------------
This work was was made possible in part by NSF (DEB-1011504) and the American Bryological and Lichenological Society.
Here is the abstract:
"Cryptogams are notorious for their paucity of morphological characters when compared with higher plants and animals. As a result, an understanding of molecular data and what they can reveal in terms of evolution is perhaps more crucial in these organisms than in many others. Workshop participants will explore principles of molecular phylogenetics and learn basic protocols for running phylogenetic analyses. The main objectives will be (1) to promote an understanding of how events in the course of molecular sequence evolution affect phylogenetic inference, (2) to explore the advantages and disadvantages of different phylogenetic methods, and (3) to facilitate sound research into the phylogenetic history of life. The workshop will include both lecture and discussion. Participants are invited to bring their own data sets for more detailed evaluation at the conclusion of the workshop."
For those scheduled to attend, I look forward to seeing you there! For those not attending, I hope to see you at a future workshop!
- Brendan
----------------------------------------------
This work was was made possible in part by NSF (DEB-1011504) and the American Bryological and Lichenological Society.
Friday, June 17, 2011
Writing a Phycas Script
For a while I have been wary of phylogenetic results supported only by Bayesian analyses, because of the so-called 'star-tree paradox' that haunts MrBayes and even some other programs like it. As I have mentioned previously, one of the best features of the Bayesian phylogenetic program Phycas is that it gives one the opportunity to allow polytomies in the trees sampled as part of the posterior (which can often deflate the inflated posterior probability values seen with programs like MrBayes). The specific command for this is:
"mcmc.allow_polytomies = True"
To run Phycas, it is best to write a script to go with a standard NEXUS-formatted sequence alignment. There is some basic information on how to install and run Phycas in my previous post:
http://squamules.blogspot.com/2011/06/installing-and-running-phycas.html
However, that post does not go into any of the details of scripting for Phycas. Recently, I ran a multigene analysis with mtSSU, ITS1, 5.8S, and ITS2 in different partitions, with a different evolutionary model for each. Here is what my Phycas script looked like:
from phycas import *
setMasterSeed(98765)
mcmc.data_source = 'Input_file_name.nex'
mcmc.out.log = 'Output_file_name.log'
mcmc.out.log.mode = REPLACE
mcmc.allow_polytomies = True
mcmc.polytomy_prior = False
mcmc.topo_prior_C = 1.0
mcmc.out.trees.prefix = 'Output_file_name'
mcmc.out.params.prefix = 'Output_file_name'
mcmc.ncycles = 50000
mcmc.sample_every = 10
# Set up the K80+I model for 5pt8S
model.type="hky"
model.state_freqs = [0.25, 0.25, 0.25, 0.25]
model.fix_freqs = True
model.kappa = 2.0
model.kappa_prior = BetaPrime(1.0, 1.0)
model.pinvar_model = True
# Save the K80+I model for 5pt8S
m3 = model()
# Set up the GTR+I model for mtSSU
model.type="gtr"
model.state_freqs = [0.3338, 0.1493, 0.1983, 0.3187]
model.fix_freqs = False
model.relrates = [1.4783, 5.8050, 3.3222, 0.6768, 7.6674, 1.0000]
model.pinvar_model = True
# Save the GTR+I model for mtSSU
m1 = model()
# Set up the HKY+G model for ITS1
model.type="hky"
model.state_freqs = [0.1487, 0.3566, 0.2704, 0.2244]
model.fix_freqs = False
model.kappa = 2.0
model.kappa_prior = BetaPrime(1.0, 1.0)
model.num_rates = 4
model.gamma_shape = 0.5
model.gamma_shape_prior = Exponential(1.0)
model.pinvar_model = False
# Save the HKY+G model for ITS1
m2 = model()
# Set up the HKY+G model for ITS2
model.state_freqs = [0.1419, 0.3069, 0.3199, 0.2314]
# Save the HKY+G model for ITS2
m4 = model()
# Define partition subsets
mtssu = subset(1, 1080)
its1 = subset(1081, 1607)
fivept8S = subset(1608, 1768)
its2 = subset(1769, 2041)
# Assign partition models to subsets
partition.addSubset(mtssu, m1, "mtSSU")
partition.addSubset(its1, m2, "ITS1")
partition.addSubset(fivept8S, m3, "5pt8S")
partition.addSubset(its2, m4, "ITS2")
partition()
# Start the run
mcmc()
# Summarize the posterior
sumt.trees = 'trees.t'
sumt.burnin = 500
sumt.tree_credible_prob = 1.0
sumt()
Although I have some notes within the script, please see the Phycas manual for instructions on what each of the individual commands does. Hopefully more people will be using Phycas (and allowing polytomies!) in the future!
- Brendan
Note: One question that I had about running Phycas was how to define exclusion sets; however, Phycas apparently can read the EXSET line of the ASSUMPTIONS block of the NEXUS file in the same way that Mesquite, MacClade, and PAUP* can.
"mcmc.allow_polytomies = True"
To run Phycas, it is best to write a script to go with a standard NEXUS-formatted sequence alignment. There is some basic information on how to install and run Phycas in my previous post:
http://squamules.blogspot.com/2011/06/installing-and-running-phycas.html
However, that post does not go into any of the details of scripting for Phycas. Recently, I ran a multigene analysis with mtSSU, ITS1, 5.8S, and ITS2 in different partitions, with a different evolutionary model for each. Here is what my Phycas script looked like:
from phycas import *
setMasterSeed(98765)
mcmc.data_source = 'Input_file_name.nex'
mcmc.out.log = 'Output_file_name.log'
mcmc.out.log.mode = REPLACE
mcmc.allow_polytomies = True
mcmc.polytomy_prior = False
mcmc.topo_prior_C = 1.0
mcmc.out.trees.prefix = 'Output_file_name'
mcmc.out.params.prefix = 'Output_file_name'
mcmc.ncycles = 50000
mcmc.sample_every = 10
# Set up the K80+I model for 5pt8S
model.type="hky"
model.state_freqs = [0.25, 0.25, 0.25, 0.25]
model.fix_freqs = True
model.kappa = 2.0
model.kappa_prior = BetaPrime(1.0, 1.0)
model.pinvar_model = True
# Save the K80+I model for 5pt8S
m3 = model()
# Set up the GTR+I model for mtSSU
model.type="gtr"
model.state_freqs = [0.3338, 0.1493, 0.1983, 0.3187]
model.fix_freqs = False
model.relrates = [1.4783, 5.8050, 3.3222, 0.6768, 7.6674, 1.0000]
model.pinvar_model = True
# Save the GTR+I model for mtSSU
m1 = model()
# Set up the HKY+G model for ITS1
model.type="hky"
model.state_freqs = [0.1487, 0.3566, 0.2704, 0.2244]
model.fix_freqs = False
model.kappa = 2.0
model.kappa_prior = BetaPrime(1.0, 1.0)
model.num_rates = 4
model.gamma_shape = 0.5
model.gamma_shape_prior = Exponential(1.0)
model.pinvar_model = False
# Save the HKY+G model for ITS1
m2 = model()
# Set up the HKY+G model for ITS2
model.state_freqs = [0.1419, 0.3069, 0.3199, 0.2314]
# Save the HKY+G model for ITS2
m4 = model()
# Define partition subsets
mtssu = subset(1, 1080)
its1 = subset(1081, 1607)
fivept8S = subset(1608, 1768)
its2 = subset(1769, 2041)
# Assign partition models to subsets
partition.addSubset(mtssu, m1, "mtSSU")
partition.addSubset(its1, m2, "ITS1")
partition.addSubset(fivept8S, m3, "5pt8S")
partition.addSubset(its2, m4, "ITS2")
partition()
# Start the run
mcmc()
# Summarize the posterior
sumt.trees = 'trees.t'
sumt.burnin = 500
sumt.tree_credible_prob = 1.0
sumt()
Although I have some notes within the script, please see the Phycas manual for instructions on what each of the individual commands does. Hopefully more people will be using Phycas (and allowing polytomies!) in the future!
- Brendan
Note: One question that I had about running Phycas was how to define exclusion sets; however, Phycas apparently can read the EXSET line of the ASSUMPTIONS block of the NEXUS file in the same way that Mesquite, MacClade, and PAUP* can.
Tuesday, June 14, 2011
Installing and Running Phycas
Phycas has recently earned a high spot on my short list of favorite computer programs for phylogenetics. Phycas is the amazing program that can run a Bayesian phylogenetic inference without being susceptible to the 'star-tree paradox' because it allows for the existence of polytomies in the sampled trees.
From an academic perspective, Phycas is actually a pretty easy program to run and install. Still, some additional notes on tricks and tips for running it were beneficial to one of my colleagues who was really having trouble getting it to go. Here were my instructions for installing Phycas on a Windows machine:
1) Install Python 2.7. [I use the Enthought Python Distribution, available here: http://www.enthought.com/products/epd.php. Everything is bundled together so components like SciPy, NumPy, etc., never need to be installed individually and the different versions of the components are all guaranteed to play well together.]
2) Follow the instructions here:
http://hydrodictyon.eeb.uconn. edu/projects/phycas/index.php/ Telling_Windows_where_to_find_ Python
to append Python27 (different from the versions they have listed there) to your PATH (I guess if your PATH is truly empty then you will just leave out the semi-colon; otherwise, keep whatever's already in your PATH in there, but just add ;C:\Python27 to the end of it it). [It might also be important to make sure that the PYTHON-STARTUP environmental variable says C:\Python27 (if you have that variable... mine was still set to 2.6, meaning that the wrong version of Python would likely open up by default), and that this is all being done for the system level and the user level... I was only doing it for the user level for a while and it got me mixed up.]
3) Do the 4-step Phycas installation as outlined on the "Windows XP/Windows Vista/Windows 7" section of this website (the manual itself is apparently wrong, so be careful here).
4) For your own particular analysis, put the NEXUS-formatted alignment file ('.nex') and the phycas script file ('.py') on the Desktop where you have the shortcut to the '.bat' file. [For more on writing a Phycas script, stay turned to this blog!]
5) Drag and drop the phycas script file (.py) onto 'Shortcut to phycas.bat'.
I'll have another blog post that goes more into the details of Phycas scripting, but I hope this post helps jump-start some of those eager to deflate their inflated posterior probability values!
- Brendan
From an academic perspective, Phycas is actually a pretty easy program to run and install. Still, some additional notes on tricks and tips for running it were beneficial to one of my colleagues who was really having trouble getting it to go. Here were my instructions for installing Phycas on a Windows machine:
1) Install Python 2.7. [I use the Enthought Python Distribution, available here: http://www.enthought.com/products/epd.php. Everything is bundled together so components like SciPy, NumPy, etc., never need to be installed individually and the different versions of the components are all guaranteed to play well together.]
2) Follow the instructions here:
http://hydrodictyon.eeb.uconn.
to append Python27 (different from the versions they have listed there) to your PATH (I guess if your PATH is truly empty then you will just leave out the semi-colon; otherwise, keep whatever's already in your PATH in there, but just add ;C:\Python27 to the end of it it). [It might also be important to make sure that the PYTHON-STARTUP environmental variable says C:\Python27 (if you have that variable... mine was still set to 2.6, meaning that the wrong version of Python would likely open up by default), and that this is all being done for the system level and the user level... I was only doing it for the user level for a while and it got me mixed up.]
3) Do the 4-step Phycas installation as outlined on the "Windows XP/Windows Vista/Windows 7" section of this website (the manual itself is apparently wrong, so be careful here).
4) For your own particular analysis, put the NEXUS-formatted alignment file ('.nex') and the phycas script file ('.py') on the Desktop where you have the shortcut to the '.bat' file. [For more on writing a Phycas script, stay turned to this blog!]
5) Drag and drop the phycas script file (.py) onto 'Shortcut to phycas.bat'.
I'll have another blog post that goes more into the details of Phycas scripting, but I hope this post helps jump-start some of those eager to deflate their inflated posterior probability values!
- Brendan
Monday, June 13, 2011
A New Home
I have relocated and am finally settled in New York! I will be doing my post-doctoral work at the New York Botanical Garden (NYBG) using advanced bioinformatics tools to study lichen ecology and evolution with Richard C. Harris and James C. Lendemer. Here is my new address:
International Plant Science Center
The New York Botanical Garden
2900 Southern Blvd.
Bronx, NY 10458-5126
NYBG has an amazing research program that is not typical for a botanical garden. The departments within the International Plant Science Center include the following: Institute of Systematic Botany, Cullman Program for Molecular Systematics, Plant Research Laboratory, NY Plant Genomics Consortium, Steere Herbarium, Graduate Studies, Mertz Library, and NYBG Press (plus more). My research will likely connect in some way with all of these departments, which is why NYBG is a perfect environment for my post-doctoral research.
Those who have been following my research closely might say "obviously you study diverse organisms from across the tree of life, but since when do you study plants?" Since the concept of "plants" once included fungi, the "plant" research program at NYBG (which began taking shape over 100 years ago) provides amazing resources for the study of fungal biology as well. Of course, I will specifically be focusing my energy on the lichen-forming fungi.
Please stay tuned for more on my research as it unfolds in New York City!
- Brendan
International Plant Science Center
The New York Botanical Garden
2900 Southern Blvd.
Bronx, NY 10458-5126
NYBG has an amazing research program that is not typical for a botanical garden. The departments within the International Plant Science Center include the following: Institute of Systematic Botany, Cullman Program for Molecular Systematics, Plant Research Laboratory, NY Plant Genomics Consortium, Steere Herbarium, Graduate Studies, Mertz Library, and NYBG Press (plus more). My research will likely connect in some way with all of these departments, which is why NYBG is a perfect environment for my post-doctoral research.
Those who have been following my research closely might say "obviously you study diverse organisms from across the tree of life, but since when do you study plants?" Since the concept of "plants" once included fungi, the "plant" research program at NYBG (which began taking shape over 100 years ago) provides amazing resources for the study of fungal biology as well. Of course, I will specifically be focusing my energy on the lichen-forming fungi.
Please stay tuned for more on my research as it unfolds in New York City!
- Brendan
Monday, May 23, 2011
Graduation/Dissertation
Upon my graduation, I thought I'd give a little update. The thesis defense last month was successful and the final draft of my dissertation has been submitted! My thesis was on the communities of non-photoautotrophic bacteria associated with lichens. This research broke a lot of new ground in terms of using high-throughput pyrosequencing and cutting-edge bioinformatics to examine the phylogenetic, ecological, and functional complexity of the lichen microbiome. My hope is that my dissertation will contribute to science through both the data that I have generated and the set of tools that I have developed. I have talked a little about some of the bioinformatics tools that I have developed in previous posts here, but I will continue to post additional elements of my thesis, especially as they are published in peer-reviewed journals.
The graduation itself gave me a final chance to reflect on the great opportunities that have been made available to me at Duke University. Here you can see me chillin' out after the hooding ceremony:
This week I'm in the process of moving up to NY to start my postdoctoral research on the systematics of lichen-forming fungi at the New York Botanical Garden. I will be working with Dr. Richard C. Harris and James Lendemer, focusing on long-standing problems in Eastern North American lichen taxonomy, while testing a number of specific ecological and biogeographical hypotheses. Addressing many of these issues will require cutting-edge bioinformatics tools that I have developed or am in the process of developing with a network of collaborators in diverse fields.
- Brendan
The graduation itself gave me a final chance to reflect on the great opportunities that have been made available to me at Duke University. Here you can see me chillin' out after the hooding ceremony:
This week I'm in the process of moving up to NY to start my postdoctoral research on the systematics of lichen-forming fungi at the New York Botanical Garden. I will be working with Dr. Richard C. Harris and James Lendemer, focusing on long-standing problems in Eastern North American lichen taxonomy, while testing a number of specific ecological and biogeographical hypotheses. Addressing many of these issues will require cutting-edge bioinformatics tools that I have developed or am in the process of developing with a network of collaborators in diverse fields.
- Brendan
Monday, May 16, 2011
Investing in Science
Several months back I worked with a group of other student members of the Botanical Society of America (BSA) to draft an open letter to lawmakers to express our hope that policymakers in Washington, DC, would sustain a national commitment to invest in our nation's scientific research, development, and education. This has now evolved into a petition. Please read the email below from the BSA student representatives to find out more about this effort:
"
Attention Students: Ask Lawmakers to Support Science Education and Research
First, we'd like to thank all of you who have taken action and responded to this call. Please take the extra step and ask your friends to consider doing so as well.
The end of the academic year is quickly approaching. Before you venture off for the summer research season or for a short break from classes, you have one more assignment. I am writing to ask that you join with other science students from across the country to sign an online petition to lawmakers. This statement reminds our elected leaders that scientific research and education are keys to our future and asks them to continue to make important investments in the scientific programs that will support your education and preparation for future careers in research, teaching, or the myriad fields that grow and benefit from scientific research.
We already have more than 2,750 signatures, but we would like to have more than 5,000 by the end of May. So, if you have not already signed this online petition, please do so today at http://www.aibs.org/public- policy/science_students_ letter.html. You may also sign the petition and encourage your friends to sign via Facebook -- http://www.facebook.com/pages/ Students-Sign-the-Open-Letter- to-Policymakers-About- Investments-in-Science/ 183684855001704.
Thank you for your time and support!
Sincerely,
Botanical Society of America Student Representatives, Marian Chau (University of Hawai`i at Manoa) and Rachel Meyer (New York Botanical Garden)
"
"
Attention Students: Ask Lawmakers to Support Science Education and Research
First, we'd like to thank all of you who have taken action and responded to this call. Please take the extra step and ask your friends to consider doing so as well.
The end of the academic year is quickly approaching. Before you venture off for the summer research season or for a short break from classes, you have one more assignment. I am writing to ask that you join with other science students from across the country to sign an online petition to lawmakers. This statement reminds our elected leaders that scientific research and education are keys to our future and asks them to continue to make important investments in the scientific programs that will support your education and preparation for future careers in research, teaching, or the myriad fields that grow and benefit from scientific research.
We already have more than 2,750 signatures, but we would like to have more than 5,000 by the end of May. So, if you have not already signed this online petition, please do so today at http://www.aibs.org/public-
Thank you for your time and support!
Sincerely,
Botanical Society of America Student Representatives, Marian Chau (University of Hawai`i at Manoa) and Rachel Meyer (New York Botanical Garden)
"
Friday, May 13, 2011
Punctelia eganii Hodkinson & Lendemer, sp. nov., a rare chemical oddity
Just today I had an article published in the journal Opuscula Philolichenum in which I, along with James Lendemer, describe a new species in the genus Punctelia that has only been found once and contains a chemical compound (lichexanthone) that has otherwise not been seen in this genus. The species was collected along the Alabama River in historic Monroe County, Alabama, near Monroeville (childhood home of Harper Lee and Truman Capote; "the literary capital of Alabama").
Under an ultraviolet light, it makes pin-pricks of bright light across the surface (due to the localized presence of lichexanthone). The species is named Punctelia eganii, after Dr. Bob Egan of the University of Nebraska at Omaha, who first collected the species and brought it to our attention. The paper ends with a discussion of 'chemotaxonomy' and the evolving views on the role of secondary chemistry in lichen taxonomy.
- Brendan
--------------------------------------------------------------
Reference:
Hodkinson, B. P., and J. C. Lendemer. 2011. Punctelia eganii, a new species in the P. rudecta group with a novel secondary compound for the genus. Opuscula Philolichenum 9: 35-38.
Download publication (PDF file)
Under an ultraviolet light, it makes pin-pricks of bright light across the surface (due to the localized presence of lichexanthone). The species is named Punctelia eganii, after Dr. Bob Egan of the University of Nebraska at Omaha, who first collected the species and brought it to our attention. The paper ends with a discussion of 'chemotaxonomy' and the evolving views on the role of secondary chemistry in lichen taxonomy.
- Brendan
--------------------------------------------------------------
Reference:
Hodkinson, B. P., and J. C. Lendemer. 2011. Punctelia eganii, a new species in the P. rudecta group with a novel secondary compound for the genus. Opuscula Philolichenum 9: 35-38.
Download publication (PDF file)
Tuesday, May 10, 2011
PICS-Ord Tricks
Those who have ever used the R statistical package know that it can be a bit tricky, especially if you're first starting out.
The R-based PICS-Ord program (developed to recode ambiguously-aligned regions for phylogenetic analyses; see here and here and here for more information) was written to be as simple as possible while remaining flexible and adjustable. A set of pretty comprehensive instructions was produced to facilitate analyses and provide recommendations for basic use (see 'manual.pdf' available here). The manual is where everyone should look first for help with PICS-Ord. However, those who are not experts in R and/or the command line may benefit from some additional information on how to implement PICS-Ord without really having to know much background information on the R statistical package. Therefore, the goal of this post is to detail one relatively simple way of implementing the PICS-Ord program on a PC.
Here are instructions for one way of running PICS-Ord on Microsoft Windows. Note that these instructions will only work with installations of the Windows version of R (try 2.12.0; the most recent version of R did not work at the time that this was last updated) [ http://cran.r-project.org/bin/ windows/base/old/2.12.0/ ] and the Ngila Windows executable [ http://scit.us/projects/ngila/ ] (for the latter, choose the option of putting Ngila in your PATH upon installation).
1) Place picsord.R (found in the picsord.zip archive) in the same folder as ngila.exe (probably C:\\"Program Files"\Ngila\bin\).
[Note: If Ngila is not in your path, you can go into the picsord.R file and change "ngila" (the one in quotes) to "ngila.exe" in the first line that is not preceded by a hash mark (#), or you can type out the full path (but this will not be necessary if you are following the rest of this procedure). Some users have manually edited picsord.R by deleting the first line (the line specifying the location of Rscript); this seems unnecessary but may be worthwhile on certain machines.]
2) Save/copy the input fasta file to a directory (e.g., your home directory, Desktop, or C:\).
3) In the Command Prompt window, use the 'cd' or 'chdir' command to navigate to the directory in which the input fasta file is stored (or simply put the input fasta file in the home directory so that navigation is not necessary), then type the following (this line can be modified based on the version of R being used or the specific location on the drive where Rscript and picsord.R are located):
C:\\"Program Files"\R\R-2.12.0\bin\Rscript.exe C:\\"Program Files"\Ngila\bin\picsord.R input.fas > output.phy
Multiple regions can be processed by running the command in step 3 separately for each region, or one can use the .bat file that comes as part of the PICS-Ord package (use of the .bat file is outlined in the manual). After this, the phylip-formatted PICS-Ord alignment portions can be pasted at the end of the original nucleotide alignment alongside the unambiguously-aligned sites. Please see the manual for further recommendations regarding implementation.
-Brendan
References:
Lücking, R., B. P. Hodkinson, A. Stamatakis, and R. A. Cartwright. 2011. PICS-Ord: Unlimited Coding of Ambiguous Regions by Pairwise Identity and Cost Scores Ordination. BMC Bioinformatics 12: 10.
Download publication (PDF file)
Download R-based PICS-Ord program (zipped program package)
View program wiki (website)
The R-based PICS-Ord program (developed to recode ambiguously-aligned regions for phylogenetic analyses; see here and here and here for more information) was written to be as simple as possible while remaining flexible and adjustable. A set of pretty comprehensive instructions was produced to facilitate analyses and provide recommendations for basic use (see 'manual.pdf' available here). The manual is where everyone should look first for help with PICS-Ord. However, those who are not experts in R and/or the command line may benefit from some additional information on how to implement PICS-Ord without really having to know much background information on the R statistical package. Therefore, the goal of this post is to detail one relatively simple way of implementing the PICS-Ord program on a PC.
Here are instructions for one way of running PICS-Ord on Microsoft Windows. Note that these instructions will only work with installations of the Windows version of R (try 2.12.0; the most recent version of R did not work at the time that this was last updated) [ http://cran.r-project.org/bin/
1) Place picsord.R (found in the picsord.zip archive) in the same folder as ngila.exe (probably C:\\"Program Files"\Ngila\bin\).
[Note: If Ngila is not in your path, you can go into the picsord.R file and change "ngila" (the one in quotes) to "ngila.exe" in the first line that is not preceded by a hash mark (#), or you can type out the full path (but this will not be necessary if you are following the rest of this procedure). Some users have manually edited picsord.R by deleting the first line (the line specifying the location of Rscript); this seems unnecessary but may be worthwhile on certain machines.]
2) Save/copy the input fasta file to a directory (e.g., your home directory, Desktop, or C:\).
3) In the Command Prompt window, use the 'cd' or 'chdir' command to navigate to the directory in which the input fasta file is stored (or simply put the input fasta file in the home directory so that navigation is not necessary), then type the following (this line can be modified based on the version of R being used or the specific location on the drive where Rscript and picsord.R are located):
C:\\"Program Files"\R\R-2.12.0\bin\Rscript.exe C:\\"Program Files"\Ngila\bin\picsord.R input.fas > output.phy
After this, the output phylip file should appear in the working directory.
Multiple regions can be processed by running the command in step 3 separately for each region, or one can use the .bat file that comes as part of the PICS-Ord package (use of the .bat file is outlined in the manual). After this, the phylip-formatted PICS-Ord alignment portions can be pasted at the end of the original nucleotide alignment alongside the unambiguously-aligned sites. Please see the manual for further recommendations regarding implementation.
-Brendan
References:
Lücking, R., B. P. Hodkinson, A. Stamatakis, and R. A. Cartwright. 2011. PICS-Ord: Unlimited Coding of Ambiguous Regions by Pairwise Identity and Cost Scores Ordination. BMC Bioinformatics 12: 10.
Download publication (PDF file)
Download R-based PICS-Ord program (zipped program package)
View program wiki (website)
Subscribe to:
Posts (Atom)