☆ 4.6 Article

Distance-Based Phylogenetic Placement with Statistical Support

BIOLOGY-BASEL (2022)

Journal

BIOLOGY-BASEL

Volume 11, Issue 8, Pages -

Publisher

MDPI

DOI: 10.3390/biology11081212

Keywords

phylogenetic placement; statistical support; distance-based phylogenetic inference; bootstrapping

Funding

National Institute of Health [1R35GM142725]

Ask authors/readers for more resources

Protocol

Community support

Reagent

Community support

Automated Summary New
Abstract

Phylogenetic identification of unknown sequences through tree placement is commonly used in ecological studies. This article addresses the issue of uncertainty in placements obtained from incomplete and noisy data. Nonparametric bootstrapping is found to be the most accurate method for measuring support, and an efficient linear algebraic formulation for bootstrapping is presented. The article also compares the accuracy of maximum likelihood support values and distance-based methods in different applications and datasets.

Phylogenetic identification of unknown sequences by placing them on a tree is routinely attempted in modern ecological studies. Such placements are often obtained from incomplete and noisy data, making it essential to augment the results with some notion of uncertainty. While the standard likelihood-based methods designed for placement naturally provide such measures of uncertainty, the newer and more scalable distance-based methods lack this crucial feature. Here, we adopt several parametric and nonparametric sampling methods for measuring the support of phylogenetic placements that have been obtained with the use of distances. Comparing the alternative strategies, we conclude that nonparametric bootstrapping is more accurate than the alternatives. We go on to show how bootstrapping can be performed efficiently using a linear algebraic formulation that makes it up to 30 times faster and implement this optimized version as part of the distance-based placement software APPLES. By examining a wide range of applications, we show that the relative accuracy of maximum likelihood (ML) support values as compared to distance-based methods depends on the application and the dataset. ML is advantageous for fragmentary queries, while distance-based support values are more accurate for full-length and multi-gene datasets. With the quantification of uncertainty, our work fills a crucial gap that prevents the broader adoption of distance-based placement tools.

Distance-Based Phylogenetic Placement with Statistical Support

Journal

BIOLOGY-BASEL

Publisher

MDPI

Keywords

Categories

Funding

Ask authors/readers for more resources

Protocol

Reagent

Authors

I am an author on this paper

Reviews

Primary Rating

Secondary Ratings

Novelty

Significance

Scientific rigor

Rate this paper

Recommended

Distance-Based Phylogenetic Placement with Statistical Support

Journal

BIOLOGY-BASEL

Publisher

MDPI

Keywords

Categories

Funding

Ask authors/readers for more resources

Protocol

Reagent

Authors

I am an author on this paper

Reviews

Primary Rating

Secondary Ratings

Novelty

Significance

Scientific rigor

Rate this paper

Recommended

Export Citation

Share Paper