Statistical Inconsistency of Error-correction Objectives for Perfect Phylogenies
Statistical Inconsistency of Error-correction Objectives for Perfect Phylogenies
Satas, G.; Myers, M. A.; Shah, S. P.
AbstractThe binary perfect phylogeny, in which each mutation arises exactly once on an evolutionary tree and is never lost, is a well-studied idealized phylogenetic model. When observed data has errors, a common approach to phylogeny inference is to seek a tree that minimizes the number of error corrections ("flips'') needed to fit a perfect phylogeny. These objectives draw on an intuitive justification: minimizing implied errors should prefer the true tree in expectation. We test this assumption using a generative model with independent errors and prove that error-correction objectives are statistically inconsistent for all positive error rates: in expectation, the minimum-cost tree need not be the true tree. Our proof is constructive and yields counterexamples involving any tree topology and any positive error rates, demonstrating the ubiquity of the phenomenon. The core problem is that tree topologies can explain observed mutation patterns with fewer errors than actually occurred, and differ in their ability to do so, introducing systematic bias. This mechanism is distinct from previously identified sources of inconsistency such as homoplasy or incomplete lineage sorting, since the error-free setting is trivially consistent for perfect phylogenies. We investigate how often this failure may occur in practice. Simulations calibrated to error rates from single-cell sequencing data show that an incorrect tree is preferred over the true tree in a substantial fraction of cases (over 50% in some settings) with rates increasing with tree size. Moreover, winning trees are not random but share specific topological features. Notably, at error rates typical of single-cell sequencing data, trees with deeper, more imbalanced topologies are consistently favored over more balanced ones. These results demonstrate that inconsistency is not a theoretical edge case, and that understanding when and how it arises is important when interpreting results in practice.