MétaCan
Menu
Back to cohort
Record W2086293112 · doi:10.1093/sysbio/syu070

The Unsolved Challenge to Phylogenetic Correlation Tests for Categorical Characters

2014· article· en· W2086293112 on OpenAlexaffabout
Wayne P. Maddison, Richard G. FitzJohn

Bibliographic record

VenueSystematic Biology · 2014
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicGenetic diversity and population structure
Canadian institutionsUniversity of British Columbia
Fundersnot available
KeywordsBoulevardColumbia universityBiologyArchaeologyBiodiversityCategorical variableGeographyEcologyMedia studiesSociology

Abstract

fetched live from OpenAlex

When comparative biologists observe that animal species living in caves also tend to have reduced eyes, they may see such correlation as evidence that the traits are adaptively or functionally linked: for instance, selection to maintain eye function is relaxed when light is unavailable. Such cross-species correlations cannot give definitive tests of evolutionary mechanisms, but nonetheless offer important insights into biological relationships among traits in realms as diverse as ecology (e.g., Paradis et al. 1998; Purvis et al. 2000) and genomics (e.g., von Mering et al. 2002; Barker and Pagel 2005). However, the last few decades have taught us that among-species correlative tests should take into account evolutionary relationships (Felsenstein 1985; Ridley 1989; Harvey and Pagel 1991). If phylogeny is not taken into account, an interpreted correlation may have a trivial explanation different from the biological relationship we claim. There is a correlation among species in the distribution of fur and bones in the middle ear—species with fur also have three bones in the middle ear, and vice versa. These two traits are characteristics of mammals, and absent outside the mammals. Using their shared distribution as evidence of an interesting biological relationship between fur and middle ear bones would be considered a mistake, however, for reasons understood long ago by Darwin (1872): We may often falsely attribute to correlated variation structures which are common to whole groups of species, and which in truth are simply due to inheritance; for an ancient progenitor may have acquired through natural selection some one modification in structure, and, after thousands of generations, some other and independent modification; and these two modifications, having been transmitted to a whole group of descendants with diverse habits, would naturally be thought to be in some necessary manner correlated. Not all is healthy with the paradigm, however. For categorical characters a special concern has been raised: commonly used and well-respected methods (e.g., Pagel 1994), do not eliminate pseudoreplication, as they are susceptible to an effect from a single evolutionary event (Maddison 1990, 2000; Read and Nee 1995; Ridley and Grafen 1996; Grafen and Ridley 1997). As a result, a significant statistical association between traits inferred by these methods can mean very little in some circumstances, misleading biologists about the relationship between traits. Although this concern has been raised, there has been little effect on the practice of comparative studies, we think because the issue has not been well understood. We do not, alas, have a solution. Our purpose here is to explore the issue in depth in part to bring caution to comparative studies, in part to characterize how our methods should behave in hopes of provoking appropriate solutions. It is useful to begin by outlining the limits of what we can hope to learn from comparative data. Comparative data may indicate a correlation between variable X and variable Y⁠, but this could be caused by the influence of a third variable—a streamlined body may not lead to tolerance to anoxia among mammals, but an aquatic habitat could lead to both. With comparative data alone, we cannot rule out the existence of a third variable that influences the two variables of interest. Credible mechanisms and controlled, manipulative experiments are needed to support precise causal hypotheses (Westoby 1999). This means that any comparative method, no matter how robustly applied, must be satisfied with a relatively weak conclusion: that the two variables of interest appear to be part of the same adaptive/functional network, causally linked either directly, or indirectly through other variables. When teased out of patterns in cross-species data, even this weaker conclusion is likely interesting and satisfying to biologists. This article is not simply a reminder that “correlation does not imply causation”. Rather, it is to highlight the fact that in cross-species data, a correlation may not even give us the weaker conclusion—it may not even imply that variables are part of the same functional network. The perfect co-distribution of fur and three middle-ear bones among species would not be expected to occur by chance in a universe in which each species were independently evolved, but that is not the universe we live in. Instead, as Darwin explains, such co-distribution can be interpreted as a simple consequence of independent origin of two traits in a lineage, followed by co-inheritance by descendant species. It is not important whether this is described as no correlation (compatible with a null hypothesis involving simple phylogenetic descent) or as a correlation explained by coincidence. In either perspective, the pattern does not provide evidence for an adaptive or functional relationship among the traits. Our target, and that of most comparative biologists, is not merely a general sense of correlation that reaches statistical significance, but rather a special sense of correlation that gives evidence for an adaptive or functional relationship between the traits. If we were satisfied with the former, we might as well do a Fisher's exact test on the species without regard to the phylogeny. If we want the latter, we should design our statistical tests to reveal adaptively or functionally interesting correlations. What we mean by “adaptive/functional relationship” (or “adaptive/functional network”) is a set of influences among traits that arise from how the traits act or behave. Thus, we mean that one trait influences the other, or a third trait influences both, through some genomic, developmental, physiological, ecological, or other effect, whether immediate or through adaptation in evolutionary time. The influences need not involve adaptive function or selection pressure; merely that there is an effect. The effect need not be specified, and indeed a comparative test alone cannot reveal the precise nature of the adaptive/functional relationship. However, a comparative test on its own can provide evidence that there exists an adaptive/functional relationship of some sort, even without manipulative experiments or information about mechanism. Many biologists seek such evidence from comparative data. We want our comparative tests to respond to patterns that would be difficult to explain unless the traits were part of the same adaptive/functional network, and respond only to such patterns. We would like to know therefore whether a positive result from a comparative test justifiably supports the conclusion that the variables are part of the same adaptive/functional network, or whether we are being misled by coincidence or biased sampling. The problem of concern in this paper is that tests such as Pagel's (1994) and Maddison's (1990) in some circumstances appear to give such support, when in fact the comparative data offers no support for a biologically interesting association (Maddison 1990; Read and Nee 1995; Ridley and Grafen 1996; Grafen and Ridley 1997). In order to consider the performance of particular comparative methods, we outline four scenarios to serve as litmus tests (Fig. 1). We argue that two of the scenarios (Fig. 1a,b) provide good evidence for an adaptive/functional relationship among the variables, while the other two (Fig. 1c,d) do not. Thus, we argue, comparative methods should indicate a correlation for Figure 1a,b (or at least, for scenarios of these types with sufficiently many origins), while they should not for Figure 1c,d. Four scenarios for the evolution of states of characters X and Y⁠. In each, the same phylogeny is mirrored to show X at left and Y at right. State 0 = white; State 1 = black. (a) Replicated co-distribution. Replicated (a) and provide good evidence for an interesting adaptive/functional relationship between X and and do not. Figure a in which comparative methods should indicate a pattern of of the in X and in Y⁠, Although the may be few to statistical on the test patterns of this would provide evidence for an It is to that X and Y are adaptively or functionally linked in some even indirectly through a third should the association between X and Y in independent Figure in which an adaptively or functionally interesting correlation can be In each of an origin of the in there are of in Y⁠. This that in Y can be or on the in but that the effect is necessary This pattern of us that statistical methods need to be to the particular biological hypotheses as different might different patterns (Maddison As with it is to that X and Y are linked in some Figure an with X and Y having a to a particular like fur and middle ear bones in mammals. This pattern be a has one the perfect it offers no on its for an adaptive/functional relationship. These traits could have of the long to mammals, each by its own independent and been each by its own in the There is no to think that any two of these traits are correlated to each other any to any one of the thousands of other in the that would have in this of the can and do independently a a single origin of variables no evidence for an interesting it also no evidence one (Westoby 1999). have a functional in The fact that there is only a single origin of among species with is not evidence a functional between the good can be for a functional relationship between and but it from and other data, rather from their co-distribution. Figure a there is in one but not the there is only a single origin of the of but of in Y that same can be that special is to Y in that but this does not imply an interesting correlation with Many other in the likely in the of the and were to with one of could as well be the of There is no good to any or adaptive/functional between X and their correlation could be coincidence. For a from the to the has at in outside mammals, in the et al. we attribute this of in to their we to the evolutionary we should have no that the evolution of would be with an of evolution of for methods for correlated trait evolution is to phylogenetic pseudoreplication, we would hope that they would no significant association between traits in (Fig. and the (Fig. that the association is not It might therefore as a that the most used phylogenetic to test for correlation of categorical Pagel's (1994) to a significant association between traits in the of Pagel (1994) we used (Maddison and to each with species. used to through the and at the of species set the of the outside the to the to This to two characters and with to in Figure such of characters for each of the We used the in to each or of evolution of each on the other In all of the for Pagel's a (Fig. Pagel's (1994) test to like and like light of of correlated independent Pagel's (1994) also correlation in the test we used the same For the X we used the as described in Figure For the Y we used (Maddison and to a on the with this Y states were evolved, its states outside the by X were set to Although this is not an evolutionary it is the same as a in which the Y variable at the with 0 and very and its of in the only is that in that the would have with Y in but in our it might not. Using to Pagel (1994) we that of the in a significant result with (Fig. Thus, Pagel's test is susceptible to significant from the of a single in one of the in and the tests the same which we that test is also susceptible to the of a single the and cannot rule out variables simply with of interest. It also give a significant result with the test is to on a single The correlation test of et al. a simple of in a the of that with states 1 in X and Y⁠, 0 in both, and 0 in one and 1 in the In the tend to show long of that 1 in X and Y⁠, and that have 0 in This is interpreted by the of et al. as a of an The methods of Pagel and et al. there are significant correlations in Figure even there is only a single evolutionary in one or is and correlations can be for interesting It that we as as we thought in for Our to these is to the as merely the they are what they were to Thus, one might that Pagel's test a significant correlation in because its are If a distribution is with the of evolution by a correlation a with the of we could simply the test to the and not like of X and Y in Figure cannot be with Pagel's of evolution this has not been Pagel's test would not be at it would simply be However, even the test could be for the of its that does not a solution. We need to be to when the is and to a test to appropriate Although all are to we that most biologists do not that this and to an of the need for phylogenetic is that these tests are in fact to to we two characters at it is very that their would on the same This is a However, do we characters could variation many and with as long as one for of the many however, is not what is in comparative only a few characters are considered in of some may have been on functional without of their may have been because we a trait a and what effect it on its we may traits that we know a characterize of special interest. If this is the we cannot be to see Thus, our significant may from When we to as evidence for at part of our may from an that is or of characters on which to for that to is in evolutionary correlation tests not be the only that However, even of characters been and or the been in the data, we argue that one should the pattern as evidence for an adaptive/functional Our very of of and that there should be many traits the that an and were by its many of characters that would be species Many of these would be sufficiently to fur and middle ear or as to be for any This is not an effect of being that there by be variables with of there are likely or thousands of other traits with that could be as even we have not can we our functional or adaptive on one of these we hope that other biologists do not the other traits we the other simple from common does not that there should be many other traits with X in Figure unless these traits have some special relationship to It is indeed difficult to explain Figure 1a,b without an adaptive/functional relationship. The problem with or the as correlated evolution is that it would to the to interesting significant by comparative methods, between of for the same or between one and a with a of might argue that most of these would be for of a between the we have used such as and or and middle ear because they a causal However, the of are and there are many characters to many could the test of such an all of the on the of the with the comparative pattern to support an adaptive/functional correlation between variables. In this article we are only with what support is by the comparative patterns. explanation for the of Maddison's (1990) and Pagel's (1994) tests is that they or of as when in fact they can common this of is at the of the that This by Read and Nee and Grafen and Ridley In our methods should be to as and that the a are Pagel's (1994) to the (Fig. The for the of among the and to test the hypothesis that the of of one variable do not on the of the other the for scenarios of states and The of an of Y from 0 to 1 is to the in which it whether it in the of 0 or in the of 1 in However, different of a to 1 in Y occur in the of 1 in the does not to whether these of 1 in X are or could all be from the same As out by Grafen and Ridley the different of 1 in X are they the same of and there is not as evidence as there to be for a Read and Nee to this as of It is that this of is simply to the as to variables a However, the explanation no to and we that this effect would even the problem of were we would even the Pagel's (1994) test were to be with characters as X in 1c,d. the problem of of it would that a rather different to is When we these with we tend to three our concern that biologists may be evidence for consider to be a problem and by while are that significant in Figure do indicate biologically correlations. We have a the last we have that adaptive/functional merely from the that fur and are correlated in is little different from two data and can we simply by being and the problem a good should be of the of and the that they would not a significant result on such patterns. good in to a statistical simply the data to see how many are to the If there is only a single the result should be If there are sufficiently many the result should be However, even we how many there what would be our rule for Figure four scenarios that highlight we cannot an simply by to In Figure there are two of in X and Y⁠. two for significance, and What the is not (Fig. that the to for what the independent of X are on the (Fig. we with that there were three of in the of in Y could be explained by an variable that in the by the of X in to rule out this Four scenarios for the evolution of states of characters X and Y⁠. In each, the same phylogeny is mirrored to show X at left and Y at right. State 0 = white; State 1 = black. (a) co-distribution. Replicated but in phylogenetic Replicated In each of the scenarios of Figure we that tests like Pagel's (1994) or Maddison's (1990) would give support to a statistical but we also know that the result is to an by the of each of the the tests our us how are these and whether to consider any correlation As the of independent these would but how we do not We have no to to these methods, even a as to whether there is in that are that are not needed to It that a could be by to tests a of It would be to have all the from the of in a single The we outline tests of correlation between and are by other comparative as et al. that their method, which whether and are or in the of a particular from significant result can arise from a single and can only that the on that or any other that might be with it (Maddison et al. that a significant relationship between body and in inferred with the could as be explained by an in to the Thus, in our to methods to that of the information in the used by the methods we have the for This of to patterns that our are not to categorical data. tests for correlations between variables be susceptible as of independent can be misled by a single but this is considered a of the and can be by a that as a and Purvis 1997). the is expected to be to the because a significant result many that show an evolutionary event of in independent on phylogenetic (e.g., and and methods should be to our as they are to the et al. However, we need to be that comparative methods between variables could in be If a were to an association between the in one variable and the of in a a single of could lead it to an we that any comparative that to the effect of a rather the effect of a be susceptible to Our of the problem of would have been precise we a in a we We are that a can be because even our methods we can data that give evidence (Fig. 1a,b) from data that give no evidence (Fig. The of and Nee 1995; 2000) is one solution. It independent of species or to see the in one a in a It can the of being by a single origin of a by of that in the states of variables and Nee any patterns are on of in they cannot be explained by a single in a In this of However, the only a of or of the data (Felsenstein and would likely have to correlations and Ridley In there are many to (Maddison some of which could the may be an for but it is not an to would be a that the phylogeny into a of in each of which there is a of states of variables, and a Pagel (1994) test or other test in each a single of the test might be misled by a single the many a of correlation in the we would have that the correlation is not likely due to This the into a of It could be to a single group with of the traits of but it could also be by the of in which the two characters could be (e.g., et al. Such would from of information not all of the in the can be and there is an to the As satisfying would be a single test or of would the data as as but would give to Figure 1a,b for having the of X in three and Figure for having the of X all arise from a single evolutionary The to by information from the whole and significant in like Figure 1c,d. that the problem with Figure is the of a third by one to a may be to third characters Grafen and Ridley a with characters that but this has been used only to explain pseudoreplication, not to a statistical that would provide a solution. methods (e.g., et al. et al. could be to the effect of variables. It is to us whether of variables would provide a null and some or all of the Such a may need to account for the fact that our characters might have been or not by chance but in part on their the of by biologists not be The and us that a good likely a sense of among Thus, one would be to seek a categorical of independent has a to between the categorical states arise by on This has to be and as a but is a to to the of the the from to phylogenetic methods of correlation in a from species as they were independent to evolutionary as However, the of the methods, and the to that this to a phylogenetic It may be that we are to the of in which we live our that it is difficult for us to a phylogenetic and explain this by to the from a to an is to be by and a this in a perspective, only to be are misled by our its own is a different sense of The that in our of to be into is in fact simply a natural that is in (Fig. on in with a phylogenetic on among Our of a or is in We are to species directly, in the of we but in fact a in is the the species. on a in and of from last a or the which is in the is by phylogeny. comparative methods they the between species as the between from one living species to In we our to the our on patterns of rather patterns of The phylogenetic in comparative methods the paradigm, the of In this they the natural between species, the evolutionary to the common and (Fig. in some the methods here Pagel do not the in phylogenetic or as they provide of whether on or not, of of as methods also of as the of species. It that they the same mistake, to of rather to evolutionary This we is the of Read and Nee by selection of characters and the biological of the lead our methods to significant association coincidence is an that many of other traits in the could provide good for the evolution of a very little can be by tests in such We need methods for categorical characters that and of evidence from independent phylogenetic As long as we (or our to see an association between X and Y when we at Figure we have to the phylogenetic Our methods do not phylogenetic as as we think they we have their to such a effect of shared we are not as as we When we learn how to methods to we have to the phylogenetic We see Figure as other a natural in with no of interesting to be explained by evolutionary When the of a trait in a is by a evolutionary we might be to that there is an interesting adaptive or functional relationship between However, the trait only the association can be to coincidence followed by because any other of the same is as an explanation for the We have for decades that we need to phylogeny to independent evolutionary to and methods for correlation of categorical characters (Maddison 1990; Pagel et al. indicate association even one of the traits only This could be to a simple of with the being with the of evolution We however, that even were the to such they would be susceptible to the of single evolutionary would lead to of and Nee by many the even when they occur all with the same of one of the traits. problem could be with our characters in that characterize biologists may be not to when there is only a single scenarios with a few or show that the problem could often be and difficult to is a problem categorical as it also methods to and characters (Maddison et al. We need to our methods, and it may a rather different to We and for important the on this us to on comparative the of in mammals. We and two for many on the there is no that they would with (or we This by a from the and of and a from

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.226
metaresearch head score (Gemma)0.640
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Theoretical or conceptual · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.774
Threshold uncertainty score0.954

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.2260.640
Meta-epidemiology (narrow)0.0020.003
Meta-epidemiology (broad)0.0070.004
Bibliometrics0.0050.011
Science and technology studies0.0070.020
Scholarly communication0.0070.014
Open science0.0130.010
Research integrity0.0070.022
Insufficient payload (model declined to judge)0.0060.002

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.019
GPT teacher head0.262
Teacher spread0.243 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designTheoretical or conceptual
DomainMethods
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations510
Published2014
Admission routes2
Has abstractyes

Explore more

Same venueSystematic BiologySame topicGenetic diversity and population structureFrench-language works237,207