No silver bullet: different soil handling techniques are useful for different research questions, exhibit differential type I and II error rates, and are sensitive to sampling intensity
Notice bibliographique
Résumé
… inferences from all studies (e.g. soil chemistry, soil microbial community composition), not just soil biota effects studies, are nearly certain to be invalid if they are derived by performing tests on mixtures of soil from multiple experimental units. … without exception, the MSS approach of mixing together soils from multiple experimental units is fatally flawed. We suggest a fulsome discussion is warranted before such strong and potentially influential statements be accepted as facts. We question whether a blanket prohibition on sample pooling is supported by their simulation, feasible, or philosophically justified. Here we address three specific issues: (1) different research questions and goals require different methodologies; (2) MSS and ISS impact both Type I and Type II error rates; and (3) sampling intensity matters. Whether soils collected from two regions have differential effects on plant growth, mediated by soil biota, is the key type of research question addressed in Reinhart & Rinella (2016). Their argument against the MSS approach was buttressed by a simulation addressing Type I error rates in both the MSS and ISS soil handling scenarios. Although their question has a narrow scope, we agree with the authors’ assertion that pooling subsamples is widespread in numerous types of plant–soil studies. We further suggest sample pooling occurs in nearly all subdisciplines of ecology and the environmental sciences (e.g. multiple subplots to estimate community structure, aggregating many small cores for estimates of soil fertility, using multiple leaves for chemical determination, etc.). However, not all studies which pool samples necessarily ask research questions of the type addressed by Reinhart & Rinella (2016). The implied assertion of Reinhart & Rinella (2016) is that studies using the MSS approach necessarily are focused in testing whether soil biota effects vary among regions. Of the articles they cited as using the MSS approach, this objective was true for some (Van der Putten et al., 1993; Nijjer et al., 2007; Felker-Quinn et al., 2011; Pendergast et al., 2013; Rodríguez-Echeverría et al., 2013; Yang et al., 2013; Gundale et al., 2014; Hilbig & Allen, 2015) but not others (Pizano et al., 2014; Larios & Suding, 2015). Additionally, many authors explicitly recognized the MSS approach reduces spatial variation in microbe abundances and/or species composition (Rodríguez-Echeverría et al., 2013; Yang et al., 2013; Pizano et al., 2014; Larios & Suding, 2015), indicating that researchers recognize the difference between the two approaches and its implications on hypothesis-testing. Even this nonrandom sample of articles, drawn from the Reinhart & Rinella (2016) reference list, demonstrates MSS is used to address multiple types of questions, and in some cases, spatial homogenization is desired, rather than a confounding effect. However, use of MSS in some conditions does not necessarily warrant its use in all conditions. As we explain later, understanding the specific research question being asked will be critical to determining the more appropriate experimental design. We generally agree with Reinhart & Rinella (2016) that when testing whether regions differ in some characteristic (e.g. soil pathogen densities), multiple independent samples taken from each region are a statistical and philosophical necessity. Similarly, if one wanted to test whether microbial community structure differed among regions, ISS is preferred over MSS (Talbot et al., 2014). However, if one wanted to test whether the average pathogen density found in each of two regions differentially effects plant growth, then creating soils of average pathogen density becomes the methodological necessity, which can be most easily done using an MSS approach. Similarly, if one wanted to test whether the microbial species pools of two regions differentially effect plant growth (e.g. Karst et al., 2015), then it is critical to try and ensure soil treatments include as much of the regional species pool as possible, something for which MSS is appropriate. Further, an MSS approach may be more appropriate than ISS when testing whether soil amendments (e.g. mycorrhizal inoculum) enhance restoration, reclamation or agricultural production, as consistency among samples can be a prerequisite to developing an effective and evidence-based management strategy. The specific needs of a given project will determine proper methodological decisions, including choosing between ISS, MMS, or a combined approach. We suggest it is inaccurate to imply that studies involving soils of different regions are only, or even principally, focused on testing regional differences. This view is not in contrast to the ideas presented in Reinhart & Rinella (2016), though we emphasize the need to ensure specificity in the research question before reaching a conclusion about the appropriate methodology. There is no methodological silver bullet of a single best answer. The critical finding that emerged from 100 simulations presented by Reinhart & Rinella (2016) was that MSS resulted in a c. 50% Type I error rate, while ISS had a 3% Type I error rate. We thank Reinhart & Rinella (2016) for including their code in their Supporting Information, and for answering simulation-related questions as we developed this paper, allowing us to rerun their simulations (Supporting Information Notes S1; Mathematica 9, Wolfram Research, Champaign, IL, USA). To increase precision in interpreting results, we ran 1000 rather than 100 simulations, and confirm their original findings (Type I error rates: MSS = 58%; ISS = 5%). We noticed, however, that their Fig. 2(a) (Reinhart & Rinella, 2016) indicates that though ISS generated approximately the expected Type I error rates, it also resulted in large confidence intervals around their point estimates of the response ratio of regional differences of soil pathogen impacts on plant growth. Reinhart & Rinella (2016) state that the difference in confidence intervals is due to contributions to residual variance. Nonetheless, we question whether the simulation and parameter set defined by Reinhart & Rinella (2016) could detect biologically meaningful differences among regions. Or more simply, though the ISS had low Type I error rates, were its Type II error rates acceptable? To explore Type II error rates, we modified the original data Reinhart & Rinella (2016) used to create their pathogen density probability distribution (SI1). Specifically, they created their probability distribution based upon three nonzero values (38.4, 9.6, 317.5) of pathogen densities (propagules g−1 soil) in each of three sites. In our first simulation, we used these values to create the probability distribution of one region, and increased them by 10% to create a probability distribution for a second region (42.2, 10.6, 349.3). We conducted all subsequent analyses as per their original code (except with 1000 simulated experiments), and estimated Type II error as the number of simulated experiments in which confidence intervals of the response ratio overlapped zero. This was then repeated with a defined difference among regions of 40% (53.8, 13.4, 444.5) and 900% (384.0, 96.0, 3175.0). Conventional practice would suggest effective methods should allow for the detection of moderate effect sizes with a power of 80% (Cohen, 1988). In all cases, ISS had a very high Type II error rate (Fig. 1a), including a 93% Type II error rate when there was a 900% difference in the underlying mean pathogen density probability distributions between regions. The failure of ISS to detect large differences does not equate to MSS being a preferred method. Though Type II error rates using MSS were c. 50% lower than ISS, they were still too high to be of value in ecological studies (Fig. 1a). Further, hidden beneath the Type II error rates are cases where a significant difference was found among regions, but in a direction opposite to that found in the underlying probability distribution. Examination of the simulated experiments found these errors of direction occurred approximately equally for MSS and ISS and their frequency decreased with increased difference between regions. Neither ISS nor MSS had a Type II error rate approaching 20%, and thus neither method represented a defensible balance between Type I and Type II error rates. One logical conclusion is both methods are fundamentally flawed and should be avoided in future studies. An alternative conclusion is our results, and those of Reinhart & Rinella (2016), reflect the data structure and simulation parameters used, rather than evidence that ISS and MSS are flawed methodologies. We identify two key components that could cause these high Type II error rates: (1) the pathogen–plant growth regression used in the simulation; and (2) the number of samples (i.e. sites) within regions, which we refer to as sampling intensity. Here, we will not discuss the former, other than indicate that the exact shape (linear vs nonlinear) and fit of this regression will be highly species- and site-specific, with consequences for Type I and II error rates, regardless of the soil handling methodology used. Instead, we focus on the more critical issue of how inadequate sampling intensity can result in a failure to provide a robust statistical test of a hypothesis, regardless of what soil handling method is used. ‘… for a given sample size, n, the value of α is related inversely to the value of β. That is, lower probabilities of committing a Type I error are associated with higher probabilities of committing a Type II error; and the only way to reduce both types of error simultaneously is to increase n’ (Zar, 1996). Thus, we explored the impacts of sampling intensity on Type I and II error rates for both ISS and MSS. The original simulation used by Reinhart & Rinella (2016) sampled soils from 10 sites within each of two regions. Further, the probability of a nonzero value of pathogen density at a given site was 30%, meaning that, on average, seven of 10 sites in a region had no simulated pathogens while nonzero pathogen densities were drawn from a normal distribution with parameters set as described earlier. To determine whether Type I and II error rates associated with the ISS and MSS methodologies varied with sampling intensity, we repeated our prior simulations with different levels of sampling intensity: 30% occurrence and 10 sites; 100% occurrence and 10 sites; and 100% occurrence and 100 sites. Increased sampling intensity had very clear effects on error rates (Fig. 1), with increases in the number of sites with nonzero pathogen densities reducing Type II error for both MSS and ISS, and reducing Type I rates for MSS. Despite these improvements, Type II error rates for ISS with 100 sites remain disconcertingly high (power of only 25% with a 40% difference in underlying probability distributions). We found 500 sites per region were needed in the ISS approach to reach a power of 80%. The MSS had substantially higher power in nearly all scenarios (Fig. 1). The effects of sampling intensity on both types of error rates are not surprising, as increased sample size increases the likelihood that the sampled values reflect the underlying statistical probability distribution. Low sampling intensity can result in population estimates substantially different from the mean values of the underlying probability distribution, which we observed occurred frequently in the 30% occurrence and 10 sites simulation. Consequently, with low sampling intensity, though estimates of pathogen densities from two regions may come from the same probability distribution, the actual sample populations may differ greatly in characteristics (mean, standard deviation, etc.). The degree of potential mismatch between the probability and sample distributions will be a function of the data structure (e.g. coefficient of variation). The strikingly large number of samples required to obtain sufficient power using the ISS methods imposes obvious logistical complications. Increased sample numbers enhance the complexity of field sampling, as well as in the size of the resulting glasshouse/bioassay-type studies. We emphasize that the specific research question being addressed will inform whether or not increased costs are justified, and may lead to the development of a hybrid-type approach which balances costs and benefits of each method. We recognize that the specifics of our modified simulation depend upon the data structure and parameter sets used, with the same limitation of scope as found in Reinhart & Rinella (2016). For those conducting field-based research in these areas, we reinforce the call for using power analyses to estimate necessary sampling intensity to detect a desired effect size (Cohen, 1988). Soil handling decisions are of great importance to many facets of ecology, including in the context of plant–soil biota interactions. The relative benefits and costs of different approaches will be dependent upon the specific research questions being asked, the feasibility of different experimental designs, and numerous other technical details. For example, ISS designs allow for measures of within-site variation, which is critical to answer certain research questions. However, ISS also has reduced statistical power and increased logistical complexity relative to MSS. Thus, MSS is recommended when accurate estimates of within-site variation are not needed to answer a specific research question. More broadly, an even more fundamental concern needs to be an evaluation of sampling design in relation to testing biological effects and inference. It is critical that any sampling program ensure that the statistical distribution of samples adequately reflects the underlying biologically determined probability distribution. If undersampled, then any downstream statistical analyses will have insufficient power to detect biological differences in the samples, rendering discussion of ISS vs MSS moot. It is self-evident that these decisions will be highly species and system specific, and contrary to the conclusion of Reinhart & Rinella (2016), we believe no blanket recommendation is justifiable. The authors thank Kurt Reinhart and Matthew Rinella for assisting them in understanding the specifics of their simulation, the underlying data structure, and interpretation of their results. The authors thank Colleen Cassady St Clair for helpful discussion which strengthened this manuscript. J.F.C., J.K., G.J.P. and N.E. identified the key problems in the source manuscript, T.B., J.A.C. and G.J.P. modified and ran the simulations, T.B., J.F.C. and J.A.C analyzed the simulation data, J.F.C. wrote the manuscript, T.B. prepared the figure, and all authors edited the text. Please note: Wiley Blackwell are not responsible for the content or functionality of any Supporting Information supplied by the authors. Any queries (other than missing material) should be directed to the New Phytologist Central Office. Please note: The publisher is not responsible for the content or functionality of any supporting information supplied by the authors. Any queries (other than missing content) should be directed to the corresponding author for the article.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,000 | 0,001 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,001 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,001 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».