Poor reproducibility and inference in hydrogen-stable-isotope studies of avian movement: A reply to Wunder et al. (2009)
Bibliographic record
Abstract
Poor reproducibility and inference in hydrogen-stable-isotope studies of avian movement: A reply to Wunder et al. (2009).—In Smith et al. (2009), we tested the assumption that measurements of hydrogen stable isotope ratios in feather samples (δDf) are reproducible among independent analysis events in which feathers are equilibrated and analyzed concurrently with keratin standards (Wassenaar and Hobson 2003). For nine independent sample groups of raptor body feathers, we documented poor measurement reproducibility, with systematic error (i.e., bias) of often large magnitude and variable direction, as well as considerable random error (i.e., imprecision) in paired measurements of adjacent subsamples from a single feather. As we reported, for eight of these sample groups, initial and repeated analyses occurred at a single lab (Environment Canada's Stable Isotope Hydrology and Ecology Laboratory in Saskatoon, Saskatchewn; hereafter “EC lab”), providing robust documentation of poor δDf measurement reproducibility within a lab. A ninth group comprised samples for which initial and repeated analyses occurred at different labs (initial analysis at the EC lab and repeated analysis at the Colorado Plateau Stable Isotope Laboratory in Flagstaff, Arizona [hereafter “CPSI lab”]), providing ancillary documentation of poor measurement reproducibility between labs. Measurement precision decreased outside the calibration range of keratin standards (greater than -100‰), compared with measurements inside this range (-190‰ to -100‰) (Smith et al. 2009: fig. 2). In their letter, Wunder et al. (2009) nicely summarize some of the complexities of δDf analysis, most importantly (1) the lack of internationally accepted reference standards of a material comparable to feathers, (2) the pressing need for additional keratin working standards with high δD values that would more thoroughly bracket the range of natural δD values in bird feathers, and (3) the need for standardized analytical protocols among isotopic laboratories. We agree completely with Wunder et al. (2009) that researchers should be cognizant of these complexities, inform themselves of the analytical protocols used by the laboratory analyzing their samples (e.g., the types and number of standards), and carefully interpret δDf values outside the keratin standard calibration range. Despite this common ground, Wunder et al. (2009) make three broad criticisms of Smith et al. (2009) with which we disagree: (1) that we provided analytical detail insufficient for study replication and failed to engage the laboratories that analyzed our samples, (2) that our design failed to appropriately consider two important sources of variation (i.e., intra-feather variation and the presence of δDf measurements outside the calibration range of keratin standards), and (3) that our results contradict a substantive body of work regarding δDf measurement error. Here, we hope to clarify the main points of Smith et al. (2009) and respond to Wunder et al.'s (2009) primary criticisms, which do not lead us to alter our original conclusion of poor δDf measurement reproducibility. We discuss the effect that poor reproducibility has on inference in stable-isotope studies in the context of a recently advanced probabilistic framework for geographic assignment (Wunder and Norris 2008a, b; Wunder 2010). We also suggest some avenues toward a potential solution to the problem of poor reproducibility and encourage future practitioners in this field to more carefully consider this problem when designing studies and interacting with labs. Finally, we advise that researchers more carefully qualify their claims of the value of information that stable-isotope studies provide regarding migratory origins and connectivity at spatial scales relevant to conservation or management, because predicted origins may be biased (Smith et al. 2009) or have low geographic specificity (Meehan et al. 2001, Kelly et al. 2002, Wunder et al. 2009: fig. 1). Why the reproducibility of keratin standards may not be equivalent to that of feathers.—In. Smith et al. (2009), we described measurement error with the metric of reproducibility, the difference between repeated measurements of the same feather when one or more analytical conditions have changed (i.e., independent analysis events). We reported reproducibility for a group of samples with summary statistics where the mean (of differences) described the average systematic shift in δDf from an initial to a repeated measurement and the standard deviation (of differences) described the variability in the magnitude of this shift (Smith et al. 2009: fig. 1). Poor reproducibility was characterized by considerable systematic error (represented by a large mean) or random error (represented by a large SD), or both. Systematic and random errors have different implications for inferences of migratory connectivity that rely on measurements of δDf. Systematic error shifts the entire spatial distribution of predicted origins, whereas random error reduces the geographic specificity of predictions. In most work to date, δDf measurement error has been described by the precision (e.g., SD) of homogenized keratin standards within a single analysis (i.e., repeatability) or asymptotically over time (i.e., reproducibility). However, keratin standards typically are developed from materials that have been homogenized to exhibit high reproducibility. By contrast, nonhomogenized feather samples from wild birds lack this desirable quality. Therefore, the precision of δD measurements from keratin standards likely underestimates that of feather δD measurements. Although standard repeatability and reproducibility represent important metrics for quality assurance and quality control, they do not necessarily describe the reproducibility of nonhomogenized feather materials accurately, which is a separate metric that must be assessed independently. Thus, satisfactory repeatability, or even reproducibility, of standards does not dismiss the poor reproducibility of feather measurements documented in Smith et al. (2009), because geographic assignments are made from measurements of feathers, not standards. Analytical disclosure and communication.—Wunder et al. (2009) asserted that we failed to provide sufficient methodological detail for study replication. We agree that some ambiguity existed in the analytical details presented in Smith et al. (2009). We suggest that much of this ambiguity resulted from the editorial removal of lab names from Smith et al. (2009), a point on which we also strongly disagreed with The Auk's editors. Although we clearly indicated that samples were analyzed by only two labs, with one lab analyzing most of the samples, we did not explicitly quantify this division. In fact, 95% (402 of 422) of the δDf measurements in Smith et al. (2009) were completed at the EC lab (including all measurements resummarized below), with the remaining 20 measurements (the repeat analysis from group “NA2”) completed at the CPSI lab using the same published protocols and keratin standards as the EC lab, as we were informed by the contributors of those data (see Acknowledgments in Smith et al. 2009). With lab identities and the distribution of samples between labs now disclosed, readers should find sufficient detail for replication in our original manuscript, for two reasons. First, the discussion of reproducibility in Smith et al. (2009) focused nearly exclusively on results from the eight sample groups analyzed only at the EC lab and the comparison between laboratories was only a marginal consideration. Second, because Smith et al. (2009) primarily assessed δDf measurement reproducibility at the EC lab, our reference to the two publications of Wassenaar and Hobson (2003, 2006) that detail the exact laboratory procedures and three keratin standards used to measure δDf at this lab accords with Wunder et al.'s (2009) statement that referencing published laboratory techniques is sufficient when the work is carried out by a single lab. Our description of laboratory methods is comparable to such descriptions in recent manuscripts involving the co-authors of Wunder et al.'s (2009) letter that used the EC lab for δDf measurements (e.g., Hobson et al. 2009, Langin et al. 2009, Paritte and Kelly 2009). More troubling is Wunder et al.'s (2009) claim that we failed to communicate two important lines of information to the laboratories involved in Smith et al. (2009): that we believed they were producing “questionable” data and that the data were to be used in a publication related to reproducibility. We acknowledge that we did not have contact with CPSI lab personnel, for three reasons: (1) the 20 samples (<5% of all samples) analyzed there represented only a marginal component of our analysis and discussion, (2) data from CPSI were contributed independently of our analyses at the EC lab by an outside party that was in close contact with the CPSI lab, and (3) the CPSI lab has used Wassenaar and Hobson (2003, 2006) as primary references for laboratory protocols (e.g., Paxton et al. 2007). By contrast, however, we communicated regularly, directly, and honestly with the EC lab director, Len Wassenaar, and Keith Hobson, a long-time associate of this lab, about sample preparation, the use of analytical standards, results, reanalyses, and problems with reproducibility. This communication began at the outset of the study, continued through the analysis of data, and included the disclosure of our intent to pursue publication and the results of preliminary analyses indicating poor reproducibility. Given this history, Wunder et al.'s (2009) claim that we failed to interact with laboratory personnel to understand or interpret our results seems disingenuous. Study design and presentation of results.—Wunder et al. (2009) suggested that the results of Smith et al. (2009) are ambiguous because our study design did not account for (1) systematic δDf changes along the length of a single feather and (2) imprecise measurement of δDf outside the calibration range of keratin standards. As pointed out by Smith et al. (2009) and reiterated by Wunder et al. (2009), true replicate measurement of nonhomogenized feather samples is impossible, because feather material is destroyed during analysis. Given this physical reality, differences between replicate measurements of biological samples could result from either measurement error or real biological variation within samples. Thus, biological variation must be accounted for to adequately assess reproducibility. Wunder et al. (2009) claim that we failed to acknowledge that biological intra-feather variation confounds estimates of reproducibility, ignoring our work on this topic (Smith et al. 2008). On the contrary, the complication of intra-feather variation was the preeminent consideration in the discussion of Smith et al. (2009). Below, we provide further clarification that intra-feather variation is minor compared with the poor reproducibility observed in Smith et al. (2009). Wunder et al. (2009) incorrectly contend that we defined poor reproducibility as the widening pattern of residuals outside the calibration range of keratin standards (Smith et al. 2009: fig. 2). We agree with Wunder et al. (2009) that not distinguishing between samples inside and outside the calibration range was an oversight on our part. However, we clearly identified poor reproducibility as the considerable systematic and random error present in our entire data set (Smith et al. 2009: fig. 1), and not simply the decrease in precision we observed outside the calibration range of keratin standards (Smith et al. 2009: fig. 2). The decrease in measurement precision outside the calibration range was a secondary result and an unsurprising consequence of applying a normalizing equation to δDf values outside the range of values on which the calibration regression was based (i.e., -190‰ to -100‰). Likewise, our suggestion to expand the isotopic range of keratin standards to include more positive values was an obvious solution to the problem, although doing so is not a trivial task, as Wunder et al. (2009) explicate. More importantly, expanding the isotopic range of keratin standards would decrease only the random error of δDf measurements currently outside the calibration range to the relatively imprecise levels observed inside the calibration range; it would have no effect on the larger problem of systematic error, which occurred both inside and outside the keratin standard calibration range. In a previous publication (Smith et al. 2008), we estimated the biological magnitude of intra-feather variation for the three most common species represented in Smith et al. (2009), independent of the confounding effect of measurement reproducibility. Specifically, all samples in the previous study were run in a continuous laboratory-analysis event, with the additional safeguard of random interspersion of samples. Smith et al. (2008) documented consistent differences in δDf between adjacent longitudinal subsamples of body feathers, the magnitude of which varied to some extent among species. Feathers of Merlins (Falco columbarius) and Sharp-shinned Hawks (Accipiter striatus) showed similar differences, with more negative δDf values in distal feather material than in proximal feather material (least squares mean ± SE, combined for the two species: -9.68 ± 1.08‰; n = 29). The same pattern appeared in Red-tailed Hawk (Buteo jamaicensis) feathers, but the difference was less pronounced (-3.00 ± 1.41‰; n = 17). The magnitude and direction of these differences serve as an expectation against which reproducibility can be assessed for samples from these three species in Smith et al.'s (2009) data set. That is, if intrafeather variation in δDf were driving the poor reproducibility observed in the δDf measurement of equivalent feather subsamples in Smith et al. (2009), differences in repeated δDf measurements in these three species should average near zero once adjusted for the intra-feather variation described above. This is not the case (Fig. 1), which suggests that some factor other than the magnitude of intrafeather variation observed in raptor body feathers is responsible for poor reproducibility. To compare reproducibility inside and outside the calibration range of keratin standards, we simply distinguished, also in Figure 1, between samples that were inside (i.e., average δDf of repeated measurements greater than or equal to -100‰) and outside (i.e., average δDf of repeated measurements greater than -100‰) this range. In general, poor reproducibility of nonhomogenized feather material existed whether the analysis occurred inside or outside the calibration range (Fig. 1). Systematic error was present and often severe regardless of how the data were summarized, with mean adjusted differences between original and repeated analyses ranging from -6.14‰ to 16.78‰ within the calibration range and from -8.93‰ to 37.41‰ outside of the calibration range. Random error was larger outside the calibration range (Fig. 1), as we noted previously (Smith et al. 2009: fig. 2). Ignoring a substantive body of work.—Wunder et al. (2009) claim that our work ignores a substantive body of work on δDf measurement error. We disagree, for there is currently a paucity of literature concerning the reproducibility of nonhomogenized feather material. Certainly, the repeatability and reproducibility of homogenized keratin standards have been well reported (Wunder and Norris 2008b, Wunder et al. 2009), but the extent to which nonhomogenized δDf measurements exhibit this same reproducibility remains largely untested outside of Smith et al. (2009). Intra-feather variation within a single analysis, which should reflect real biological variation in nonhomogenized feather material that is not confounded by the problem ofreproducibility, has been studied to a limited extent (e.g., Wassenaar and Hobson 2006, Smith et al. 2008). However, these studies provide no information about measurement reproducibility among independent laboratory events. Thus, aside from Smith et al. (2009), we know of only a single, small inter-laboratory comparison of 18 nonhomogenized passerine feathers (Wassenaar 2008) that permits an assessment of δDf measurement reproducibility. Although Wassenaar (2008) did not quantify reproducibility, he presented a graph (fig. 2.5) from which he inferred good comparability of δDf measurements among labs. Nonetheless, a close inspection of this figure consistent systematic differences in δDf measurements between some on average and to an lack of intra-feather variation in passerine feathers et al. Wassenaar and Hobson 2006, Langin et al. 2007). Thus, although we do not that researchers have failed to consider isotopic variation or δD measurement error, we suggest that (Smith et al. 2009) is the published to adequately reproducibility in measurements of nonhomogenized feather material using the analysis by isotope laboratories (i.e., Wassenaar and Hobson 2003). between two δDf measurements from adjacent of a single raptor feather for eight independent sample groups during independent analysis events at Canada's Stable Isotope Hydrology and Ecology Laboratory in Saskatoon, Poor reproducibility is present for all sample systematic error is present whether samples are or of the calibration range of keratin standards used for δDf value (see whereas random error is large in both but larger of the calibration range. groups feathers from three species of Sharp-shinned and Red-tailed from the data described in Smith et al. (2009). in δDf measurements are as initial measurement repeated adjusted for natural intra-feather variation in these species (Smith et al. The at no systematic error (i.e., bias) between initial and repeated measurements. the and the and of the and the and are indicated in The along the to those in and figure of Smith et al. (2009). group species Red-tailed Sharp-shinned Merlins and Sharp-shinned Sharp-shinned Merlins and Sharp-shinned Sharp-shinned and reproducibility within the probabilistic framework of Wunder Wunder and (Wunder and Norris 2008a, b; Wunder have to the problem of geographic assignment a probabilistic framework that is well to and independent sources of on predictions. We find this framework a over previous for the origins of migratory However, Wunder et al. using the reproducibility of keratin standards as an of measurement error (e.g., Wunder and Norris 2008b, Wunder 2010). geographic assignment is based on measurements of feathers, not standards, and homogenized standard reproducibility is not necessarily of nonhomogenized feather reproducibility, we suggest that the reproducibility of nonhomogenized feather material must be estimated and this probabilistic framework concurrently but the reproducibility of keratin standards. The large and variable of systematic error documented by Smith et (2009) the of a distribution for the reproducibility of nonhomogenized feather material within the probabilistic an of the driving systematic error between analysis events. the effect of random error in δDf measurements on of be assessed using the probabilistic systematic error, by the variation (e.g., SD) the average of repeated δDf measurements from the same for intra-feather an is comparable to how Wunder and Norris the reproducibility of keratin standards and how Hobson et al. (2009) from independent analysis events are from a large number of a distribution of standard can be and Norris to represent feather δD reproducibility. However, this to the problem of systematic error among analytical events. of how the reproducibility of nonhomogenized feather material is the negative of δDf measurement error on the or specificity of inferences regarding migratory origins and connectivity be larger than has previously been reported once this error has been For this we the conclusion that δDf measurement error has minor on the specificity of geographic assignment from studies that based the of measurement error on the relatively by reproducibility of keratin standards (Wunder and Norris 2008b, Wunder 2010). for and δDf measurement important point on which Smith et al. (2009) and Wunder et al. (2009) and which we hope has been during this is that the of analyzing a single of nonhomogenized feather to represent the isotopic of an that of reproducibility account for real biological variation within feathers we have with 1). nonhomogenized feather as the material on which inferences of migratory connectivity are measurement reproducibility for nonhomogenized feathers for intrafeather variation on a for the relatively minor of intra-feather variation on Smith et al.'s (2009) results, we that Figure robust documentation of the of protocols for feather and analysis to reproducible results for the measurement of δD in nonhomogenized feather material. does this Given that homogenized keratin standards reproducibility to the of keratin standard measurement error on geographic assignment (e.g., Wunder and Norris we suggest that may be an toward δDf measurement reproducibility, the likely in time and potential (1) material is to keratin standards, in with the of for stable-isotope analyses and (2) replicate samples are from the same feather feathers from the same that would for intra-feather or (3) is to quantify the true reproducibility of δDf measurements independent of biological variation in feathers, which would measurement error to be within of how feathers are and we that measurement reproducibility is we can understand the with which we can migratory origins and connectivity using we suggest that although it may be to assess reproducibility in analysis events that are in but are independent (e.g., the to the in the δDf measurement reproducibility also should be when replicate measurements at different of the different δD and to laboratory and between different laboratories protocols and and probabilistic et al. point to the value of for a regression of δDf predicted δD in fig. as that these an for to the geographic assignment of This to the published in and Smith fig. as a that pattern must not be with of based on from calibration data such as that presented in and Smith have suggested for that the spatial of this is of an with to a of range (Meehan et al. 2001, Kelly et al. This is because it is the predicted origins of that must inform inferences of migratory origins and connectivity (Wunder 2010). The probabilistic framework by Wunder et al. (2009) has the potential to the low geographic specificity that stable in feathers to (1) studies of migratory origins or connectivity are at a spatial scales that inform conservation or (i.e., bird conservation or and (2) is using probabilistic stable-isotope studies have to at these scales for or assignments Wunder 2010). This is not the result of in which is but from our present of the variability that is to the isotope framework be in the problems with reproducibility that were identified in Smith et al. (2009) and further in this However, it is our that the variability of stable in et al. et al. and to our to and that describe how feathers this variability in and Smith 2006, Wunder and Norris 2008a, Wunder to result in low specificity of geographic assignments using stable once the of reproducibility has been Although The Auk's that presentation of the analyses to this claim was the of this we encourage readers to this claim with their data using the methods in Wunder to all of the error that is to the between in feathers and (e.g., the variation to the data set of and Smith fig. as fig. in Wunder et al. with with Wunder et al.'s (2009) claims that Smith et al. (2009) was or a for On the contrary, we hope that Smith et al. (2009) and this greater in studies that to the sources of variation that most the specificity with which migratory origins can be We agree that the variation in δDf with different sources (e.g., Wunder is an important in those sources that are most However, the publications of Hobson and Wassenaar and et al. studies we have their considerable have studies to our of such as (1) sources of variation in δDf (2) spatial and variability in the distribution of δD in and (3) the and of if problems of δDf measurement reproducibility are probabilistic assignment of an have low geographic specificity our of the between δDf and when assignments are made with a of Thus, we that researchers should discuss their inferences of migratory connectivity with the of that a field in which independent of results is and that other that likely the specificity of predicted origins, described largely within the probabilistic framework of Wunder We Wunder et al. and for us to clarify our on the reproducibility of feather δD measurements and by Wunder et al. (2009). and two provided on also provided on of The This was Smith was a at
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".