Population‐level anemia prevalence rates may be rendered inaccurate when hemoglobin is measured in pooled capillary blood or with the <scp>HemoCue</scp>® 301 device
Bibliographic record
Abstract
Anemia, defined as a low hemoglobin concentration, impacts ~one third of women of reproductive age (WRA) and children globally, with the highest rates in low-resource settings.1 Venous blood measured with hematology autoanalyzers are currently recommended by the World Health Organization for hemoglobin measurement; however, this is not always feasible. Portable hemoglobinometers, such as the HemoCue® (Angelholm, Sweden), are often used in large field studies or anemia surveys due to their ease of use and their ability to assess hemoglobin in a single-drop of capillary blood.1 However, there is growing concern that single-drop capillary blood introduces too much variability (random error) to hemoglobin estimates, which leads to inaccurate estimation of population-level anemia prevalence.2, 3 Pooled capillary blood offers a possible alternate blood source to single-drop capillary blood, but there is limited and mixed evidence on whether pooled capillary specimens have the same inherent issues as single-drop capillary blood4; thus, more research in this area is urgently required. The recent HEmoglobin MEeasurement (HEME) multicountry study5 evaluated repeated hemoglobin measurements using a variety of analytical methods and blood sources. Using data from our Cambodian HEME study site, we had the unique opportunity to evaluate the use of pooled capillary blood for hemoglobin measurement, and offer novel commentary for its use as compared with gold standard methods. In brief, pooled capillary blood was collected using contact-activated lancets (BD Medical, USA). After wiping away the first drop of blood, a 1 mL EDTA vacutainer (Greiner Bio-One, Austria) was placed at the base of the puncture site to collect the pooled sample (~250 μL blood; ~8–15 drops) within 2 min. Venous blood was collected concurrently in a 2 mL EDTA vacutainer (BD Medical, USA). Hemoglobin concentrations were measured with three HemoCue® models (Hb 201+, 301, and 801) using pooled capillary blood and venous blood; venous blood was further evaluated in a hematology autoanalyzer (Sysmex XN-1000, Sysmex Corp., Japan). All HemoCue® devices performed well by manufacturer standards and reference materials (HemoTrol® QC solution) tested within acceptable ranges. We calculated the mean difference (95% CI) in hemoglobin measurements across methods. Repeated measures from a total of 36 participants (18 WRA and 18 children aged 12–59 months) were included in this assessment, and compared with a repeated measures ANOVA (with Bonferroni correction for any multiple comparisons); see Table 1. Across WRA and children (data pooled), mean (95% CI) hemoglobin concentrations using pooled capillary blood measured in the HemoCue® models (201+, 301, and 801) were 3.5 (1.5, 5.5) g/L, 9.2 (7.4, 11) g/L, and 5.1 (2.6, 7.5) g/L higher, respectively, as compared with the gold standard method (venous blood via the autoanalyzer) (Figure 1). Separate data for WRA and children, and the corresponding anemia prevalence rates for each group are also presented in Table 1. This overestimation of hemoglobin concentrations driven by the use of pooled capillary blood is further evidenced in Figure 2; across WRA and children (data pooled), mean (95% CI) hemoglobin concentrations using pooled capillary blood measured in the HemoCue® models (201+, 301, and 801) were 4.4 (1.6, 7.1) g/L, 1.7 (−0.8, 4.2) g/L, and 2.8 (−0.05, 5.7) g/L higher, respectively, as compared with venous blood from the same individuals repeated in the same HemoCue® models. Ultimately, the use of pooled capillary blood resulted in a systematic overestimation of hemoglobin concentrations and concurrent underestimation of the prevalence of population-level anemia; the most predominant underestimation of anemia found via use of the HemoCue® 301 device. These findings suggest that population-level anemia prevalence rates may be rendered inaccurate when hemoglobin is measured in pooled capillary blood or with the HemoCue® 301 device. Our findings are aligned with much of the previous literature, which reports that differences in analytical methods can substantially influence hemoglobin results.6, 7 This may be driven by many factors, such as operator-experience, humidity, and the time between collection and measurement8-10; these factors typically vary across protocols and settings. Of note, blood collection and quantification techniques in the current study were performed as per a detailed protocol that was developed by anemia experts of the larger HEME study,5 which aimed to minimize potential biases (e.g., using only one comprehensively-trained operator, detailed specimen handling, time between collection, and analysis <1 min). However, this study reinforces caution against the comparison of hemoglobin estimates or population-level anemia prevalence rates across surveys or timepoints when different analytical approaches have been used. Furthermore, a crucial gap in current literature is that most hemoglobin method-comparison studies cannot tease out whether differences across methods are due to the blood source or the analytical device; we address this gap in our investigation. We found that hemoglobin concentrations measured with pooled capillary blood were systematically higher than when measured with venous blood via the autoanalyzer or with the same HemoCue® devices (suggesting that differences were due to the blood source). However, we note some uncertainty as to whether this was a result of biological differences in the blood source or simply due to the higher variability in capillary blood as compared with venous blood (as evidenced by other studies2, 3, 11). Regardless, we caution against use of pooled capillary blood specimens due to these apparent inaccuracies. Finally, while overestimation was observed across all three HemoCue® devices, the greatest inaccuracies in estimation of anemia prevalence was observed in the 301 device. Others have reported similar issues of overestimating hemoglobin concentration with the HemoCue® 301 device,10 which is particularly problematic as it is the most commonly used model for use in National Demographic and Health Surveys. If 301 devices are consistently and systematically overestimating hemoglobin concentrations, this would result in a global underestimation of anemia prevalence. It has recently been suggested that this type of systematic bias could potentially be adjusted for with the validation of each HemoCue® device before the start of a survey2; however, there is no global consensus on this approach and whether it is appropriate for all settings and populations. Further, this approach poses challenges if the directionality of the observed systematic bias differs across the range of low and high hemoglobin concentrations, as we have reported in a previous study.12 Overall, these finding are timely and of critical relevance considering the recent release of the 2024 World Health Organization guidelines on hemoglobin measurement and the revised thresholds for anemia diagnosis (https://www.who.int/publications/i/item/9789240088542). Countries on the verge of implementing DHS surveys or large anemia studies may be conflicted on what blood source and analytical device to use, given a lack of evidence or guidance regarding the use of pooled capillary blood specimens. Further research is needed to determine if procedures for the collection of pooled capillary blood could be further optimized to reduce variability and overestimation. At this time, we echo sentiments to follow the gold standard methods, which is the collection of venous blood for hemoglobin measurement with use of an automated hematology analyzer. CDK and HK conceptualized the investigation; KMC and BAW assisted with field implementation, data and blood specimen collection, and formal analysis; AC assisted with field implementation and blood specimen collection; HK supervised research staff in the field; KMC, BAW, and CDK contributed to writing the original draft; HK and AC contributed to review and editing. The authors have accepted responsibility for the entire content of this manuscript and approved its submission. We thank Ngik Rem for his assistance with training and study implementation in Cambodia. The study was funded by the U.S. Agency for International Development (USAID) under the terms of contract 7200AA18C00070 awarded to JSI Research & Training Institute. The contents are the responsibility of the authors, and do not necessarily reflect the views of USAID or the US Government. CDK is supported by a Michael Smith Foundation for Health Research Scholar Award and holds a Canada Research Chair in Micronutrients and Human Health. BAW was supported by a Michael Smith Foundation for Health Research (#180216). The authors declare no conflict of interest. Data is available upon reasonable request to the principal investigator (CD Karakochuk).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.014 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.009 | 0.006 |
| Insufficient payload (model declined to judge) | 0.011 | 0.008 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".