Defining ferritin clinical decision limits to improve diagnosis and treatment of iron deficiency: A modified Delphi study
Bibliographic record
Abstract
BACKGROUND: Iron deficiency is highly prevalent worldwide and is an issue of health inequity. Despite its high prevalence, uncertainty on the clinical applicability and evidence-base of iron-related lab test cut-offs remains. In particular, current ferritin decision limits for the diagnosis of iron deficiency may not be clinically appropriate nor scientifically grounded. METHODS: A modified Delphi study was conducted with various clinical experts who manage iron deficiency across Canada. Statements about ferritin decision limits were generated by a steering committee, then distributed to the expert panel to vote on agreement with the aim of achieving consensus and acquiring feedback on the presented statements. Consensus was reached after two rounds, which was defined as 70% of experts rating their agreement for a statement as 5 or higher on a Likert scale from 1 to 7. RESULTS: Twenty-six clinical experts across 10 different specialties took part in the study. Consensus was achieved on 28 ferritin decision limit statements in various populations (including patients with multiple comorbid conditions, pediatric patients, and pregnant patients). For example, there was consensus that a ferritin <30 μg/L rules in iron deficiency in all adult patients (age ≥ 18 years) and warrants iron replacement therapy. CONCLUSION: Consensus statements generated through this study corresponded with current evidence-based literature and guidelines. These statements provide clarity to facilitate clinical decisions around the appropriate detection and management of iron deficiency.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".