Statistical methods used in the calculation of geriatric reference intervals: a systematic review
Bibliographic record
Abstract
BACKGROUND: Geriatric reference intervals (RIs) are not commonly available and are rarely used. It is difficult to select a reference population from a cohort with a high degree of morbidity. Also important are the statistical approaches used to determine health-associated reference values. It is the aim of this study to examine the statistical methods used in the calculation of geriatric RIs. METHODS: A search was conducted on EMBASE and Medline for articles between January 1989 and January 2014. Studies were selected if they: 1) were English primary articles; 2) performed a clinical chemistry test on a blood fraction; 3) had a population sub-group consisting of individuals ≥65 years of age; and 4) calculated a RI for the subgroup ≥65 years of age. RESULTS: There were 64 articles identified, of which 78.1% described the RI calculation method used. RI calculation was performed by non-parametric (21.9%), parametric (42.2%), robust (3.1%), or other (17.2%) methods. Outlier detection (SD, Grubb's test, Tukey's fence, Dixon) was infrequently used and although most studies performed partitioning, only 57.8% tested the statistical significance of the partitions. Few studies (17.2%) reported confidence intervals for the RI estimates. Overall, only 14.1% of studies provided RI estimates which followed the CLSI guideline EP28-A3c. CONCLUSIONS: Statistical methods for RI calculation and partitioning varied considerably between studies and many failed to provide adequate descriptions of these methods. Challenges in analyses arose from insufficient sample sizes and heterogeneity in the elderly population. Geriatric RIs, although present in the literature, may not be properly calculated and should be carefully considered before applying them for clinical care.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.084 | 0.298 |
| Meta-epidemiology (narrow) | 0.003 | 0.002 |
| Meta-epidemiology (broad) | 0.013 | 0.017 |
| Bibliometrics | 0.031 | 0.028 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.005 | 0.006 |
| Open science | 0.005 | 0.003 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".