Single-Site vs Multisite Bone Density Measurement for Fracture Prediction
Bibliographic record
Abstract
BACKGROUND: Bone density measurement with dual-energy x-ray absorptiometry is widely used for fracture risk assessment. Discordance between measurement sites is common, but it is unclear how this affects fracture prediction. METHODS: We performed a historical cohort study among 16 505 women 50 years or older at the time of baseline dual-energy x-ray absorptiometry of the spine and hip (mean +/- SD observation period, 3.2 +/- 1.5 years). The study population was drawn from a database that contains all clinical dual-energy x-ray absorptiometry test results for the province of Manitoba, Canada. Each subject's longitudinal health service record was assessed for the presence of fracture codes after bone density testing. The likelihood ratio test was used to assess the improvement in fracture prediction from Cox proportional hazards models using bone density covariates from a single site or from combined sites. RESULTS: Age-adjusted hazard ratios (HRs) per standard deviation for osteoporotic fracture ranged from 1.61 (95% confidence interval [CI], 1.39-1.87) for the lumbar spine to 1.85 (95% CI, 1.70-2.01) for the total hip, with intermediate values for the femur neck (HR, 1.76 [95% CI, 1.62-1.92]) and trochanter (HR, 1.77 [95% CI, 1.63-1.92]). For fracture prediction, use of the minimum bone density measurement was no better than use of a hip measurement alone. When the total hip measurement was included in a fracture prediction model for the overall population, none of the other measurements added substantial information. The spine was the most useful site for the prediction of spine fractures alone. CONCLUSIONS: Proximal femur bone density measurements consistently outperformed lumbar spine measurements for global fracture prediction. In this cohort, the total hip was the best site for overall fracture assessment.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".