How well does the minimum data set measure healthcare use? a validation study
Bibliographic record
Abstract
BACKGROUND: To improve care, planners require accurate information about nursing home (NH) residents and their healthcare use. We evaluated how accurately measures of resident user status and healthcare use were captured in the Minimum Data Set (MDS) versus administrative data. METHODS: This retrospective observational cohort study was conducted on all NH residents (N = 8832) from Winnipeg, Manitoba, Canada, between April 1, 2011 and March 31, 2013. Six study measures exist. NH user status (newly admitted NH residents, those who transferred from one NH to another, and those who died) was measured using both MDS and administrative data. Rates of in-patient hospitalizations, emergency department (ED) visits without subsequent hospitalization, and physician examinations were also measured in each data source. We calculated the sensitivity, specificity, positive and negative predictive values (PPV, NPV), and overall agreement (kappa, κ) of each measure as captured by MDS using administrative data as the reference source. Also for each measure, logistic regression tested if the level of disagreement between data systems was associated with resident age and sex plus NH owner-operator status. RESULTS: MDS accurately identified newly admitted residents (κ = 0.97), those who transferred between NHs (κ = 0.90), and those who died (κ = 0.95). Measures of healthcare use were captured less accurately by MDS, with high levels of both under-reporting and false positives (e.g., for in-patient hospitalizations sensitivity = 0.58, PPV = 0.45), and moderate overall agreement levels (e.g., κ = 0.39 for ED visits). Disagreement was sometimes greater for younger males, and for residents living in for-profit NHs. CONCLUSIONS: MDS can be used as a stand-alone tool to accurately capture basic measures of NH use (admission, transfer, and death), and by proxy NH length of stay. As compared to administrative data, MDS does not accurately capture NH resident healthcare use. Research investigating these and other healthcare transitions by NH residents requires a combination of the MDS and administrative data systems.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.112 | 0.270 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.002 | 0.003 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".