The impact of multimorbidity level and functional limitations on the accuracy of using self-reported survey data compared to administrative data to measure general practitioner and specialist visits in community-living adults
Bibliographic record
Abstract
Abstract Background Researchers often use survey data to study the effect of health and social variables on physician use, but how self-reported physician use compares to administrative data, the gold standard, in particular within the context of multimorbidity and functional limitations remains unclear. We examine whether multimorbidity and functional limitations are related to agreement between self-reported and administrative data for physician use. Methods Cross-sectional data from 52,854 Ontario participants of the Canadian Community Health Survey linked to administrative data were used to assess agreement on physician use. The number of general practitioner (GP) and specialist visits in the previous year was assessed using both data sources; multimorbidity and functional limitation were from self-report. Results Fewer participants self-reported GP visits (84.8%) compared to administrative data (89.1%), but more self-reported specialist visits (69.2% vs. 64.9%). Sensitivity was higher for GP visits (≥90% for all multimorbidity levels) compared to specialist visits (approximately 75% for 0 to 90% for 4+ chronic conditions). Specificity started higher for GP than specialist visits but decreased more swiftly with multimorbidity level; in both cases, specificity levels fell below 50%. Functional limitations, age and sex did not impact the patterns of sensitivity and specificity seen across level of multimorbidity. Conclusions Countries around the world collect health surveys to inform health policy and planning, but the extent to which these are linked with administrative, or similar, data are limited. Our study illustrates the potential for misclassification of physician use in self-report data and the need for sensitivity analyses or other corrections.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.098 | 0.280 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".