Small area variation in the utilization of common medical tests and consultations in Ontario, Canada
Bibliographic record
Abstract
ABSTRACT ObjectiveEvidence of regional variation in the utilization of medical tests and procedures has raised concern surrounding the potential overuse of unnecessary care. Such overuse is detrimental as it may lead to overdiagnosis, and the resulting overtreatment of indolent disease, inefficient use of resources, and rising healthcare costs. The purpose of this study was to explore small area variation in rates of commonly used laboratory, imaging, and cardiac tests and specialist consultations, and to identify factors associated with rate variations. ApproachThis is a population-based cross-sectional study in Ontario, Canada using linked, administrative databases from the Institute for Clinical Evaluative Sciences (ICES). The study population was all adults aged 40 to 75 as of January 1, 2008. We measured the age- and sex-standardized rates of 36 laboratory, imaging and cardiac tests and specialist consultations across 97 health regions in 2008 using physician and laboratory billing data. The list of tests and consultations was chosen through discussion with primary care physicians to identify procedures that are commonly used and potentially overused in primary care settings. We compared the small area rates to the Ontario rate. We calculated small area variation statistics, including the extremal quotient (EQ), coefficient of variation (CV) and systematic component of variance (SCV), for each test and consult. We used multivariable regression models to identify factors associated with health area utilization rates. ResultsAt minimum, a 10-fold difference was observed in the rates of each test and consult across the 97 health regions in Ontario, with the extremal quotients ranging from 13.6 to 54.9. When ranked in highest to lowest variation using the SCV, the tests and procedures with the greatest small area variation were limb computed tomography (EQ=49.6, CV=23.3, SCV=38.9), ferritin blood tests (EQ=42.7, CV=33.4, SCV=36.8) and vitamin B12 blood tests (EQ=40.9, CV=35.9, SCV=36.0). The test with the smallest variation was knee imaging (EQ=13.6, CV=2.1, SCV=1.7). ConclusionWe observed substantial variation across Ontario in the utilization of 36 medical tests and consultations. These findings may indicate problems with access to care in areas with low utilization, or overuse of potentially inappropriate or unnecessary medical care in areas with high utilization. Ongoing analyses are exploring determinants of area-level utilization to better understand the observed rate variations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.007 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.005 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".