Do doctors who order more routine medical tests diagnose more cancers? A population‐based study from Ontario Canada
Bibliographic record
Abstract
BACKGROUND: The overuse of medical tests leads to higher costs, wasting of resources, and the potential for overdiagnosis of disease. This study was designed to determine whether the patients of family doctors who order more routine medical tests are diagnosed with more cancers. METHOD: A retrospective population-based cross-sectional study using administrative health care data in Ontario Canada. We investigated the ordering of 23 routine laboratories and imaging tests 2008-20012 by 6849 Ontario family physicians on their 4.9 million rostered adult patients. We compared physicians' test utilization and calculated case-mix adjusted observed to expected (O:E) utilization ratios to categorize physicians as Typical, Higher or Lower testers. Age-sex standardized rates (cases/10 000 patient years) and Rate Ratios were determined for cancers of the thyroid, prostate, breast, lymphoma, kidney, melanoma, uterus, ovary, lung, esophagus, and pancreas for each tester group. RESULTS: There was wide variation in the use of the 23 tests by Ontario physicians. 26% and 24% of physicians were deemed Higher Testers for laboratory and imaging tests, while 41% and 38% were Typical Testers. The patients of higher test users were diagnosed with more cancers of thyroid (laboratory [RR 1.61, 95% CI 1.39-1.87] and imaging [RR 2.08, 95% CI 0.88-2.30]) and prostate (laboratory [RR 1.10, 95% CI 1.03-1.18] and imaging [RR 1.05, 95% CI 1.00-1.10]). CONCLUSION: There is a wide variation in the ordering of routine and common medical tests among Ontario family doctors. The patients of higher testers were diagnosed with more thyroid and prostate cancers.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.063 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".