Variation in medical practice: getting the balance right
Bibliographic record
Abstract
Contemporary clinical practice is characterized by its complexity as the volume and diversity of medical interventions, whether they are drugs, procedures or diagnostic tests, are increasing and threaten to overwhelm our capacity to deliver patient-centred care. Consider some statistics: the average American citizen can expect to undergo seven operations in their lifetime, 10% will undergo an MRI scan annually (three times higher than the rate in neighbouring Canada) and 50% of Medicare beneficiaries are prescribed five or more medications. In Ireland, one-fifth of the whole population aged over 70 years are taking long-term Proton Pump Inhibitor (PPI) therapy.1–3 The consequences of this phenomenon for patients in terms of benefit (increase quantity and quality of life) versus harm (medicalization of a person, side effects of therapies and costs to the health service budget) give rise to questions concerning the epidemiology of health care utilization and how it differs between and within countries. Seminal work carried out by John Wennberg, a health services researcher and epidemiologist who developed the Dartmouth Atlas Health Project (www.dartmouthatlas.org), has produced an emerging science that examines variation in medical practice and raises important questions about what constitutes ‘appropriate’ health care. This editorial outlines the taxonomy of medical practice variation with clinical examples showing how it relates to family medicine. Medical practice variation may be grouped into three categories each with different implications for patients, clinicians and policy makers.4
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.073 | 0.241 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.004 | 0.002 |
| Bibliometrics | 0.007 | 0.006 |
| Science and technology studies | 0.009 | 0.047 |
| Scholarly communication | 0.028 | 0.065 |
| Open science | 0.004 | 0.017 |
| Research integrity | 0.020 | 0.025 |
| Insufficient payload (model declined to judge) | 0.009 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".