Effect of a practice-based strategy on test ordering performance of primary care physicians: a randomized trial.
Bibliographic record
Abstract
CONTEXT: Numbers of diagnostic tests ordered by primary care physicians are growing and many of these tests seem to be unnecessary according to established, evidence-based guidelines. An innovative strategy that focused on clinical problems and associated tests was developed. OBJECTIVE: To determine the effects of a multifaceted strategy aimed at improving the performance of primary care physicians' test ordering. DESIGN: Multicenter, randomized controlled trial with a balanced, incomplete block design and randomization at group level. Thirteen groups of primary care physicians underwent the strategy for 3 clinical problems (arm A; cardiovascular topics, upper and lower abdominal complaints), while 13 other groups underwent the strategy for 3 other clinical problems (arm B; chronic obstructive pulmonary disease and asthma, general complaints, degenerative joint complaints). Each arm acted as a control for the other. SETTING: Primary care physician groups in 5 regions in the Netherlands with diagnostic centers recruited from May to September 1998. STUDY PARTICIPANTS: Twenty-six primary care physician groups, including 174 primary care physicians. INTERVENTION: During the 6 months of intervention, physicians discussed 3 consecutive, personal feedback reports in 3 small group meetings, related them to 3 evidence-based clinical guidelines, and made plans for change. MAIN OUTCOME MEASURE: According to existing national, evidence-based guidelines, a decrease in the total numbers of tests ordered per clinical problem, and of some defined inappropriate tests, is considered a quality improvement. RESULTS: For clinical problems allocated to arm A, the mean total number of requested tests per 6 months per physician was reduced from baseline to follow-up by 12% among physicians in the arm A intervention, but was unchanged in the arm B control, with a mean reduction of 67 more tests per physician per 6 months in arm A than in arm B (P =.01). For clinical problems allocated to arm B, the mean total number of requested tests per 6 months per physician was reduced from baseline to follow-up by 8% among physicians in the arm B intervention, and by 3% in the arm A control, with a mean reduction of 28 more tests per physician per 6 months in arm B than in arm A (P =.22). Physicians in arm A had a significant reduction in mean total number of inappropriate tests ordered for problems allocated to arm A, whereas the reduction in inappropriate test ordered physicians in arm B for problems allocated to arm B was not statistically significant. CONCLUSION: In this study, a practice-based, multifaceted strategy using guidelines, feedback, and social interaction resulted in modest improvements in test ordering by primary care physicians.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.011 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.002 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.004 | 0.002 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".