Patient-Oriented Evidence that Matters (POEMs) Suggest Potential Clinical Topics for the Choosing Wisely Campaign
Bibliographic record
Abstract
OBJECTIVE: We propose a method of identifying clinical topics for campaigns like Choosing Wisely. METHODS: In the context of an ongoing continuing medication education program, we analyzed ratings on every patient-oriented evidence that matters (POEM) synopsis delivered in 2012 and 2013. Given the objective of the Choosing Wisely campaign, we focused this analysis on 1 specific item in the validated questionnaire used by physicians to rate POEMs. This questionnaire item is about "avoiding an unnecessary diagnostic test or treatment." For each POEM, we calculated frequencies and proportions for this item, then we identified the 20 POEMs that were most commonly associated with this item in 2012 and 2013. Finally, we determined whether the clinical topic of each of these POEMs was mentioned in the Choosing Wisely master list. RESULTS: In 2012 and 2013 we received 506,809 completed questionnaires (or ratings) linked to 530 POEMs, for an average of 956 ratings per POEM. In 59% of these POEMs (n = 312), the most commonly expected type of health benefit was "avoiding an unnecessary diagnostic test or treatment." We then identified the top 20 POEMs most commonly associated with this item in each year by ranking all 312 POEMs from the top down. The clinical topic addressed by 29 of these 40 POEMs was not addressed in the Choosing Wisely master list. These topics fell into 3 categories: diagnostic tests, medical interventions, and surgical interventions. CONCLUSION: "Big data" can identify clinical topics relevant to campaigns such as Choosing Wisely. This process represents a new way to inform the expert panel approach.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.027 | 0.017 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.004 | 0.001 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.003 | 0.001 |
| Research integrity | 0.000 | 0.004 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".