Development and validation of search filters to identify articles on deprescribing in Medline and Embase
Bibliographic record
Abstract
BACKGROUND: Deprescribing literature has been increasing rapidly. Our aim was to develop and validate search filters to identify articles on deprescribing in Medline via PubMed and in Embase via Embase.com . METHODS: Articles published from 2011 to 2020 in a core set of eight journals (covering fields of interest for deprescribing, such as geriatrics, pharmacology and primary care) formed a reference set. Each article was screened independently in duplicate and classified as relevant or non-relevant to deprescribing. Relevant terms were identified by term frequency analysis in a 70% subset of the reference set. Selected title and abstract terms, MeSH terms and Emtree terms were combined to develop two highly sensitive filters for Medline via Pubmed and Embase via Embase.com . The filters were validated against the remaining 30% of the reference set. Sensitivity, specificity and precision were calculated with their 95% confidence intervals (95% CI). RESULTS: A total of 23,741 articles were aggregated in the reference set, and 224 were classified as relevant to deprescribing. A total of 34 terms and 4 MeSH terms were identified to develop the Medline search filter. A total of 27 terms and 6 Emtree terms were identified to develop the Embase search filter. The sensitivity was 92% (95% CI: 83-97%) in Medline via Pubmed and 91% (95% CI: 82-96%) in Embase via Embase.com . CONCLUSIONS: These are the first deprescribing search filters that have been developed objectively and validated. These filters can be used in search strategies for future deprescribing reviews. Further prospective studies are needed to assess their effectiveness and efficiency when used in systematic reviews.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.589 | 0.464 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.006 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".