Naming Specific Adverse Effects Improves Relative Recall for Search Filters Identifying Literature on Surgical Interventions in MEDLINE and Embase
Bibliographic record
Abstract
A Review of: Golder, S., Wright, K., & Loke, Y.K. (2018). The development of search filters for adverse effects of surgical interventions in MEDLINE and Embase. Health Information and Libraries Journal, 35(2), 121-129. https://doi.org/10.1111/hir.12213 Abstract Objective – “To develop and validate search filters for MEDLINE and Embase for the adverse effects of surgical interventions” (p.121). Design – From a universe of systematic reviews, the authors created “an unselected cohort…where relevant articles are not chosen because of the presence of adverse effects terms” (p.123). The studies referenced in the cohort reviews were extracted to create an overall citation set. From this, three equal-sized sets of studies were created by random selection, and used for: development of a filter (identifying search terms); evaluation of the filter (testing how well it worked); and validation of the filter (assessing how well it retrieved relevant studies). Setting – Systematic reviews of adverse effects from the Database of Abstracts of Reviews of Effects (DARE), published in 2014. Subjects – 358 studies derived from the references of 19 systematic reviews (352 available in MEDLINE, 348 available in Embase). Methods – Word and phrase frequency analysis was performed on the development set of articles to identify a list of terms, starting with the term creating the highest recall from titles and abstracts of articles, and continuing until adding new search terms produced no more new records recalled. The search strategy thus developed was then tested on the evaluation set of articles. In this case, using the strategy recalled all of the articles which could be obtained using generic search terms; however, adding specific search terms (such as the MeSH term “surgical site infection”) improved recall. Finally, the strategy incorporating both generic and specific search terms for adverse effects was used on the validation set of articles. Search strategies used are included in the article, as is a list in the discussion section of MeSH and Embase indexing terms specific to or suggesting adverse effects. Main Results – “In each case the addition of specific adverse effects terms could have improved the recall of the searches” (p. 127). This was true for all six cases (development, evaluation and validation study sets, for each of MEDLINE and Embase) in which specific terms were added to searches using generic terms, and recall percentages compared. Conclusion – While no filter can deliver 100% of items in a given standard set of studies on adverse effects (since title and abstract fields may not contain any indication of relevance to the topic), adding specific adverse effects terms to generic ones while developing filters is shown to improve recall for surgery-related adverse effects (similarly to drug-related adverse effects). The use of filters requires user engagement and critical analysis; at the same time, deploying well-constructed filters can have many benefits, including: helping users, especially clinicians, get a search started; managing a large and unwieldy set of citations retrieved; and to suggest new search strategies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.334 | 0.715 |
| Meta-epidemiology (narrow) | 0.003 | 0.002 |
| Meta-epidemiology (broad) | 0.010 | 0.017 |
| Bibliometrics | 0.050 | 0.029 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.009 | 0.015 |
| Open science | 0.004 | 0.010 |
| Research integrity | 0.004 | 0.002 |
| Insufficient payload (model declined to judge) | 0.009 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".