Title-plus-abstract versus title-only first-level screening approach: a case study using a systematic review of dietary patterns and sarcopenia risk to compare screening performance
Bibliographic record
Abstract
BACKGROUND: Conducting a systematic review is a time- and resource-intensive multi-step process. Enhancing efficiency without sacrificing accuracy and rigor during the screening phase of a systematic review is of interest among the scientific community. METHODS: This case study compares the screening performance of a title-only (Ti/O) screening approach to the more conventional title-plus-abstract (Ti + Ab) screening approach. Both Ti/O and Ti + Ab screening approaches were performed simultaneously during first-level screening of a systematic review investigating the relationship between dietary patterns and risk factors and incidence of sarcopenia. The qualitative and quantitative performance of each screening approach was compared against the final results of studies included in the systematic review, published elsewhere, which used the standard Ti + Ab approach. A statistical analysis was conducted, and contingency tables were used to compare each screening approach in terms of false inclusions and false exclusions and subsequent sensitivity, specificity, accuracy, and positive predictive power. RESULTS: Thirty-eight citations were included in the final analysis, published elsewhere. The current case study found that the Ti/O first-level screening approach correctly identified 22 citations and falsely excluded 16 citations, most often due to titles lacking a clear indicator of study design or outcomes relevant to the systematic review eligibility criteria. The Ti + Ab approach correctly identified 36 citations and falsely excluded 2 citations due to limited population and intervention descriptions in the abstract. Our analysis revealed that the performance of the Ti + Ab first-level screening was statistically different compared to the average performance of both approaches (Chi-squared: 5.21, p value 0.0225) while the Ti/O approach was not (chi-squared: 2.92, p value 0.0874). The predictive power of the first-level screening was 14.3% and 25.5% for the Ti/O and Ti + Ab approaches, respectively. In terms of sensitivity, 57.9% of studies were correctly identified at the first-level screening stage using the Ti/O approach versus 94.7% by the Ti + Ab approach. CONCLUSIONS: In the current case study comparing two screening approaches, the Ti + Ab screening approach captured more relevant studies compared to the Ti/O approach by including a higher number of accurately eligible citations. Ti/O screening may increase the likelihood of missing evidence leading to evidence selection bias. SYSTEMATIC REVIEW REGISTRATION: PROSPERO Protocol Number: CRD42020172655.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.095 | 0.023 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.026 | 0.003 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.002 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.009 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".