Google Scholar Could Be Used as a Stand-Alone Resource for Systematic Reviews
Bibliographic record
Abstract
A Review of: Gehanno, J. F., Rollin, L., & Darmoni, S. (2013). Is the coverage of Google Scholar enough to be used alone for systematic reviews. BMC Medical Informatics and Decision Making, 13(1): 7. doi: 10.1186/1472-6947-13-7 Abstract Objective – To determine if Google Scholar (GS) is sensitive enough to be used as the sole search tool for systematic reviews. Design – Citation analysis. Setting – Biomedical literature. Subjects – Original studies included in 29 systematic reviews published in the Cochrane Library or JAMA. Methods – The authors searched MEDLINE for any systematic reviews published in the 2008 and 2009 issues of JAMA or in the July 8, 2009 issue of the Cochrane Database of Systematic Reviews. They chose 29 systematic reviews for the study and included these reviews in a gold standard database created specifically for this project. The authors searched GS for the title of each of the original references for the 29 reviews. They computed and noted the recall of GS for each reference. Main Results – The authors searched GS for 738 original studies with a 100% recall rate. They also made a side discovery of a number of major errors in the bibliographic references. Conclusion – Researchers could use GS as a stand-alone database for systematic reviews or meta-analyses. With a couple improvements to the rate of positive predictive values and advanced search features, GS could become the leading medical bibliographic database. Conclusion – Researchers could use GS as a stand-alone database for systematic reviews or meta-analyses. With a couple improvements to the rate of positive predictive values and advanced search features, GS could become the leading medical bibliographic database.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.072 | 0.336 |
| Meta-epidemiology (narrow) | 0.005 | 0.006 |
| Meta-epidemiology (broad) | 0.018 | 0.010 |
| Bibliometrics | 0.111 | 0.113 |
| Science and technology studies | 0.002 | 0.005 |
| Scholarly communication | 0.022 | 0.022 |
| Open science | 0.009 | 0.020 |
| Research integrity | 0.013 | 0.008 |
| Insufficient payload (model declined to judge) | 0.250 | 0.164 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".