MétaCan
Menu
Back to cohort
Record W1564344853 · doi:10.18438/b8d328

Google Scholar Out-Performs Many Subscription Databases when Keyword Searching

2010· article· en· W1564344853 on OpenAlexaffvenue
Giovanna Badia

Bibliographic record

VenueEvidence Based Library and Information Practice · 2010
Typearticle
Languageen
FieldPsychology
TopicHealth and Well-being Studies
Canadian institutionsMcGill University Health Centre
Fundersnot available
KeywordsInformation retrievalComputer scienceCitationBibliographic databaseRelevance (law)ScopusDatabaseLibrary scienceWorld Wide WebMEDLINE

Abstract

fetched live from OpenAlex

A Review of: Walters, W. H. (2009). Google Scholar search performance: Comparative recall and precision. portal: Libraries and the Academy, 9(1), 5-24. Objective – To compare the search performance (i.e., recall and precision) of Google Scholar with that of 11 other bibliographic databases when using a keyword search to find references on later-life migration. Design – Comparative database evaluation. Setting – Not stated in the article. It appears from the author’s affiliation that this research took place in an academic institution of higher learning. Subjects – Twelve databases were compared: Google Scholar, Academic Search Elite, AgeLine, ArticleFirst, EconLit, Geobase, Medline, PAIS International, Popline, Social Sciences Abstracts, Social Sciences Citation Index, and SocIndex. Methods – The relevant literature on later-life migration was pre-identified as a set of 155 journal articles published from 1990 to 2000. The author selected these articles from database searches, citation tracking, journal scans, and consultations with social sciences colleagues. Each database was evaluated with regards to its performance in finding references to these 155 papers. Elderly and migration were the keywords used to conduct the searches in each of the 12 databases, since these were the words that were the most frequently used in the titles of the 155 relevant articles. The search was performed in the most basic search interface of each database that allowed limiting results by the needed publication dates (1990-2000). Search results were sorted by relevance when possible (for 9 out of the 12 databases), and by date when the relevance sorting option was not available. Recall and precision statistics were then calculated from the search results. Recall is the number of relevant results obtained in the database for a search topic, divided by all the potential results which can be obtained on that topic (in this case, 155 references). Precision is the number of relevant results obtained in the database for a search topic, divided by the total number of results that were obtained in the database on that topic. Main Results – Google Scholar and AgeLine obtained the largest number of results (20,400 and 311 hits respectively) for the keyword search, elderly and migration. Database performance was evaluated with regards to the recall and precision of its search results. Google Scholar and AgeLine also obtained the largest total number of relevant search results out of all the potential results that could be obtained on later-life migration (41/155 and 35/155 respectively). No individual database produced the highest recall for every set of search results listed, i.e., for the first 10 hits, the first 20 hits, etc. However, Google Scholar was always in the top four databases regardless of the number of search results displayed. Its recall rate was consistently higher than all the other databases when over 56 search results were examined, while Medline out-performed the others within the first set of 50 results. To exclude the effects of database coverage, the author calculated the number of relevant references obtained as a percentage of all the relevant references included in each database, rather than as a percentage of all 155 relevant references from 1990-2000 that exist on the topic. Google Scholar ranked fourth place, with 44% of the relevant references found. Ageline and Medline tied for first place with 74%. For precision, Google Scholar ranked eighth among the 12 databases when the complete set of search results was examined, but ranked third within the first 20 search results listed. Within the first 20, 55% of the search results were relevant. This precision rate put Google Scholar in third place, after Medline (80%) and Academic Search Elite (70%). Google Scholar’s precision and recall statistics may have been positively affected by its search for a keyword in the full-text content of indexed articles, rather than just searching in the bibliographic records as is the case for the other 11 databases. The author re-calculated the recall and precision rates for a title search in Google Scholar using the same keywords, elderly and migration. Compared to the standard search on the same topic, there was almost no difference in recall or precision when a title search was performed and the first 50 results were viewed. Conclusion – Database search performance differs significantly from one field to another so that a comparative study using a different search topic might produce different search results from those summarized above. Nevertheless, Google Scholar out-performs many subscription databases – in terms of recall and precision – when using keyword searches for some topics, as was the case for the multidisciplinary topic of later-life migration. Google Scholar’s recall and precision rates were high within the first 10 to 100 search results examined. According to the author, “these findings suggest that a searcher who is unwilling to search multiple databases or to adopt a sophisticated search strategy is likely to achieve better than average recall and precision by using Google Scholar” (p. 16). The author concludes the paper by discussing the relevancy of search results obtained by undergraduate students. All of the 155 relevant journal articles on the topic of later-life migration were pre-selected based on an expert critique of the complete articles, rather than by looking at only the titles or abstracts of references as most searchers do. Instructors and librarians may wish to support the use of databases that increase students’ contact with high-quality research documents (i.e.., documents that are authoritative, well written, contain a strong analysis, or demonstrate quality in other ways). The study’s findings indicate that Google Scholar is an example of one such database, since it obtained a large number of references to the relevant papers on the topic searched.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.002
metaresearch head score (Gemma)0.003
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesScholarly communication, Insufficient payload (model declined to judge)
Consensus categoriesInsufficient payload (model declined to judge)
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.842
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0020.003
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0010.000
Scholarly communication0.0000.193
Open science0.0000.000
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0020.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.040
GPT teacher head0.345
Teacher spread0.305 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designNot applicable
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations2
Published2010
Admission routes2
Has abstractyes

Explore more

Same venueEvidence Based Library and Information PracticeSame topicHealth and Well-being StudiesFrench-language works237,207