MétaCan
Menu
← Back to cohort
Record W6967124126 · doi:10.5281/zenodo.10035580

The Use of Bibliometrics in the Social Sciences and Humanities

2004· report· en· W6967124126 on OpenAlexaboutno aff

Bibliographic record

VenueZenodo (CERN European Organization for Nuclear Research) · 2004
Typereport
Languageen
Field
Topic
Canadian institutionsnot available
Fundersnot available
KeywordsBibliometricsCitation indexCitationCitation analysisIndex (typography)BenchmarkingScience Citation IndexRepresentation (politics)

Abstract

fetched live from OpenAlex

The Social Sciences and Humanities Research Council (SSHRC) asked Science-Metrix to identify current practices in bibliometric evaluation of research in the social sciences and humanities (SSH). The resulting study involves a critical review of the literature in order to identify the specific characteristics of the SSH and their effects on the use of bibliometrics for evaluating and mapping research. In addition, this report presents an overview of methods of research benchmarking and mapping and identification of emerging SSH fields. This part of the report is particularly relevant because of the need to exercise considerable caution when using bibliometrics to evaluate and map SSH research. This report shows that bibliometrics must be used with care and caution in a number of SSH disciplines. Knowledge dissemination media in the SSH are different from those in the natural sciences and engineering (NSE), particularly because of the much greater role of books in the SSH. Articles account for 45% to 70% of research output in the social sciences and for 20% to 35% in the humanities, depending on the discipline. Bibliometric analyses that focus solely on research published in journals may not give an accurate representation of SSH research output. In addition, bibliometric analyses reflect the biases of the databases used. For example, the Social Science Citation Index (SSCI) and the Arts and Humanities Citation Index (AHCI) of Thomson ISI over-represent research output published in English. Original findings produced by this study show that the bias results in an estimated 20-25% over-representation of English material in the two databases. Findings from the scientific literature support those of Science-Metrix. In order to benchmark national performances and identify Canada's strengths in SSH, it is possible to use research articles published in journals representing disciplines where this medium of communication is popular, such as economics. For other disciplines, journal-based bibliometric analysis may be used with due caution and databases can be built in order to factor in other knowledge dissemination media. However, one must be wary of conducting comparative analyses of SSH disciplines without taking into account the effects of the knowledge dissemination media of each discipline on the bibliometric tools being used. Bibliometric methods have not yet been refined to the point where they can serve to identify emerging fields. In this regard, the methods with the greatest potential are co-citation analysis, co-word analysis and bibliographic coupling. However, their usefulness for policy development has been challenged. It is therefore preferable to combine bibliometrics with research monitoring and even peer review for identifying emerging fields. Another approach is to track the development of bibliometric methods, which nonetheless show promise on many fronts. In short, bibliometrics must be used carefully for SSH research evaluation. Furthermore, each discipline has its own specific characteristics, so bibliometrics is to be applied differently in each case. This report presents original findings that will help in determining how bibliometric analysis should be applied to the various SSH disciplines. It is possible to adopt at least two possible attitudes toward the challenge of offsetting the limitations of bibliometrics: a passive one (laissez faire) or a proactive one (interventionism). Given current trends such as the increased publication of articles and open access, the laissez faire approach may be the most effective way of enhancing the validity of SSH bibliometric analysis. The interventionist approach focuses on creating and optimizing databases such as the Common CV System.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.289
metaresearch head score (Gemma)0.560
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Bibliometrics
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Review · Consensus signal: none
Teacher disagreement score0.842
Threshold uncertainty score0.877

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.2890.560
Meta-epidemiology (narrow)0.0020.002
Meta-epidemiology (broad)0.0060.003
Bibliometrics0.1580.221
Science and technology studies0.0070.012
Scholarly communication0.0260.021
Open science0.0040.014
Research integrity0.0030.005
Insufficient payload (model declined to judge)0.0040.002

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.352
GPT teacher head0.357
Teacher spread0.006 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designNot applicable
Domainnot available
GenreReview

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2004
Admission routes1
Has abstractyes

Explore more

Same venueZenodo (CERN European Organization for Nuclear Research)→French-language works237,207→