Bibliometrics and the Study of Academic Knowledge Circulation
Bibliographic record
Abstract
Quantitative science studies have been used to facilitate the systematic analysis of the digital tide of literature over the last few decades. This chapter reviews two applications related to science mapping of scholarly communication. Science mapping includes visualisations to analyse potential relationships among people, organisations and concepts. Here, we review two methods: (a) co-authorship: to unveil the scientific connections of actors in the field; (b) and co-word analysis: to reveal the main themes and topics and their cognitive structure. We also present an exploratory case of study for the knowledge circulation literature. Based on a sample of 3,900+ documents published between 1996 and 2021, we found a global average annual growth of the literature of approximately 17 per cent, particularly since 2009. We then segmented the science mapping analysis into two periods: 1996–2008 and 2009–21 to track changes over time. Co-authorship analysis showed a growing institutional collaboration between periods, first led by USA–Asia Pacific institutions, then by the UK and Canada. However, the co-word analysis showed consistency between periods of the conceptual communities concerning knowledge management via Information Technology (IT) and miscellaneous in the history of medicine and research/methodologies . Relevant conceptual communities in the first period were displaced by communities such as e-learning and education economics ; research methods in public policy and sustainability sciences and environmental policy . Beginners, seasoned scholars and other actors could use our findings to identify key institutions and topics in the knowledge circulation research landscape.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.032 | 0.176 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.117 | 0.205 |
| Science and technology studies | 0.002 | 0.004 |
| Scholarly communication | 0.011 | 0.013 |
| Open science | 0.001 | 0.004 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.006 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".