Quality or quantity? Questions on the growth of global scientific production
Bibliographic record
Abstract
Quality or quantity? Questions on the growth of global scientific production Dear Editor,In recent decades, science has been able to answer questions that were thought impossible to understand or solve.The global scientific production has been gradually increasing in the last two decades, as new fields of research have been created and more and more novel hypotheses have been proposed.In particular, science is characterized by being rigorous, relevant and reproducible to provide real solutions to problems that afflict humanity.For this reason, it has always emphasized the need to focus on prioritizing quality of research over quantity [1].Recently, Scimago Journal & Country Rank (SJR) released the latest metrics for the year 2021 [2].It is very curious that compared to 2020, there was an exponential increase in the number of citable documents (corresponding mainly to original articles and revisions) in all disciplines, especially among some countries that make up the top five countries with the highest productivity (USA, China, UK, Germany and Japan, respectively).We decided to analyze the productivity of these countries and their metrics globally over the last ten years.To do so, we compared the citable documents produced by each country, total citations, citations per document, self-citations, and adjusted the citations by excluding self-citations (Table 1).We found that China, since 2010, increased its productivity by approximately 40,000 to 50,000 citable documents per year (year 2010: 344,328 to year 2021: 860,012), with this value increasing (year 2020: 771,730 to year 2021: 860,012).Over the last three years, the increase per year ranges from approximately 70,000 to 90,000 citable documents.When calculating the average output from 2010 to 2016, China obtained 438,233 documents per year, compared to the USA and UK, which obtained 663,178 and 197,400, respectively.When recalculating the average for China until 2021, an increase of more than 100,000 documents per year was observed (438,233 to 545,863); while the USA and the UK had an increase of approximately 20,000 and 12,000 documents, respectively.However, according to SJR, the country with the highest number of citations per document is the USA (29.3), followed by the UK (27.0) and China (11.6).China has the highest percentage of self-citations (57.8%),
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.019 | 0.098 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.005 | 0.012 |
| Scholarly communication | 0.010 | 0.015 |
| Open science | 0.003 | 0.004 |
| Research integrity | 0.063 | 0.040 |
| Insufficient payload (model declined to judge) | 0.012 | 0.006 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".