Definition and Measurement of Core Collaborators in Scientific Research Collaboration: A Measurement Methodology Based on the H-Index
Bibliographic record
Abstract
[Purpose/Significance] Different collaborators play different roles and assume corresponding responsibilities in scientific research collaboration. Distinguishing the different roles in research collaborators is important for the evaluation of research talent and human resources allocation. Previous studies have defined the roles of collaborators from multiple perspectives, both qualitative and quantitative, but lack a simple and efficient way to identify core collaborators. In this paper, we use the number of collaborations to identify core collaborators in scientists' collaborative relationships based on the H-index measure, which is a very easy to calculate and intuitively understandable method. [Method/Process] Using the OpenAlex database as a data source, an empirical analysis of approximately 5.05 million journal papers in the field of computing in China and G7 countries over 20 years (2000-2021) was conducted. First, the core collaborators of highly productive scientists were studied and their collaboration characteristics were analyzed from the perspective of size and share. Second, based on the H-index fitting formula proposed by previous authors, a formula for estimating the number of core collaborators based on the number of publications and the average number of collaborators per article was proposed. Finally, the formula was used to compare the differences between the theoretical and actual values of the number of core collaborators across countries. [Results/Conclusions] The study found that in terms of size and proportion of core collaborators, China had the highest average total number of collaborators among highly productive scientists, followed by the USA, Germany and the UK, while Italy had the lowest. The number of core collaborators was generally 3-7 across countries, with China and Italy having a higher rate of cooperation and the UK, France and Canada having a lower rate of cooperation. In terms of the number of core collaborators as a percentage, no country has more than 10%, with Italy having the highest percentage of core collaborators at 7.42%, followed by Japan, France and Canada, while the US has the lowest percentage of core collaborators. In terms of the total number of collaborators, there is no significant difference between China, the US and Germany, while there is a significant difference among all five other countries. In terms of the number of core collaborators, China is not significantly different from Italy and is significantly different from all other six countries. The number of core collaborators can be estimated by using the formula of the product of the number of publications and the power of the average number of collaborators per article, which has a good fit of 0.8 or more. Among China and the G7, the US, Germany and the UK have a lower proportion of core collaborators, with more frequent mobility and exchange of talent, while Italy, Japan and China have a higher proportion of core collaborators, indicating a lack of talent mobility and a relative consolidation of research collaboration.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.020 | 0.080 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.010 | 0.014 |
| Science and technology studies | 0.002 | 0.003 |
| Scholarly communication | 0.003 | 0.007 |
| Open science | 0.002 | 0.005 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".