Large-Scale Comparison of Authorship, Citations, and Tweets of Web of Science Authors
Bibliographic record
Abstract
Social media platforms are increasingly part of the academic workflow. However, there is a lack of research that examines these activities, particularly at the author level. This paper explores the activity of researchers in the Twittersphere by analyzing a large database of Web of Science authors systematically identified on Twitter using data from Altmetric.com. Using this information, this paper explores and compares patterns of tweeted and self-tweeted publications with other academic activities, such as citations, self-citations, and authorship at the author level. This paper also compares the thematic orientation among these different activities by analyzing the similarity of the research topics of the publications tweeted, cited, and authored. The results show that the productivity and impact of researchers, as defined by conventional bibliometric indicators, are not correlated to their popularity on the Twitter platform and that scholars generally tend to tweet about topics closely related to the publications they author and cite. These findings suggest that social media metrics capture a broader aspect of the academic workflow that is most likely related to science communication, dissemination, and engagement with wider audiences and that differs from conventional forms of impact as captured by citations. Areas for further exploration are also proposed.1
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.045 | 0.134 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.091 | 0.310 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".