Group-Based Trajectory Modeling of Citations in Scholarly Literature:\n Dynamic Qualities of "Transient" and "Sticky Knowledge Claims"
Bibliographic record
Abstract
Group-based Trajectory Modeling (GBTM) is applied to the citation curves of\narticles in six journals and to all citable items in a single field of science\n(Virology, 24 journals), in order to distinguish among the developmental\ntrajectories in subpopulations. Can highly-cited citation patterns be\ndistinguished in an early phase as "fast-breaking" papers? Can "late bloomers"\nor "sleeping beauties" be identified? Most interesting, we find differences\nbetween "sticky knowledge claims" that continue to be cited more than ten years\nafter publication, and "transient knowledge claims" that show a decay pattern\nafter reaching a peak within a few years. Only papers following the trajectory\nof a "sticky knowledge claim" can be expected to have a sustained impact. These\nfindings raise questions about indicators of "excellence" that use aggregated\ncitation rates after two or three years (e.g., impact factors). Because\naggregated citation curves can also be composites of the two patterns,\n5th-order polynomials (with four bending points) are needed to capture citation\ncurves precisely. For the journals under study, the most frequently cited\ngroups were furthermore much smaller than ten percent. Although GBTM has proved\na useful method for investigating differences among citation trajectories, the\nmethodology does not enable us to define a percentage of highly-cited papers\ninductively across different fields and journals. Using multinomial logistic\nregression, we conclude that predictor variables such as journal names, number\nof authors, etc., do not affect the stickiness of knowledge claims in terms of\ncitations, but only the levels of aggregated citations (that are\nfield-specific).\n
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.003 |
| Open science | 0.003 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".