How should we measure research impact in paediatric cardiology?
Bibliographic record
Abstract
This edition of Cardiology in the Young includes two interesting analyses examining research publications in our field.In the first, Sew 1 et al use citation frequency to evaluate the most cited manuscripts and authors in paediatric cardiology.Their analysis provides a powerful review spanning from the 1939 landmark publication by Robert Gross and John Hubbard 2 on surgical ligation of a patent arterial duct to the 2014 work by Ariane Marelli 3 and colleagues on the lifetime prevalence of CHD in Quebec, Canada.The second manuscript by Loomba and colleagues 4 uses impact factor 5 and the h-index 6,7 to evaluate citation potential for research published in the paediatric cardiology-focused journals Cardiology in the Young, Pediatric Cardiology, and Annals of Pediatric Cardiology, compared to two adult-centric cardiology journals, Circulation, and the Journal of the American College of Cardiology.During the 2014 calendar year, 45 paediatric cardiology manuscripts were published in Circulation and JACC while 783 were published in the3 paediatric cardiology journals.Articles published in Circulation and JACC were cited more frequently; however, the h-index was higher for paediatric cardiology journals, and perhaps most noteworthy, amongst the 50 most cited paediatric cardiology manuscripts, 28 were published in paediatric cardiology journals compared to 22 in Circulation or JACC.The authors make a compelling case that publication in paediatric cardiology journals in no way diminishes the potential for citation, and that authors should use other factors, most importantly potential readership, when deciding where to submit research findings.These papers beg the question -how should the impact of research be measured?The analyses cited above rely on several metrics -citation frequency, impact factor, and the h-index. 7hat do these measures represent and are they reliable estimates of research quality or impact?At the most basic level, the number of peer-reviewed published manuscripts and the number of times manuscripts are cited seem like reasonable measures of the impact of a researcher.However, even at this level, as Sew and colleagues suggest, publication biases in favour of work from certain countries or institutions, more common conditions, positive findings, agreed-upon practice, readership interest, etc., can strongly influence the ability to publish.Even amongst researchers with similar number of publications, different researchers may have had different levels of impact, where a few individuals made critical contributions, and others were also included as co-authors but could never have conceived of or completed the research independently.Research and researchers also differ in how much original content to incorporate in manuscripts, with some studies yielding multiple divided or duplicative publications.Citation indices were designed to quantify the impact of research beyond number of publications, but are also imperfect measures.The journal "impact factor", represents the average number of citations for manuscripts published by a journal over the preceding 2 years, 5 but this measure will vary based on how large a specialty is.For example, adult cardiology is a much larger field than paediatric cardiology, so adult-centric manuscripts are inherently likely to be cited more frequently.Circulation, which publishes many more adult-centric manuscripts than paediatric, has a 2020 impact factor of 23.6 meaning that the average manuscript published in Circulation in 2018 and 2019 received 23.6 citations.By comparison, there is no paediatric cardiology journal with an impact factor above 2.0.Yet, as Loomba and colleagues demonstrate, 28 out of our field's 50 most cited manuscripts were published in these "lower impact" journals.In general, it is difficult to compare the number of citations as a measure of impact in fields of differing sizes.There can also be cultural differences amongst specialties, such as the common practice of naming operations after individuals, and then always citing first reports, both less common in medical specialties.Similar to the impact factor, the h-index is used to compare journals but is also applicable at the author level.The h-index is defined as the maximum value of h such that the given author/journal has published h papers that have each been cited at least h times. 6It was designed to overcome biases from extreme citations (i.e., impacting the average number of citations), but is itself skewed by the overall volume of publications.As Loomba et al demonstrate, the h-index was actually higher for the paediatric cardiology specialty journals.This finding is slightly misleading as only 45 manuscripts were published in Circulation or JACC, meaning the maximum attainable h-index was 45, while there were 783 manuscripts published in paediatric cardiology journals.Despite nuances and limitations, Loomba and
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.240 | 0.678 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.006 | 0.004 |
| Bibliometrics | 0.043 | 0.055 |
| Science and technology studies | 0.003 | 0.011 |
| Scholarly communication | 0.023 | 0.030 |
| Open science | 0.005 | 0.009 |
| Research integrity | 0.005 | 0.009 |
| Insufficient payload (model declined to judge) | 0.005 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".