Measuring research success via bibliometrics: where they fit and how they help and hinder
Bibliographic record
Abstract
What does ‘success’ in a researcher look like? While such judgements should rightly be made from the perspective of external sources – as opposed to the researcher themselves – how should these judgements be made, on what basis and for what purpose? These considerations are important – a researcher's apparent success influences their personal reputation, their competitiveness for future research grants and awards and, consequently, their career prospects. This sense of success also influences elements of a scholar's identity as represented by their employer – in marketing and profiling – and also personally in how they view themselves as scholars and interact with others (Kamler & Thomson 2006). Research success is then important – but it is also elusive. Researchers' own scholarly identities are complex fusions of perceived contributions, successes and failures (Kamler & Thomson 2006, Clark & Thompson 2015). Although central to how individual scholars are represented, perceived, funded and treated, impressions about research success can be both subjective and objective. Indeed, the science of how to measure research success – bibliometrics – is now evolving on many fronts. Many researchers are more aware of bibliometrics through the increased prominence of research networking and websites such as ResearchGate and Google Scholar – who give platforms and information to scholars and their professional networks to describe and harness their individual performance metrics. These include various research profiles/scores, citation counts, h-indexes (not to mention i10-index and g-index) and subtle exhortations or direct invitations to publish in journals with high impact factors and 5-year impact factors. Indeed, metrics to measure have become so pervasive that researchers can expend such time and effort searching, updating, profiling and benchmarking that their actual research suffers. Metrics are important to and in nursing but should also be approached with caution. (Smith & Hazleton 2011, Molzahn & Clark 2015). These metrics serve a useful purpose and give an indication of research performance and esteem but marked variations exist across disciplines and as with many metrics, some types are more susceptible to ‘creative’ manipulation, mis-interpretation and selective use that can skew markedly the meaning of these data. For instance, h-indices and citation counts vary tremendously across disciplines and depending on which citation indexing and search service is used, for example, the Web of Knowledge (Thomson Reuters), Scopus (Elsevier) or Google Scholar (De Groote & Raszewski 2012). Thus, it is not uncommon to see ‘inflated’ or highly selective data being used to maximize the status, influence, impact and, no doubt in many instances, ego of an individual researcher, particularly when they are seeking tenure, promotion or research funding opportunities. H-indexes, citation counts and impact factors increasingly influence decisions about promotion, funding and publishing but can also be misleading (Johnstone 2007, Ketefian & Freda 2009, Gallagher 2011, Polit & Northam 2011, Fitzpatrick & Madigan 2013). For example, a researcher may have 10,000 citations but an h-index of only 10 as only 10 of his/her papers have been cited at least 10 times. This can happen when one paper has been cited thousands of times in relation to comparably very much lower citation rates for other papers. To attain a high h-index one must have been named as an author on a large number of papers. This raises important and challenging issues about how to reconcile quality, quantity and visibility in publishing decisions. There have, for example, been instances of Nobel Prize winners in scientific fields with relatively low h-index due to them having published one or very few seminal papers and many other papers that were not deemed important and, consequently, not highly cited. Furthermore, the h-index does not distinguish the relative contributions of authors in multi-author articles and tends to favour authors who have been scholar longer as the index varies significantly depending on longevity of publishing. Finally, the h-index can never decrease, which can be a problem as it does not indicate changes in the productivity and influence of a scholar. Impact factors are similarly prominent but also prone to problems. For example, impact factors of journals and citations can vary widely across disciplines and when evaluating these bibliometrics it is advisable to assess them in the discipline itself rather than against different disciplines. Alternative measures, such as the Source Normalized Impact per Paper (SNIP) and Eigenfactor, can give arguably as reliable measures that are more comparable across disciplines. This is important in nursing as journal impact factors are comparably lower than other disciplines. Other important considerations in choosing a journal to submit are usually the reputation of the journal and the relevance of the journal content to the discipline (Clark & Thompson 2012). Peer, mentors and colleagues can give important guidance into the relative status of journals. High impact factor journals usually have higher rejection rates and tend to favour systematic reviews, methodology and state of the science papers and controversial opinion pieces as these are most often highly cited (Fitzpatrick & Madigan 2013). While scientific significance should be the prime consideration, citation impact and media coverage are undoubtedly assuming increasing importance in success measures. However, an over-reliance on selective bibliometric indicators of context is ill-advised. A recent notable example is a Nobel Prize-winning biologist declaring a boycott of some of the top science journals such as Science and Nature because the pressure to publish in such journals distorts the scientific process by encouraging researchers to cut corners and pursue expedience and fashion over form, function and long-term priorities (The Guardian 2013). Indeed, the Nobel Laureate mentions that the prestige of publishing in major journals has led the Chinese Academy of Sciences to pay successful authors the equivalent of £18,000 ($30,000) (The Guardian 2013). Moreover, for all those work or comparing researchers metrics across disciplines, it is important to recognize that because collaboration patterns vary across disciplines, metrics tend to vary too. (Stefaniak 2001, Lee & Bozeman 2005) For example, in some disciplines, like physics, chemistry and biomedicine, more authors are involved in studies and success is far more diffused (Vaughan & Shaw 2005). One first author may become very well-known for a study that involved the contributions of hundreds of research and staff contributors and tens of authors at multiple sites. This has led to the use of measures that are particularly useful because they are adjusted to be more comparable across different fields. For example, the SNIP measures contextual citation impact by weighting citations using the total number of citations in a subject field. As such, the impact of a single citation is given higher value in areas where citations are less likely. The SNIP measures contextual citation impact by ‘normalizing’ citation values, and takes a research field's citation frequency into account. It also considers immediacy and accounts for how well the field is covered by the underlying database. In addition to disciplinary-sensitive comparisons, peer comparisons can be used to compare researchers in similar ‘topic’ fields – allowing fair comparisons of like-with-like. This can allow departments to also benchmark the outputs of their researchers to similar departments. In addition, peer-comparisons can promote healthy competition between researchers in different fields, or even – arguably, more appropriately, comparisons competition with one's past performance. Researchers can use metrics then to ascertain more objectively and accurately the impact which their work has had. While metrics can give indication on past performance, these measures can also be used to guide researchers in understanding which elements of their work are most influential. Manuscript journals, titles and even their associated use of social media can all influence citation rates. Knowing more of the impact of particular publications can then allow researchers to better position their future work for impact. It is also increasingly common for selections, promotion and tenure committees to be provided with metrics by researchers seeking promotion. Even though many university procedures are still catching up with the full range of metrics that are available, providing these metrics to committees provides more an impartial and data-driven testament of a candidate's track-record. University policy and procedures should encourage the use of metrics and wherever possible incorporate standardized and measurements into their decision-making to promote equity across applicants. Nurses in academia, as indeed with those from other disciplines, can be prone to pleading ‘special case’ status to emphasize the unique nature of their discipline and the incomparability of wider or cross-cutting standards and expectations around impact. The well-known and often discussed flaws of metrics are often cast as illuminating the inherent folly of measuring research influence. Similarly, measurements can be used carelessly or surreptitiously by prominent researchers without recognizing their limitations and limits. Both these extremes are self-serving – putting the needs of the individual first. Metrics can be useful but should be used at the individual and organization level but always carefully, thoughtfully and ethically.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.083 | 0.159 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.321 | 0.186 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.011 | 0.003 |
| Open science | 0.004 | 0.001 |
| Research integrity | 0.001 | 0.004 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".