On Lampposts, Sneetches, and Stars: A Call to Go Beyond Bibliometrics for Determining Academic Value
Notice bibliographique
Résumé
Academic scholarship is something that we have always sought to quantify. Journal impact factor, CiteScore, h-index, hi-10, and the alternative metrics (Altmetrics) are terms that are familiar to seasoned academics.1 Junior scholars are often advise to use bibliometric and publication metrics to decide where to send their work.1, 2 Whether we are talking about journal-level, article-level, or person-level metrics, these are all human constructs that have arisen as ways to quantify the reach and impact of our academic work. In the paper by Boudreaux and colleagues3 in this issue of Academic Emergency Medicine, we are introduced to a new way to benchmark scholarly productivity for faculty members. Using a norm-based methodology (e.g., publications over a certain amount of time) and impact (e.g., citations that each publication gets within that group) to compare an individuals’ scholarly output against that of his or her peers. We congratulate the authors on taking on this difficult problem. So often, external letters for promotion state “Based on their productivity, I believe this person would be promoted at my institution” are at best an estimate and likely a guess. Similarly, tools that go beyond popularity contests and provide metric-based benchmarks for an individual (or even a departments productivity) are welcome. But, in this editorial we ask: Are the metrics we often use the right ones? Are they equitable for individuals across disparate systems? What are the ramifications to our academic culture when we choose to use certain metrics over others? Might we create perverse incentivizes and make ourselves vulnerable to predators? Boyer et al. articulated how we should reconsider scholarship in modern academia in the 1990s by proposing four broad categories of scholarship: 1) The scholarship of discovery (i.e., what we may see as “bench” or epidemiologic research), 2) the scholarship of application (i.e., clinical research or implementation science), 3) the scholarship of integration (i.e., where we do interdisciplinary work, or “remix” prior work via knowledge syntheses), and 4) the scholarship of teaching (i.e., which has resulted in much of the educational research and innovation that we see in medical education).4, 5 Boyer was followed by Glassick,6 as well as Hutchings and Shulman,7 who helped to formalize standards for assessing scholarship—and more specifically the scholarship of teaching, which was deemed to be particularly troublesome to define. Box 1 lists Glassick's proposed list of standards for judging scholarship and additional criteria for the scholarship of teaching by Hutchings and Shulman.7 For all forms of scholarship: Clear Goals Adequate Preparation Appropriate Methods Significant (i.e., Important) Results Effective Presentation Reflective Critique Specifically for the scholarship of teaching: Must be made public Must be available for peer review & critique according to accepted standards Must be able to be reproduced and built upon by other scholars. When talking about metrics, it brings to mind the fable about a man and a lamppost.8 The story goes that the man is on his hands and knees under a lamppost when a bystander walks by and asks him if he needs help. The man looks up and explains that he is looking for his keys. The bystander steps in to help, and after searching for a while asks: “Are you sure you lost the keys around here?” He looks up and states: “Actually, I think I dropped them at the end of the block, but it's too dark there, the light is much better over here …” This story is a cautionary tale to those who use convenient measures, rather than the measures that may be more relevant but harder to measure (e.g., using hemoglobin A1c instead of development of diabetic neuropathy; board exams scores as a surrogate marker for clinical competence, rather than ability to resuscitate a patient in the workplace).8 Convenience measures may incentivize the wrong behaviors and keep you searching under the wrong lamppost. Academia has largely embraced a quantitative approach to measuring productivity via bibliometrics (see Table 1 for partial listing). We have moved from simply counting work (e.g., pure numbers of publications, citations) to incorporating increasing levels of sophistication to determine the impact of his or her works. Newer metrics such as the one proposed by Boudreaux et al. combine metrics in new ways. Other examples of these are more advanced concepts such as the h-index1 (a person-level impact factor) and alternative metrics (publication metrics that harness social media and mainstream media uptake to quantify impact with end-users).9-11 Similarly, normalized metrics proposed by Boudreaux and colleagues may be convenient as surrogates for the quality of an academic's scholarship, rather than a true indicator of quality. CiteScore (journal-level metric) The “citability” of an author's most important works.1 In emergency medicine, this has been proposed to help judge academic performance and identify individuals of high scholarly potential.12, 13 Can be used as a measure of scholarly discussion around a given article via these online platforms.9-11 Has been correlated with later citation metrics in some fields.14 Quantitative metrics such as journal impact factor as a surrogate for quality, but it is well known that journal impact factors do not always correlate with the quality of the work contained within the journal.1, 15 While it is a reasonable inference that there must be some correlation between a paper's worth and its citations, it's important to realize this may not always be true. Controversial papers may be highly cited. And while publication and peer review process is supposed to be another check and balance to judge the quality of one's work, merely being published may also not be a high enough bar to judge high-quality work.15-17 One only need to look at the debacle over the Andrew Wakefield incident (which, per Google Scholar has been cited 31 times to date) and realize that even high-impact journals with expert editors and reviewers can be led astray.18 It should be no surprise that there are some championing the cause of evaluating scholarship via quality metrics—all of which can assist individuals to understand what makes high-quality learning materials. In the age of JAMA User's Guides19 and reporting guidelines some groups are lobbying for the development of quality assessment tools to guide individuals to determine the quality of individual works of scholarship. Interestingly, a recent flurry of work in this area has been done by those groups that are engaging in online or digital scholarship,20-23 likely because disruptive and ubiquitous resources have under the most scrutiny, resulting in the responsive quality assessment agenda.24-26 Quality measures exist such as ratings of papers within certain journals. One new, PubMed-indexed journal (Cureus.com) has the Scholarly Impact Quotient rating, which asks individual readers to rate articles based on 10-point Likert scales for six criteria: 1) clarity of background and rationale, 2) clinical importance, 3) study design and methods, 4) data analysis, 5) novelty of conclusions, and 6) quality of presentation.27 There is no doubt that metrics and quality scores abound. But at the end of the day, when we are examining the best practices for determining scholarly work, quantifying publications/citations move us away from appreciating alternative scholarship, especially in medical education and quality improvement. Theodor Seuss Geisel (a.k.a. Dr. Seuss) warned us about the ramifications of valuing certain types of phenotypic markers over others.28 In his book The Sneetches and Other Stories, he tells the cautionary tale of an opportunistic charleton (Sylvester McMonkey McBean) who sees that a group of individuals (Plain-Belly Sneetches) are desirous of a certain phenotypic marker (i.e., a star on their belly) and sells them this … at a price.28 Although a children's tale, this is very much mirrored in our current academic era where the quantification of performance metrics are en vogue. We understand the scope of this paper is measuring publications and very much agree with the authors comment in the results where they state that “… [m]any academicians may not be expected to perform research or publish papers, so including them in the database may not be appropriate”3 for the very important fact we must be cautious not to singularly use metrics that favor one type of scholarly work over others. The substantial expansion of in EM training programs into community-based health systems without research infrastructure creates a divide, effectively marking some of our centers a group with “stars,” instead of celebrating new academics for their talents. For instance, individuals toiling away setting up a new residency programs or those working in innovative frontiers may spend valuable time involved in curriculum design or administration and may not register initially on the proposed norm-based bibliometric scale. We have moved the needle on academic productivity in the areas of discovery and possibly application, but for the areas like integrative innovation or educational delivery an over emphasis on established metrics may moves us away from fostering successes. The launch of our sister journal AEM Education and Training29 or the indexing of the AAMC's MedEdPortal.org may hold promise in creating avenues for encouraging new types of scholarship. Indeed, with dedicated attention, we know it is possible to encourage educational scholars toward achieving comparable publication-based metrics,30 but this may still bias us away from important work such as curricular development, administration, or journalism. In addition, opportunistic companies abound knowing this need for publishing metrics. We must be wary that incentivizing pure numbers have given to the rise of predatory journals.31, 32 Inexperienced academics are easy prey for these notorious entities (who often write e-mails loaded with complimentary language such as: “Dear esteemed and honorific professor”) and the pressures can be so unbearable that some are more than happy to provide some “stars upon thars for just a few dollars eaches.” As academic leaders, if we persist in merely creating new mathematical formulae to combine numbers that are most easy to attain, are we not merely asking all our faculty to spend their time chasing surrogate measures that we hope represent quality? And by not measuring other important and valuable work in a meaningful ways, will this not encourage a drift away from what matters toward what counts? Ultimately, as a specialty, we should and must decide on what is important to us—and then pursue ways to quantify and qualify these things effectively. Table 2 lists the present metrics that we often use for promotions processes. Academicians are incentivized to publish more papers. May contribute to “publish or perish” or undue stressors. May encourage poor academic practices (e.g., “salami slicing”).35 May result in inexperienced scholars being lured into publishing within predatory journals. Academicians may perceive that only work which is cited is worthy. May lead to a devaluing of high quality work that May encourage junior faculty to pursue forms of scholarship that are more likely to be cited (e.g., guideline work or systematic reviews/meta-analyses) Academicians may view that certain journals are more important than others. May lead to devaluing of high-quality work which is published in a less prominent journal, but eventually is discovered by modern search technologies (e.g., Google Scholar) and is highly cited and read. Academicians may continually chase grants, instead of actually performing the work. May lead to a partitioning of academicians who exclusively write grants but may not always participate fully in the fulfillment or authorship of the works in a personally fulfilling way. May lead to academicians who “chase” trendy topics to research, swayed by granting agencies and politics, rather than choosing topics of importance to our field. Academicians may seek to supervise more trainees than they can handle. May lead to academic abuse and poor quality mentorship. Academicians may seek opportunities more aligned with external measures (e.g., national or international organizations) rather than contribute to their local academic institutions or environments. May lead to nationally or internationally renowned academicians that do not contribute to your local environment. We believe that while the present paper in our journal represents a start of a norm-based movement for bibliometric-based metrics, there is much work ahead of us in this field before we can adequately quantify the true productivity of our field. While for some academics (e.g., clinical or health services researchers) these metrics may resonate, they do not encapsulate all the metrics that reflect all types of high-value scholarship that comes from our specialty. We propose that those interested in these areas be inspired by Boudreaux et al. to better quantify and qualify the academic work that our specialty has to offer. Table 3 proposes some value-driven outcomes that could be integrated into academic merit or promotion systems to foster a more holistic academic environment for a diverse set of scholars. We hope that by listing some of these other valuable endeavors and the possible metrics that we may encourage others to quantify some of these other metrics for best comparison within emergency medicine. Ultimately, perhaps a more holistic and programmatic approach, like what training programs are doing for determining competence of residents,33, 34 would be a better way to organize metrics—assembling multiple measures of different kinds of academic merit together into an organized framework would be the best way forward. Measuring our productivity is important. This, however, is just one tool in our box to measure one's contributions. Our hope is this work doesn't become the only tool, because much of academic productivity can't be quantified in this way. As one of the newest specialties in the house of medicine, we have the chance to look at things differently and avoid doing things the “way they have always been done.” We have always prided ourselves to be an innovative and nimble specialty; if we apply the same ingenuity to academics maybe we could lead change in the way we tackle academic scholarship. We look forward to future work using new and alternative metrics and measuring tools that consider the broader definition of academic productivity. Innovation and bravery will certainly prevent all of us from being Sneetches standing beneath the wrong lamppost.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,006 | 0,153 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,002 | 0,000 |
| Bibliométrie | 0,010 | 0,007 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,001 | 0,000 |
| Intégrité de la recherche | 0,004 | 0,014 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,001 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; les deux têtes enseignantes s’accordent sur ce qui est montré ici.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».