MétaCan
Menu
Back to cohort
Record W2914474517 · doi:10.1111/acem.13707

On Lampposts, Sneetches, and Stars: A Call to Go Beyond Bibliometrics for Determining Academic Value

2019· letter· en· W2914474517 on OpenAlexaff
Teresa M. Chan, Damon Kuehl

Bibliographic record

VenueAcademic Emergency Medicine · 2019
Typeletter
Languageen
FieldMedicine
TopicHealth and Medical Research Impacts
Canadian institutionsMcMaster University
Fundersnot available
KeywordsPopularityScholarshipImpact factorBibliometricsProductivityAltmetricsAcademic institutionPromotion (chess)Metric (unit)Value (mathematics)MedicineScholarly communicationPublic relationsComputer scienceData sciencePsychologyMarketingPublishingLibrary sciencePolitical scienceLawSocial psychology

Abstract

fetched live from OpenAlex

Academic scholarship is something that we have always sought to quantify. Journal impact factor, CiteScore, h-index, hi-10, and the alternative metrics (Altmetrics) are terms that are familiar to seasoned academics.1 Junior scholars are often advise to use bibliometric and publication metrics to decide where to send their work.1, 2 Whether we are talking about journal-level, article-level, or person-level metrics, these are all human constructs that have arisen as ways to quantify the reach and impact of our academic work. In the paper by Boudreaux and colleagues3 in this issue of Academic Emergency Medicine, we are introduced to a new way to benchmark scholarly productivity for faculty members. Using a norm-based methodology (e.g., publications over a certain amount of time) and impact (e.g., citations that each publication gets within that group) to compare an individuals’ scholarly output against that of his or her peers. We congratulate the authors on taking on this difficult problem. So often, external letters for promotion state “Based on their productivity, I believe this person would be promoted at my institution” are at best an estimate and likely a guess. Similarly, tools that go beyond popularity contests and provide metric-based benchmarks for an individual (or even a departments productivity) are welcome. But, in this editorial we ask: Are the metrics we often use the right ones? Are they equitable for individuals across disparate systems? What are the ramifications to our academic culture when we choose to use certain metrics over others? Might we create perverse incentivizes and make ourselves vulnerable to predators? Boyer et al. articulated how we should reconsider scholarship in modern academia in the 1990s by proposing four broad categories of scholarship: 1) The scholarship of discovery (i.e., what we may see as “bench” or epidemiologic research), 2) the scholarship of application (i.e., clinical research or implementation science), 3) the scholarship of integration (i.e., where we do interdisciplinary work, or “remix” prior work via knowledge syntheses), and 4) the scholarship of teaching (i.e., which has resulted in much of the educational research and innovation that we see in medical education).4, 5 Boyer was followed by Glassick,6 as well as Hutchings and Shulman,7 who helped to formalize standards for assessing scholarship—and more specifically the scholarship of teaching, which was deemed to be particularly troublesome to define. Box 1 lists Glassick's proposed list of standards for judging scholarship and additional criteria for the scholarship of teaching by Hutchings and Shulman.7 For all forms of scholarship: Clear Goals Adequate Preparation Appropriate Methods Significant (i.e., Important) Results Effective Presentation Reflective Critique Specifically for the scholarship of teaching: Must be made public Must be available for peer review & critique according to accepted standards Must be able to be reproduced and built upon by other scholars. When talking about metrics, it brings to mind the fable about a man and a lamppost.8 The story goes that the man is on his hands and knees under a lamppost when a bystander walks by and asks him if he needs help. The man looks up and explains that he is looking for his keys. The bystander steps in to help, and after searching for a while asks: “Are you sure you lost the keys around here?” He looks up and states: “Actually, I think I dropped them at the end of the block, but it's too dark there, the light is much better over here …” This story is a cautionary tale to those who use convenient measures, rather than the measures that may be more relevant but harder to measure (e.g., using hemoglobin A1c instead of development of diabetic neuropathy; board exams scores as a surrogate marker for clinical competence, rather than ability to resuscitate a patient in the workplace).8 Convenience measures may incentivize the wrong behaviors and keep you searching under the wrong lamppost. Academia has largely embraced a quantitative approach to measuring productivity via bibliometrics (see Table 1 for partial listing). We have moved from simply counting work (e.g., pure numbers of publications, citations) to incorporating increasing levels of sophistication to determine the impact of his or her works. Newer metrics such as the one proposed by Boudreaux et al. combine metrics in new ways. Other examples of these are more advanced concepts such as the h-index1 (a person-level impact factor) and alternative metrics (publication metrics that harness social media and mainstream media uptake to quantify impact with end-users).9-11 Similarly, normalized metrics proposed by Boudreaux and colleagues may be convenient as surrogates for the quality of an academic's scholarship, rather than a true indicator of quality. CiteScore (journal-level metric) The “citability” of an author's most important works.1 In emergency medicine, this has been proposed to help judge academic performance and identify individuals of high scholarly potential.12, 13 Can be used as a measure of scholarly discussion around a given article via these online platforms.9-11 Has been correlated with later citation metrics in some fields.14 Quantitative metrics such as journal impact factor as a surrogate for quality, but it is well known that journal impact factors do not always correlate with the quality of the work contained within the journal.1, 15 While it is a reasonable inference that there must be some correlation between a paper's worth and its citations, it's important to realize this may not always be true. Controversial papers may be highly cited. And while publication and peer review process is supposed to be another check and balance to judge the quality of one's work, merely being published may also not be a high enough bar to judge high-quality work.15-17 One only need to look at the debacle over the Andrew Wakefield incident (which, per Google Scholar has been cited 31 times to date) and realize that even high-impact journals with expert editors and reviewers can be led astray.18 It should be no surprise that there are some championing the cause of evaluating scholarship via quality metrics—all of which can assist individuals to understand what makes high-quality learning materials. In the age of JAMA User's Guides19 and reporting guidelines some groups are lobbying for the development of quality assessment tools to guide individuals to determine the quality of individual works of scholarship. Interestingly, a recent flurry of work in this area has been done by those groups that are engaging in online or digital scholarship,20-23 likely because disruptive and ubiquitous resources have under the most scrutiny, resulting in the responsive quality assessment agenda.24-26 Quality measures exist such as ratings of papers within certain journals. One new, PubMed-indexed journal (Cureus.com) has the Scholarly Impact Quotient rating, which asks individual readers to rate articles based on 10-point Likert scales for six criteria: 1) clarity of background and rationale, 2) clinical importance, 3) study design and methods, 4) data analysis, 5) novelty of conclusions, and 6) quality of presentation.27 There is no doubt that metrics and quality scores abound. But at the end of the day, when we are examining the best practices for determining scholarly work, quantifying publications/citations move us away from appreciating alternative scholarship, especially in medical education and quality improvement. Theodor Seuss Geisel (a.k.a. Dr. Seuss) warned us about the ramifications of valuing certain types of phenotypic markers over others.28 In his book The Sneetches and Other Stories, he tells the cautionary tale of an opportunistic charleton (Sylvester McMonkey McBean) who sees that a group of individuals (Plain-Belly Sneetches) are desirous of a certain phenotypic marker (i.e., a star on their belly) and sells them this … at a price.28 Although a children's tale, this is very much mirrored in our current academic era where the quantification of performance metrics are en vogue. We understand the scope of this paper is measuring publications and very much agree with the authors comment in the results where they state that “… [m]any academicians may not be expected to perform research or publish papers, so including them in the database may not be appropriate”3 for the very important fact we must be cautious not to singularly use metrics that favor one type of scholarly work over others. The substantial expansion of in EM training programs into community-based health systems without research infrastructure creates a divide, effectively marking some of our centers a group with “stars,” instead of celebrating new academics for their talents. For instance, individuals toiling away setting up a new residency programs or those working in innovative frontiers may spend valuable time involved in curriculum design or administration and may not register initially on the proposed norm-based bibliometric scale. We have moved the needle on academic productivity in the areas of discovery and possibly application, but for the areas like integrative innovation or educational delivery an over emphasis on established metrics may moves us away from fostering successes. The launch of our sister journal AEM Education and Training29 or the indexing of the AAMC's MedEdPortal.org may hold promise in creating avenues for encouraging new types of scholarship. Indeed, with dedicated attention, we know it is possible to encourage educational scholars toward achieving comparable publication-based metrics,30 but this may still bias us away from important work such as curricular development, administration, or journalism. In addition, opportunistic companies abound knowing this need for publishing metrics. We must be wary that incentivizing pure numbers have given to the rise of predatory journals.31, 32 Inexperienced academics are easy prey for these notorious entities (who often write e-mails loaded with complimentary language such as: “Dear esteemed and honorific professor”) and the pressures can be so unbearable that some are more than happy to provide some “stars upon thars for just a few dollars eaches.” As academic leaders, if we persist in merely creating new mathematical formulae to combine numbers that are most easy to attain, are we not merely asking all our faculty to spend their time chasing surrogate measures that we hope represent quality? And by not measuring other important and valuable work in a meaningful ways, will this not encourage a drift away from what matters toward what counts? Ultimately, as a specialty, we should and must decide on what is important to us—and then pursue ways to quantify and qualify these things effectively. Table 2 lists the present metrics that we often use for promotions processes. Academicians are incentivized to publish more papers. May contribute to “publish or perish” or undue stressors. May encourage poor academic practices (e.g., “salami slicing”).35 May result in inexperienced scholars being lured into publishing within predatory journals. Academicians may perceive that only work which is cited is worthy. May lead to a devaluing of high quality work that May encourage junior faculty to pursue forms of scholarship that are more likely to be cited (e.g., guideline work or systematic reviews/meta-analyses) Academicians may view that certain journals are more important than others. May lead to devaluing of high-quality work which is published in a less prominent journal, but eventually is discovered by modern search technologies (e.g., Google Scholar) and is highly cited and read. Academicians may continually chase grants, instead of actually performing the work. May lead to a partitioning of academicians who exclusively write grants but may not always participate fully in the fulfillment or authorship of the works in a personally fulfilling way. May lead to academicians who “chase” trendy topics to research, swayed by granting agencies and politics, rather than choosing topics of importance to our field. Academicians may seek to supervise more trainees than they can handle. May lead to academic abuse and poor quality mentorship. Academicians may seek opportunities more aligned with external measures (e.g., national or international organizations) rather than contribute to their local academic institutions or environments. May lead to nationally or internationally renowned academicians that do not contribute to your local environment. We believe that while the present paper in our journal represents a start of a norm-based movement for bibliometric-based metrics, there is much work ahead of us in this field before we can adequately quantify the true productivity of our field. While for some academics (e.g., clinical or health services researchers) these metrics may resonate, they do not encapsulate all the metrics that reflect all types of high-value scholarship that comes from our specialty. We propose that those interested in these areas be inspired by Boudreaux et al. to better quantify and qualify the academic work that our specialty has to offer. Table 3 proposes some value-driven outcomes that could be integrated into academic merit or promotion systems to foster a more holistic academic environment for a diverse set of scholars. We hope that by listing some of these other valuable endeavors and the possible metrics that we may encourage others to quantify some of these other metrics for best comparison within emergency medicine. Ultimately, perhaps a more holistic and programmatic approach, like what training programs are doing for determining competence of residents,33, 34 would be a better way to organize metrics—assembling multiple measures of different kinds of academic merit together into an organized framework would be the best way forward. Measuring our productivity is important. This, however, is just one tool in our box to measure one's contributions. Our hope is this work doesn't become the only tool, because much of academic productivity can't be quantified in this way. As one of the newest specialties in the house of medicine, we have the chance to look at things differently and avoid doing things the “way they have always been done.” We have always prided ourselves to be an innovative and nimble specialty; if we apply the same ingenuity to academics maybe we could lead change in the way we tackle academic scholarship. We look forward to future work using new and alternative metrics and measuring tools that consider the broader definition of academic productivity. Innovation and bravery will certainly prevent all of us from being Sneetches standing beneath the wrong lamppost.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.006
metaresearch head score (Gemma)0.153
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Meta-epidemiology (narrow), Research integrity
Consensus categoriesResearch integrity
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Commentary · Consensus signal: Commentary
Teacher disagreement score0.147
Threshold uncertainty score0.999

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0060.153
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0020.000
Bibliometrics0.0100.007
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0010.000
Research integrity0.0040.014
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.224
GPT teacher head0.478
Teacher spread0.254 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designNot applicable
Domainnot available
GenreCommentary

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations32
Published2019
Admission routes1
Has abstractyes

Explore more

Same venueAcademic Emergency MedicineSame topicHealth and Medical Research ImpactsFrench-language works237,207