Beyond the Impact Factor in Peer-Reviewed Literature: What Really Matters
Bibliographic record
Abstract
What is an “impact factor” and why does it matter so much today in the academic environment of medicine? Does it really matter in the clinical practice of medicine and specifically in our field of plastic surgery? In the modern scientific era, there have been several attempts to describe the quality of the bountiful literature. The impact factor is currently the most widely accepted method to accomplish this. The impact factor is defined as the ratio of cited works to published works for a given journal over a 2-year period.1,2 The advantage to this method lies in its simplicity—an easily calculable figure that approximates the influence of articles in a scientific journal. Developed by the Institute for Scientific Information and Eugene Garfield in the 1960s,1 the impact factor was designed to help librarians decide to which print journals they should subscribe, given the rapidly increasing volume of scientific literature. This information overload is even magnified today, as the number of journals in existence has grown dramatically over the last several decades, with more than 100 new journals in print each year.3 With time, the impact factor has developed into a surrogate for the “scientific prestige”4 of a journal. Because researchers want their results disseminated as often and as broadly as possible, they strive to publish in journals that boast high impact factors,2 because theoretically this would reach the widest readership. As such, many scientists receive promotion and funding based on the impact factors of journals in which they publish.3,5,6 In 2021, Plastic and Reconstructive Surgery (PRS) celebrated its highest impact factor to date: 4.730. This means that in 2020, on average, a PRS article was cited over 4.7 times. Though all plastic surgery journals generally are showing rising impact factors, PRS maintains its number 1 ranking in this field and is recognized as a premier surgical journal, with a position of 29 out of 212 surgery journals. Although this milestone is widely celebrated, the impact factor does not represent the full impact of any one journal. There are several limitations in the impact factor metrics. Tallying citations as the only method to assess the quality of a journal is narrow in perspective and does not represent the true merit of the work and effort of an author, article, or journal. Timeliness and rigor of review, author reputation, credentials of the editorial board, and total number of yearly citations to all content (not just the content in the impact factor 2-year period) are notably absent from the impact factor. The impact factor does not reflect accurately the total citations of a journal when compared to a smaller journal that publishes far less content and hence can have a similar or even higher impact factor when its total citations may be significantly less. Furthermore, it does not measure the modern way in which readers (i.e., physicians) read, learn, and share journal content by other modalities, such as video, social media, email, presentations, infographics, and so on, nor does it consider the specificity or generality of the subject matter covered or the size of the field represented; typically, the narrower the focus of a journal, the more niche its readership, and the more niche the readership, the fewer opportunities there are for citations. The end result is a diminished impact factor for niche journals with a restricted readership, despite the quality of their research.4 Foreign-language publications with robust research are excluded from English-language databases and therefore excluded from consideration in the impact factor, illustrating another weakness of this metric.2 The 2-year data collection period is confining because many impactful articles are cited well after 2 years. Poignant research take time to be influential, and thus may not be fully represented in the impact factor.3 Moreover, there are several ways the impact factor can be intentionally manipulated. Journals that publish a significant number of systematic reviews and meta-analysis articles will usually have more citations and hence generate a greater impact factor simply because these article types are cited more frequently.7 A journal may choose to publish fewer “citable” articles to decrease the denominator, which naturally will increase the impact factor.7 Self-citations or the unethical practice of editors asking authors to cite their own journals more heavily can lead to enhanced impact factors.3,8 There are penalties in the publishing world for intentional, proven publication fraud; however, the permissible ways in which the impact factor can be “gamed” still influence scores and rankings. Because of the various effective search engines and the efficiency of information dissemination, impact factor no longer is the best metric to express the influence of a particular journal.2,7 Many groups have developed and supported alternative journal ranking tools, including Eigenfactor, SCImago Journal Rank, h-index, and more. Other societies and journals signed the San Francisco Declaration on Research Assessment9 to avoid journal-based metrics, largely the impact factor, “as a surrogate measure of the quality of individual research articles, to assess an individual scientist’s contributions, or in hiring, promotion, or funding decisions.” More than 20,000 individuals and organizations in almost 150 countries have signed the declaration as of this writing.9 The internet has revolutionized the flow of information. Researchers are taking advantage of new platforms, such as streaming videos (including YouTube), blog posts, community forums, open access journals, and social media. Is there another way to measure true impact than a 2-year average of citations in scientific databases? The answer is “absolutely yes”; a new impact factor must emerge to truly measure the information expansion in our medical literature today beyond the written word. The infinite amount of information at our fingertips highlights the shortcomings of citation-based bibliometrics when determining the influence of literature, and so Altmetrics were born. The Altmetric Explorer is a tool that scours the internet to provide real-time data on research output through analysis of a comprehensive list of domains, including citations but also public policy documents, blogs, social media, patents, mainstream media outlets, the Open Syllabus Project, and Wikipedia.10 The tool uses a weighted algorithm to assign value to the media through which scientific literature is shared. For example, a news outlet delivering a story on breast implant–associated illness will contribute more to the Altmetric score than a patient posting on Facebook.11 Twitter has become an important outlet for the academic community to share research, ideas, and feedback, which is now credited in Altmetric.12 These data are depicted as a vibrant donut (Fig. 1), with each color representing a different medium. The relative amount of color in the donut will change depending on the sources from which a research output has received attention. Also included in the analysis is the Altmetric Attention Score (Fig. 1). This is a weighted approximation based on the volume, sources, and authors of the article. Not coincidentally, these calculations incorporate systems to overcome the deficiencies of the impact factor. For instance, if a Tweet is shared multiple times, only the first share is counted toward the overall volume. In addition, Almetrics consider who is sharing the articles, so reputable authors contributing to an academic discussion have a greater influence on the score. Conversely, authors who continue to share articles with extreme bias, poor science, or from an irreputable journal will be flagged and their attention score adjusted accordingly.10 If a true goal of research is the dissemination of information to the populous, the Altmetric score must be part of the conversation. However, social media are rapidly changing and are prone to reporting biases based on ever-changing algorithms on those platforms, which can shift without notice. Just as with the impact factor, the Altmetric score has its own set of limitations. For example, it was created in 2012 and does not include some of today’s most important social media platforms that gained popularity or were created after its founding. That means that Altmetrics do not factor in the traffic on Instagram, Snapchat, or TikTok and still list now-defunct social media platforms (such as Google+) as part of their recipe.13 According to the Altmetric support page, the weight of certain platforms can vary, therefore adding to recent scrutiny and confusion.14,15 Finally, though the Altmetric score is a valuable tool on the article-based level, it is hard to aggregate a predictive or holistic measure for a journal based on this measure.Fig. 1.: Altmetric donut and attention score for Furnas HJ, Garza RM, Li AY, et al. Gender differences in the professional and personal lives of plastic surgeons. Plast Reconstr Surg. 2018;142:252–264.Although the impact factor and Altmetric score can each be useful and interesting, they should not be stand-alone metrics to assess the quality of a journal or the science within. Therefore, all of us in the orbit of scholarly publishing must assess these metrics thoughtfully to seek and to embrace a holistic view of the true influence of a peer-reviewed journal and the articles it publishes. We know many additional factors determine the impact of our Journal, including digital media, social media, quality and efficiency of peer review, engagement initiatives, readership numbers, and reach. As our specialty celebrates innovation, Plastic and Reconstructive Surgery strives to promote new technology and methods for information exchange. From its origin in 1946 and establishment as the leader in plastic surgery learning and innovation for over 75 years, PRS is unique in the peer-reviewed medical journal world with its rich, global multimedia resources, including video content, podcasts, evidence-based medicine, topic-based reading, written and video discussions, and social media conversations. The impact and influence of social media were recognized early in PRS, and the Journal embraced its role to deliver information not only to the entire specialty but also to interested members of the public as well. Plastic and Reconstructive Surgery prides itself on its rigorous peer reviews, quality of the editorial board, stringent acceptance rate, timeliness of publication, and reputation of the authors.2 When one reads the Journal, one is assured that articles published represent the most influential research in the field, covering both the aesthetic and reconstructive arena that formed the foundation of plastic surgery. With the widespread dissemination of scientific literature outside traditional methods, social media–based and video-based academic platforms will be even more prevalent and meaningful. Tools such as Altmetrics, and their successors, will play a vital role in measuring impact. Despite important dissent from the San Francisco Declaration on Research Assessment initiative and those who signed it, there is still an overwhelming importance on tracking citations of peer-reviewed publications in the form of the impact factor. Despite the limitations of the current, accepted systems for tracking metrics, they seem to be here to stay. It is the strong opinion of the leadership of Plastic and Reconstructive Surgery and the American Society of Plastic Surgeons that the bodies controlling popular metrics receive feedback and move toward full transparency of their algorithm, changes, and limitations. More importantly, we believe that those who read, write, and review for Plastic and Reconstructive Surgery, as well as those who determine career advancement for physicians, recognize the strengths and weaknesses of each tool used to determine a journal’s worth and, as such, consider many broad and different factors that represent a journal’s true impact.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.066 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.009 |
| Insufficient payload (model declined to judge) | 0.008 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".