Evaluating Journal Impact Factor: a systematic survey of the pros and cons, and overview of alternative measures
Bibliographic record
Abstract
Background: Journal Impact Factor (JIF) has several intrinsic flaws, which highlight its inability to adequately measure citation distributions or indicate journal quality. Despite these flaws, JIF is still widely used within the academic community, resulting in the propagation of potentially misleading information. A critical review of the usefulness of JIF is needed including an overview of the literature to identify viable alternative metrics. The objectives of this study are: (1) to assess the usefulness of JIF by compiling and comparing its advantages and disadvantages; (2) to record the differential uses of JIF within research environments; and (3) to summarize and compare viable alternative measures to JIF. Methods: Three separate literature search strategies using MEDLINE and Web of Science were completed to address the three study objectives. Each search was completed in accordance with PRISMA guidelines. Results were compiled in tabular format and analyzed based on reporting frequency. Results: For objective (1), 84 studies were included in qualitative analysis. It was found that the recorded advantages of JIF were outweighed by disadvantages (18 disadvantages vs. 9 advantages). For objective (2), 653 records were included in a qualitative analysis. JIF was found to be most commonly used in journal ranking (n = 653, 100%) and calculation of scientific research productivity (n = 367, 56.2%). For objective (3), 65 works were included in qualitative analysis. These articles revealed 45 alternatives, which includes 18 alternatives that improve on highly reported disadvantages of JIF. Conclusion: JIF has many disadvantages and is applied beyond its original intent, leading to inaccurate information. Several metrics have been identified to improve on certain disadvantages of JIF. Integrated Impact Indicator (I3) shows great promise as an alternative to JIF. However, further scientometric analysis is needed to assess its properties.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.028 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.002 | 0.006 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".