Integrative Approach to Quality Assessment of Medical Journals Using Impact Factor, Eigenfactor, and Article Influence Scores
Bibliographic record
Abstract
BACKGROUND: Impact factor (IF) is a commonly used surrogate for assessing the scientific quality of journals and articles. There is growing discontent in the medical community with the use of this quality assessment tool because of its many inherent limitations. To help address such concerns, Eigenfactor (ES) and Article Influence scores (AIS) have been devised to assess scientific impact of journals. The principal aim was to compare the temporal trends in IF, ES, and AIS on the rank order of leading medical journals over time. METHODS: The 2001 to 2008 IF, ES, AIS, and number of citable items (CI) of 35 leading medical journals were collected from the Institute of Scientific Information (ISI) and the http://www.eigenfactor.org databases. The journals were ranked based on the published 2008 ES, AIS, and IF scores. Temporal score trends and variations were analyzed. RESULTS: In general, the AIS and IF values provided similar rank orders. Using ES values resulted in large changes in the rank orders with higher ranking being assigned to journals that publish a large volume of articles. Since 2001, the IF and AIS of most journals increased significantly; however the ES increased in only 51% of the journals in the analysis. Conversely, 26% of journals experienced a downward trend in their ES, while the rest experienced no significant changes (23%). This discordance between temporal trends in IF and ES was largely driven by temporal changes in the number of CI published by the journals. CONCLUSION: The rank order of medical journals changes depending on whether IF, AIS or ES is used. All of these metrics are sensitive to the number of citable items published by journals. Consumers should thus consider all of these metrics rather than just IF alone in assessing the influence and importance of medical journals in their respective disciplines.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.020 | 0.063 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.012 | 0.040 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".