Second-Order Peer Review of the Medical Literature for Clinical Practitioners
Bibliographic record
Abstract
CONTEXT: Most articles in clinical journals are not appropriate for direct application by individual clinicians. OBJECTIVE: To create a second order of clinical peer review for journal articles to determine which articles are most relevant for specific clinical disciplines. DESIGN AND SETTING: A 2-stage prospective observational study in which research staff reviewed all issues of over 110 (number has varied slightly as new journals were added or discarded from review but number has always been over 110) clinical journals and selected each article that met critical appraisal criteria from January 2003 through the present. Practicing physicians were recruited from around the world, excluding Northern Ontario, to the McMaster Online Rating of Evidence (MORE) system and registered as raters according to their clinical disciplines. An automated system assigned each qualifying article to raters for each pertinent clinical discipline, and recorded their online assessments of the articles on 7-point scales (highest score, 7) of relevance and newsworthiness (defined as useful new information for physicians). Rated articles fed an online alerting service, the McMaster Premium Literature Service (PLUS). Physicians from Northern Ontario were invited to register with PLUS and then receive e-mail alerts about articles according to MORE system peer ratings for their own discipline. Online access by PLUS users of PLUS alerts, raters' comments, article abstracts, and full-text journal articles was automatically recorded. MAIN OUTCOME MEASURES: Clinical rater recruitment and performance. Relevance and newsworthiness of journal articles to clinical practice in the discipline of the rating physician. RESULTS: Through October 2005, MORE had 2139 clinical raters, and PLUS had 5892 articles with 45 462 relevance ratings and 44 724 newsworthiness ratings collected since 2003. On average, clinicians rated systematic review articles higher for relevance to practice than articles with original evidence and lower for useful new information. Primary care physicians rated articles lower than did specialists (P<.05). Of the 98 physicians who registered for PLUS, 88 (90%) used it on 3136 occasions during an 18-month test period. CONCLUSIONS: This demonstration project shows the feasibility and use of a post-publication clinical peer review system that differentiates published journal articles according to the interests of a broad range of clinical disciplines.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.182 | 0.297 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.029 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".