MétaCan
Menu
Back to cohort
Record W4220967648 · doi:10.1097/bpo.0000000000002136

Interobserver and Intraobserver Agreement are Unsatisfactory When Determining Abstract Study Design and Level of Evidence

2022· article· en· W4220967648 on OpenAlexaff
Neeraj M. Patel, Matthew R. Schmitz, Tracey P. Bastrom, Arvindera Ghag, Joseph A. Janicki, Indranil Kushare, Ronald W. Lewis, R. Justin Mistovich, Susan E. Nelson, Jeffrey R. Sawyer, Kelly L. Vanderhave, Maegen Wallace, Scott McKay

Bibliographic record

VenueJournal of Pediatric Orthopaedics · 2022
Typearticle
Languageen
FieldDecision Sciences
TopicMeta-analysis and systematic reviews
Canadian institutionsBC Children's Hospital
Fundersnot available
KeywordsMedicineMedical physicsNuclear medicine

Abstract

fetched live from OpenAlex

BACKGROUND: Understanding differences between types of study design (SD) and level of evidence (LOE) are important when selecting research for presentation or publication and determining its potential clinical impact. The purpose of this study was to evaluate interobserver and intraobserver reliability when assigning LOE and SD as well as quantify the impact of a commonly used reference aid on these assessments. METHODS: Thirty-six accepted abstracts from the Pediatric Orthopaedic Society of North America (POSNA) 2021 annual meeting were selected for this study. Thirteen reviewers from the POSNA Evidence-Based Practice Committee were asked to determine LOE and SD for each abstract, first without any assistance or resources. Four weeks later, abstracts were reviewed again with the guidance of the Journal of Bone and Joint Surgery (JBJS) LOE chart, which is adapted from the Oxford Centre for Evidence-Based Medicine. Interobserver and intraobserver reliability were calculated using Fleiss' kappa statistic (k). χ2 analysis was used to compare the rate of SD-LOE mismatch between the first and second round of reviews. RESULTS: Interobserver reliability for LOE improved slightly from fair (k=0.28) to moderate (k=0.43) with use of the JBJS chart. There was better agreement with increasing LOE, with the most frequent disagreement between levels 3 and 4. Interobserver reliability for SD was fair for both rounds 1 (k=0.29) and 2 (k=0.37). Similar to LOE, there was better agreement with stronger SD. Intraobserver reliability was widely variable for both LOE and SD (k=0.10 to 0.92 for both). When matching a selected SD to its associated LOE, the overall rate of correct concordance was 82% in round 1 and 92% in round 2 (P<0.001). CONCLUSION: Interobserver reliability for LOE and SD was fair to moderate at best, even among experienced reviewers. Use of the JBJS/Oxford chart mildly improved agreement on LOE and resulted in less SD-LOE mismatch, but did not affect agreement on SD. LEVEL OF EVIDENCE: Level II.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Direct model labels (unvalidated)

Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.

Model armCategoriesStudy designConfidence
gemmaMetaresearch
Domain: Evaluation · Genre: Empirical
About the Canadian research system: no · About a Canadian topic: no
Observationallow
gptMetaresearch
Domain: Methods · Genre: Empirical
About the Canadian research system: no · About a Canadian topic: no
Observationalhigh
models agreeAgreement compares identical category sets and study designs across arms.

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.589
metaresearch head score (Gemma)0.774
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: Methods · Consensus signal: Methods
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.411
Threshold uncertainty score0.507

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.5890.774
Meta-epidemiology (narrow)0.0020.002
Meta-epidemiology (broad)0.0040.004
Bibliometrics0.0160.008
Science and technology studies0.0040.007
Scholarly communication0.0060.005
Open science0.0040.006
Research integrity0.0030.003
Insufficient payload (model declined to judge)0.0030.002

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.902
GPT teacher head0.507
Teacher spread0.395 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Labeled directly by 2 models reading the full record.

Study designObservational
DomainEvaluation · Methods
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2022
Admission routes1
Has abstractyes

Explore more

Same venueJournal of Pediatric OrthopaedicsSame topicMeta-analysis and systematic reviewsCategoryMetaresearchFrench-language works237,207