MétaCan
Menu
Back to cohort
Record W4220967648 · doi:10.1097/bpo.0000000000002136

Interobserver and Intraobserver Agreement are Unsatisfactory When Determining Abstract Study Design and Level of Evidence

2022· article· en· W4220967648 on OpenAlexaff
Neeraj M. Patel, Matthew R. Schmitz, Tracey P. Bastrom, Arvindera Ghag, Joseph A. Janicki, Indranil Kushare, Ronald W. Lewis, R. Justin Mistovich, Susan E. Nelson, Jeffrey R. Sawyer, Kelly L. Vanderhave, Maegen Wallace, Scott McKay

Bibliographic record

VenueJournal of Pediatric Orthopaedics · 2022
Typearticle
Languageen
FieldDecision Sciences
TopicMeta-analysis and systematic reviews
Canadian institutionsBC Children's Hospital
Fundersnot available
KeywordsMedicineMedical physicsNuclear medicine

Abstract

fetched live from OpenAlex

BACKGROUND: Understanding differences between types of study design (SD) and level of evidence (LOE) are important when selecting research for presentation or publication and determining its potential clinical impact. The purpose of this study was to evaluate interobserver and intraobserver reliability when assigning LOE and SD as well as quantify the impact of a commonly used reference aid on these assessments. METHODS: Thirty-six accepted abstracts from the Pediatric Orthopaedic Society of North America (POSNA) 2021 annual meeting were selected for this study. Thirteen reviewers from the POSNA Evidence-Based Practice Committee were asked to determine LOE and SD for each abstract, first without any assistance or resources. Four weeks later, abstracts were reviewed again with the guidance of the Journal of Bone and Joint Surgery (JBJS) LOE chart, which is adapted from the Oxford Centre for Evidence-Based Medicine. Interobserver and intraobserver reliability were calculated using Fleiss' kappa statistic (k). χ2 analysis was used to compare the rate of SD-LOE mismatch between the first and second round of reviews. RESULTS: Interobserver reliability for LOE improved slightly from fair (k=0.28) to moderate (k=0.43) with use of the JBJS chart. There was better agreement with increasing LOE, with the most frequent disagreement between levels 3 and 4. Interobserver reliability for SD was fair for both rounds 1 (k=0.29) and 2 (k=0.37). Similar to LOE, there was better agreement with stronger SD. Intraobserver reliability was widely variable for both LOE and SD (k=0.10 to 0.92 for both). When matching a selected SD to its associated LOE, the overall rate of correct concordance was 82% in round 1 and 92% in round 2 (P<0.001). CONCLUSION: Interobserver reliability for LOE and SD was fair to moderate at best, even among experienced reviewers. Use of the JBJS/Oxford chart mildly improved agreement on LOE and resulted in less SD-LOE mismatch, but did not affect agreement on SD. LEVEL OF EVIDENCE: Level II.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Direct model labels (unvalidated)

Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.

Model armCategoriesStudy designConfidence
gemmaMetaresearch
Domain: Evaluation · Genre: Empirical
About the Canadian research system: no · About a Canadian topic: no
Observationallow
gptMetaresearch
Domain: Methods · Genre: Empirical
About the Canadian research system: no · About a Canadian topic: no
Observationalhigh
models agreeAgreement compares identical category sets and study designs across arms.

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.101
metaresearch head score (Gemma)0.015
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Insufficient payload (model declined to judge)
Consensus categoriesMetaresearch
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.087
Threshold uncertainty score0.999

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.1010.015
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0020.001
Bibliometrics0.0010.001
Science and technology studies0.0000.000
Scholarly communication0.0000.001
Open science0.0010.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0020.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.902
GPT teacher head0.507
Teacher spread0.395 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Labeled directly by 2 models reading the full record.

Study designObservational
DomainEvaluation · Methods
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2022
Admission routes1
Has abstractyes

Explore more

Same venueJournal of Pediatric OrthopaedicsSame topicMeta-analysis and systematic reviewsCategoryMetaresearchFrench-language works237,207