MétaCan
Menu
Back to cohort
Record W1487942018

Developing and Testing a Tool for the Classification of Study Designs in Systematic Reviews of Interventions and Exposures

2010· article· en· W1487942018 on OpenAlexaff
Lisa Hartling, Kenneth Bond, Krystal Harvey, Santaguida Pl, Meera Viswanathan, Dryden Dm

Bibliographic record

VenueJournal of Clinical Epidemiology · 2010
Typearticle
Languageen
FieldDecision Sciences
TopicMeta-analysis and systematic reviews
Canadian institutionsUniversity of Alberta
Fundersnot available
KeywordsInter-rater reliabilityReliability (semiconductor)Psychological interventionSystematic reviewComputer scienceClinical study designTest (biology)MedicineKappaClassification schemeMedical physicsMEDLINEData miningMachine learningStatisticsClinical trialMathematicsRating scalePathologyNursing
DOInot available

Abstract

fetched live from OpenAlex

Background Classification of study design can help provide a common language for researchers. Within a systematic review, definition of specific study designs can help guide inclusion, assess the risk of bias, pool studies, interpret results, and grade the body of evidence. However, recent research demonstrated poor reliability for an existing classification scheme. Objectives To review tools used to classify study designs; to select a tool for evaluation; to develop instructions for application of the tool to intervention/exposure studies; and to test the tool for accuracy and interrater reliability. Methods We contacted representatives from all AHRQ Evidence-based Practice Centers (EPCs), other relevant organizations, and experts in the field to identify tools used to classify study designs. Twenty-three tools were identified; 10 were relevant to our objectives. The Steering Committee ranked the 10 tools using predefined criteria. The highest-ranked tool was a design algorithm for studies of health care interventions developed, but no longer advocated, by the Cochrane Non-Randomised Studies Methods Group. This tool was used as the basis for our classification tool and was revised to encompass more study designs and to incorporate elements of other tools. A sample of 30 studies was used to test the tool. Three members of the Steering Committee developed a reference standard (i.e., the “true” classification for each study); 6 testers applied the revised tool to the studies. Interrater reliability was measured using Fleiss’ kappa (κ) and accuracy of the testers’ classification was assessed against the reference standard. Based on feedback from the testers and the reference standard committee, the tool was further revised and tested by another 6 testers using 15 studies randomly selected from the original sample. Results In the first round of testing the inter-rater reliability was fair among the testers (κ = 0.26) and the reference standard committee (κ = 0.33). Disagreements occurred at all decision points in the algorithm; revisions were made based on the feedback. The second round of testing showed improved interrater reliability (κ = 0.45, moderate agreement) with improved, but still low, accuracy. The most common disagreements were whether the study was “experimental” (5/15 studies) and whether there was a comparison (4/15 studies). In both rounds of testing, the level of agreement for testers who had completed graduate-level training was higher than for testers who had not completed training. Conclusion Potential reasons for the observed low reliability and accuracy include the lack of clarity and comprehensiveness of the tool, inadequate reporting of the studies, and variability in user characteristics. Application of a tool to classify study designs in the context of a systematic review should be accompanied by adequate training, pilot testing, and documented decision rules.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.814
metaresearch head score (Gemma)0.940
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: Methods · Consensus signal: Methods
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.186
Threshold uncertainty score0.234

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.8140.940
Meta-epidemiology (narrow)0.0100.009
Meta-epidemiology (broad)0.0160.033
Bibliometrics0.0760.053
Science and technology studies0.0090.013
Scholarly communication0.0230.029
Open science0.0110.024
Research integrity0.0130.015
Insufficient payload (model declined to judge)0.0070.003

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.983
GPT teacher head0.730
Teacher spread0.253 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designNot applicable
DomainMethods
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations34
Published2010
Admission routes1
Has abstractyes

Explore more

Same venueJournal of Clinical EpidemiologySame topicMeta-analysis and systematic reviewsFrench-language works237,207