MétaCan
Menu
Back to cohort
Record W2157427387 · doi:10.1007/s11136-010-9606-8

The COSMIN checklist for assessing the methodological quality of studies on measurement properties of health status measurement instruments: an international Delphi study

2010· article· en· W2157427387 on OpenAlexaff
Lidwine B. Mokkink, Caroline B. Terwee, Donald L. Patrick, Jordi Alonso, Paul W. Stratford, Dirk L. Knol, L.M. Bouter, Henrica C. W. de Vet

Bibliographic record

VenueQuality of Life Research · 2010
Typearticle
Languageen
FieldSocial Sciences
TopicHealth Education and Validation
Canadian institutionsMcMaster University
FundersVrije Universiteit Amsterdam
KeywordsChecklistDelphi methodCriterion validityConstruct validityReliability (semiconductor)Content validityApplied psychologyPsychologyFace validityQuality (philosophy)MedicinePsychometricsClinical psychologyStatisticsMathematics

Abstract

fetched live from OpenAlex

BACKGROUND: Aim of the COSMIN study (COnsensus-based Standards for the selection of health status Measurement INstruments) was to develop a consensus-based checklist to evaluate the methodological quality of studies on measurement properties. We present the COSMIN checklist and the agreement of the panel on the items of the checklist. METHODS: A four-round Delphi study was performed with international experts (psychologists, epidemiologists, statisticians and clinicians). Of the 91 invited experts, 57 agreed to participate (63%). Panel members were asked to rate their (dis)agreement with each proposal on a five-point scale. Consensus was considered to be reached when at least 67% of the panel members indicated 'agree' or 'strongly agree'. RESULTS: Consensus was reached on the inclusion of the following measurement properties: internal consistency, reliability, measurement error, content validity (including face validity), construct validity (including structural validity, hypotheses testing and cross-cultural validity), criterion validity, responsiveness, and interpretability. The latter was not considered a measurement property. The panel also reached consensus on how these properties should be assessed. CONCLUSIONS: The resulting COSMIN checklist could be useful when selecting a measurement instrument, peer-reviewing a manuscript, designing or reporting a study on measurement properties, or for educational purposes.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.597
metaresearch head score (Gemma)0.541
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: Methods · Consensus signal: Methods
Study designCandidate signal: Qualitative · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.403
Threshold uncertainty score0.497

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.5970.541
Meta-epidemiology (narrow)0.0030.002
Meta-epidemiology (broad)0.0040.006
Bibliometrics0.0200.008
Science and technology studies0.0050.007
Scholarly communication0.0050.005
Open science0.0060.015
Research integrity0.0040.005
Insufficient payload (model declined to judge)0.0040.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.969
GPT teacher head0.723
Teacher spread0.246 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designQualitative
DomainMethods
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations4,208
Published2010
Admission routes1
Has abstractyes

Explore more

Same venueQuality of Life ResearchSame topicHealth Education and ValidationFrench-language works237,207