MétaCan
Menu
Back to cohort
Record W3014896206 · doi:10.1002/sim.8532

STRATOS guidance document on measurement error and misclassification of variables in observational epidemiology: Part 1—Basic theory and simple methods of adjustment

2020· review· en· W3014896206 on OpenAlexafffund
Ruth H. Keogh, Pamela A. Shaw, Paul Gustafson, Raymond J. Carroll, Veronika Deffner, Kevin W. Dodd, Helmut Küchenhoff, Janet A. Tooze, Michael P. Wallace, Victor Kipnis, Laurence S. Freedman

Bibliographic record

VenueStatistics in Medicine · 2020
Typereview
Languageen
FieldMedicine
TopicNutritional Studies and Diet
Canadian institutionsUniversity of WaterlooUniversity of British Columbia
FundersNational Institutes of HealthNational Cancer InstituteMedical Research CouncilNational Institute of Allergy and Infectious DiseasesNatural Sciences and Engineering Research Council of CanadaPatient-Centered Outcomes Research Institute
KeywordsCovariateObservational errorStatisticsComputer scienceObservational studyErrors-in-variables modelsEconometricsType I and type II errorsRegressionRegression analysisSample size determinationExtrapolationCalibrationLinear regressionData miningMathematics

Abstract

fetched live from OpenAlex

Measurement error and misclassification of variables frequently occur in epidemiology and involve variables important to public health. Their presence can impact strongly on results of statistical analyses involving such variables. However, investigators commonly fail to pay attention to biases resulting from such mismeasurement. We provide, in two parts, an overview of the types of error that occur, their impacts on analytic results, and statistical methods to mitigate the biases that they cause. In this first part, we review different types of measurement error and misclassification, emphasizing the classical, linear, and Berkson models, and on the concepts of nondifferential and differential error. We describe the impacts of these types of error in covariates and in outcome variables on various analyses, including estimation and testing in regression models and estimating distributions. We outline types of ancillary studies required to provide information about such errors and discuss the implications of covariate measurement error for study design. Methods for ascertaining sample size requirements are outlined, both for ancillary studies designed to provide information about measurement error and for main studies where the exposure of interest is measured with error. We describe two of the simpler methods, regression calibration and simulation extrapolation (SIMEX), that adjust for bias in regression coefficients caused by measurement error in continuous covariates, and illustrate their use through examples drawn from the Observing Protein and Energy (OPEN) dietary validation study. Finally, we review software available for implementing these methods. The second part of the article deals with more advanced topics.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.041
metaresearch head score (Gemma)0.124
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.959
Threshold uncertainty score0.217

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0410.124
Meta-epidemiology (narrow)0.0040.003
Meta-epidemiology (broad)0.0030.004
Bibliometrics0.0080.007
Science and technology studies0.0020.003
Scholarly communication0.0050.004
Open science0.0080.005
Research integrity0.0160.011
Insufficient payload (model declined to judge)0.0330.029

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.332
GPT teacher head0.491
Teacher spread0.159 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designNot applicable
DomainMethods
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations176
Published2020
Admission routes2
Has abstractyes

Explore more

Same venueStatistics in MedicineSame topicNutritional Studies and DietFrench-language works237,207