MétaCan
Menu
Back to cohort
Record W2946571535 · doi:10.1186/s12874-019-0737-5

The utility of multivariate outlier detection techniques for data quality evaluation in large studies: an application within the ONDRI project

2019· article· en· W2946571535 on OpenAlexafffundabout
Kelly M. Sunderland, Derek Beaton, Julia Fraser, Donna Kwan, Paula McLaughlin, Manuel Montero‐Odasso, Alicia Peltsch, Frederico Pieruccini‐Faria, Demetrios J. Sahlas, Richard H. Swartz, Stephen C. Strother, Malcolm A. Binns

Bibliographic record

VenueBMC Medical Research Methodology · 2019
Typearticle
Languageen
FieldMedicine
TopicDementia and Cognitive Impairment Research
Canadian institutionsPublic Health OntarioCanada Research ChairsHealth Sciences CentreUniversity of New BrunswickUniversity of TorontoWestern UniversitySunnybrook Health Science CentreLawson Health Research InstituteParkwood InstituteUniversity of WaterlooMcMaster UniversityBaycrest Hospital
FundersOntario Ministry of Research and InnovationWestern UniversityUniversity of TorontoStrongCanadian Institutes of Health ResearchSunnybrook Research InstituteOntario Brain InstituteOntario Ministry of Research, Innovation and ScienceHeart and Stroke Foundation of Canada
KeywordsMultivariate statisticsComputer scienceOutlierMultivariate analysisData miningAnomaly detectionData scienceQuality (philosophy)Data qualityInformation retrievalMachine learningArtificial intelligenceOperations managementEngineering

Abstract

fetched live from OpenAlex

BACKGROUND: Large and complex studies are now routine, and quality assurance and quality control (QC) procedures ensure reliable results and conclusions. Standard procedures may comprise manual verification and double entry, but these labour-intensive methods often leave errors undetected. Outlier detection uses a data-driven approach to identify patterns exhibited by the majority of the data and highlights data points that deviate from these patterns. Univariate methods consider each variable independently, so observations that appear odd only when two or more variables are considered simultaneously remain undetected. We propose a data quality evaluation process that emphasizes the use of multivariate outlier detection for identifying errors, and show that univariate approaches alone are insufficient. Further, we establish an iterative process that uses multiple multivariate approaches, communication between teams, and visualization for other large-scale projects to follow. METHODS: We illustrate this process with preliminary neuropsychology and gait data for the vascular cognitive impairment cohort from the Ontario Neurodegenerative Disease Research Initiative, a multi-cohort observational study that aims to characterize biomarkers within and between five neurodegenerative diseases. Each dataset was evaluated four times: with and without covariate adjustment using two validated multivariate methods - Minimum Covariance Determinant (MCD) and Candès' Robust Principal Component Analysis (RPCA) - and results were assessed in relation to two univariate methods. Outlying participants identified by multiple multivariate analyses were compiled and communicated to the data teams for verification. RESULTS: Of 161 and 148 participants in the neuropsychology and gait data, 44 and 43 were flagged by one or both multivariate methods and errors were identified for 8 and 5 participants, respectively. MCD identified all participants with errors, while RPCA identified 6/8 and 3/5 for the neuropsychology and gait data, respectively. Both outperformed univariate approaches. Adjusting for covariates had a minor effect on the participants identified as outliers, though did affect error detection. CONCLUSIONS: Manual QC procedures are insufficient for large studies as many errors remain undetected. In these data, the MCD outperforms the RPCA for identifying errors, and both are more successful than univariate approaches. Therefore, data-driven multivariate outlier techniques are essential tools for QC as data become more complex.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.278
metaresearch head score (Gemma)0.186
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Other design · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.856
Threshold uncertainty score0.821

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.2780.186
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.001
Science and technology studies0.0000.001
Scholarly communication0.0000.000
Open science0.0010.001
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.755
GPT teacher head0.685
Teacher spread0.070 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designOther design
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations72
Published2019
Admission routes3
Has abstractyes

Explore more

Same venueBMC Medical Research MethodologySame topicDementia and Cognitive Impairment ResearchFrench-language works237,207