MétaCan
Menu
Back to cohort
Record W4391843538 · doi:10.1186/s12874-024-02161-1

Identifying biologically implausible values in big longitudinal data: an example applied to child growth data from the Brazilian food and nutrition surveillance system

2024· article· en· W4391843538 on OpenAlexafffund
Juliana Freitas de Mello e Silva, Natanael de Jesus Silva, Thaís Rangel Bousquet Carrilho, Elizabete de Jesus Pinto, Aline S. Rocha, Jéssica Pedroso, Sara Araújo da Silva, Ana Maria Spaniol, Rafaella da Costa Santin de Andrade, Gisele Ane Bortolini, Enny S. Paixão, Gilberto Kac, Rita de Cássia Ribeiro‐Silva, Maurício L. Barreto

Bibliographic record

VenueBMC Medical Research Methodology · 2024
Typearticle
Languageen
FieldMedicine
TopicObesity, Physical Activity, Diet
Canadian institutionsUniversity of British Columbia
FundersAgencia Estatal de InvestigaciónSecretaría de Ciencia y Técnica, Universidad de Buenos AiresMinistério da SaúdeConselho Nacional de Desenvolvimento Científico e TecnológicoGeneralitat de CatalunyaMichael Smith Health Research BCCentres de Recerca de CatalunyaMinisterio de Ciencia e InnovaciónWellcome TrustBill and Melinda Gates Foundation
KeywordsOutlierAnthropometryLongitudinal studyPopulationLongitudinal dataStatisticsMedicineDemographyStandard scoreMathematicsEnvironmental health

Abstract

fetched live from OpenAlex

BACKGROUND: Several strategies for identifying biologically implausible values in longitudinal anthropometric data have recently been proposed, but the suitability of these strategies for large population datasets needs to be better understood. This study evaluated the impact of removing population outliers and the additional value of identifying and removing longitudinal outliers on the trajectories of length/height and weight and on the prevalence of child growth indicators in a large longitudinal dataset of child growth data. METHODS: Length/height and weight measurements of children aged 0 to 59 months from the Brazilian Food and Nutrition Surveillance System were analyzed. Population outliers were identified using z-scores from the World Health Organization (WHO) growth charts. After identifying and removing population outliers, residuals from linear mixed-effects models were used to flag longitudinal outliers. The following cutoffs for residuals were tested to flag those: -3/+3, -4/+4, -5/+5, -6/+6. The selected child growth indicators included length/height-for-age z-scores and weight-for-age z-scores, classified according to the WHO charts. RESULTS: The dataset included 50,154,738 records from 10,775,496 children. Boys and girls had 5.74% and 5.31% of length/height and 5.19% and 4.74% of weight values flagged as population outliers, respectively. After removing those, the percentage of longitudinal outliers varied from 0.02% (<-6/>+6) to 1.47% (<-3/>+3) for length/height and from 0.07 to 1.44% for weight in boys. In girls, the percentage of longitudinal outliers varied from 0.01 to 1.50% for length/height and from 0.08 to 1.45% for weight. The initial removal of population outliers played the most substantial role in the growth trajectories as it was the first step in the cleaning process, while the additional removal of longitudinal outliers had lower influence on those, regardless of the cutoff adopted. The prevalence of the selected indicators were also affected by both population and longitudinal (to a lesser extent) outliers. CONCLUSIONS: Although both population and longitudinal outliers can detect biologically implausible values in child growth data, removing population outliers seemed more relevant in this large administrative dataset, especially in calculating summary statistics. However, both types of outliers need to be identified and removed for the proper evaluation of trajectories.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.034
metaresearch head score (Gemma)0.023
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.316
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0340.023
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0000.001
Science and technology studies0.0000.001
Scholarly communication0.0000.000
Open science0.0020.004
Research integrity0.0000.002
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.642
GPT teacher head0.520
Teacher spread0.123 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations4
Published2024
Admission routes2
Has abstractyes

Explore more

Same venueBMC Medical Research MethodologySame topicObesity, Physical Activity, DietFrench-language works237,207