MétaCan
Menu
← Back to cohort
Record W2420868261

Influential observations in weighted analyses: examples from the National Longitudinal Survey of Children and Youth (NLSCY).

2005· article· en· W2420868261 on OpenAlexaff
Jennifer J. Macnab, John J. Koval, Kathy N. Speechley, M. Karen Campbell

Bibliographic record

VenuePubMed · 2005
Typearticle
Languageen
FieldHealth Professions
TopicFood Security and Health in Diverse Populations
Canadian institutionsWestern University
Fundersnot available
KeywordsMedicineLinear regressionRegression analysisStatisticsAdolescent healthWeightingPopulationPopulation healthDemographyEnvironmental healthMathematics
DOInot available

Abstract

fetched live from OpenAlex

This paper highlights the impact of survey weights on model fit in multiple linear regression with specific reference to the National Longitudinal Survey of Children and Youth (NLSCY) and provides recommendations for the treatment of influential observations. Multiple linear regression was used to estimate the association between child and family factors in the preschool years and vocabulary development at school age. Analyses were performed with and without survey weights. The model fit was assessed by examining the distribution of the studentized residuals and the change in the regression coefficients that would occur if an observation were removed. Two summary measures of influence, Dffits and Cook's D are reported. The models were refit excluding influential observations. Weighting of the linear model resulted in previously non-influential observations having an undue influence on the estimation of the regression parameters in the weighted model. The influential observations were driven primarily by the size of the survey weight as opposed to unusual values of x and y. Researchers working with large national health surveys such as the NLSCY and the National Population Health Survey (NPHS) are advised to include a detailed influence analysis before any final conclusions are made.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.061
metaresearch head score (Gemma)0.209
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.939
Threshold uncertainty score0.324

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0610.209
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.002
Bibliometrics0.0040.009
Science and technology studies0.0010.001
Scholarly communication0.0010.002
Open science0.0010.003
Research integrity0.0010.002
Insufficient payload (model declined to judge)0.0020.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.495
GPT teacher head0.449
Teacher spread0.046 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designSimulation or modeling
DomainMethods
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations11
Published2005
Admission routes1
Has abstractyes

Explore more

Same venuePubMed→Same topicFood Security and Health in Diverse Populations→French-language works237,207→