MétaCan
Menu
Back to cohort
Record W2314719657 · doi:10.1097/ede.0000000000000172

Regression Analysis of Aggregate Continuous Data

2014· letter· en· W2314719657 on OpenAlexafffundabout
Rahim Moineddin, Marcelo L. Urquía

Bibliographic record

VenueEpidemiology · 2014
Typeletter
Languageen
FieldHealth Professions
TopicFood Security and Health in Diverse Populations
Canadian institutionsUniversity of TorontoSt. Michael's Hospital
FundersCanadian Institutes of Health Research
KeywordsCategorical variableAggregate dataMicrodata (statistics)StatisticsLinear regressionPopulationData setConfidentialityRegression analysisEconometricsAggregate (composite)Standard errorMedicineComputer scienceMathematicsEnvironmental healthCensus

Abstract

fetched live from OpenAlex

To the Editor: Individual-level statistical analyses are paramount for obtaining accurate estimates of an exposure-outcome relation in population groups. However, data privacy and confidentiality concerns have led to a conflict between ethico-legal restrictions to access microdata and scientific accuracy achievable through analyses of individual-level data.1,2 Disclosure of results is also restricted by rules protecting subjects’ confidentiality, such as small cell data suppression, rounding, and collapsing.3 Data aggregation is a major way to share data publicly while protecting the confidentiality of the subjects, but regression models for summary continuous data are lacking. We have developed a method to fill this gap. To perform linear regression analyses on a continuous aggregate outcome we need only 2 parameters: the mean and the standard deviation (SD), within strata of a set of categorical predictors. The frequencies (counts) of the combinations of the predictors can be used as weights. The SD of the raw data for all combinations of the predictors can be used to calculate the pooled variance that subsequently can be used to correct the standard errors (SE) of the estimated parameters using aggregate data. To illustrate the application of the method, we focused on the association between receipt of WIC (The Special Supplemental Nutrition Program for Women, Infants, and Children) food for the mother during this pregnancy (http://www.fns.usda.gov/wic/about-wic) and gestational weight gain. We used a subset of the 2012 Natality Public Use Births File of the National Center for Health Statistics (NCHS).4 The subset is drawn from the 2003 revision of the U.S. Standard Certificate of Live Birth and includes singleton term pregnancies (37–41 weeks gestation) of underweight (body mass index is less than 18.5) women aged 20 to 25 years, who did not complete high school, and were Medicaid recipients. Records with unknown prenatal care initiation information and “other” race/ethnicity were excluded. The final sample contains 5,270 observations and 4 variables. The aggregate dataset has 12 observations and includes the number, mean, and SD for each combination of the levels of the predictors. A technical description of the method, the SAS (SAS Institute, Cary, NC) program to analyze the data and the aggregate, and individual-level datasets are in the eAppendix (https://links.lww.com/EDE/A830; URL). The point estimates of the regression model based on the aggregate data are identical to those based on the microdata (Table). The SEs based on aggregate regression do not differ by more than 1% from those based on the individual-level regression. Including covariates, even product terms, improved the estimation.TABLE: Linear Regression Models Based on Individual-level and Aggregate DataThe main limitation is that continuous covariates cannot be accommodated. However, continuous covariates can be collapsed into categories. This method has several potential applications. It can be used to perform meta-analyses and pooled analyses of multicenter, international, or comparative studies. Perhaps more important, it can be used to analyze summary information from datasets that otherwise cannot be accessed due to data confidentiality concerns, at least until open data initiatives and validated mechanisms to share microdata are in place.5,6 The inclusion of this method in our analytic toolkit challenges us to revisit the practice of categorizing continuous endpoints. Most data repository reporting systems make summary statistics publicly available in the form of counts and proportions, even when measures are originally of continuous nature, for variables such as birthweight, body mass index, and lab tests. Categorization of continuous outcomes has some shortcomings,7,8 such as assuming risk homogeneity within groups, multiple testing, and loss of power. Categorizing continuous data also creates dissent regarding the choice of the categories and appropriateness of the cutoff points, which hampers comparison of results across studies. Using means and standard deviations across groups may avoid such shortcomings, if properly analyzed. Rahim Moineddin Department of Family and Community Medicine University of Toronto Toronto, Ontario, Canada Marcelo Luis Urquia St. Michael’s Hospital Toronto, Ontario, Canada [email protected]

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.007
metaresearch head score (Gemma)0.011
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Meta-epidemiology (narrow), Research integrity, Insufficient payload (model declined to judge)
Consensus categoriesResearch integrity
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Commentary · Consensus signal: Commentary
Teacher disagreement score0.040
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0070.011
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0030.000
Bibliometrics0.0010.001
Science and technology studies0.0010.000
Scholarly communication0.0000.000
Open science0.0010.001
Research integrity0.0040.005
Insufficient payload (model declined to judge)0.0020.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.524
GPT teacher head0.554
Teacher spread0.030 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designNot applicable
Domainnot available
GenreCommentary

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations13
Published2014
Admission routes3
Has abstractyes

Explore more

Same venueEpidemiologySame topicFood Security and Health in Diverse PopulationsFrench-language works237,207