Bibliographic record
Abstract
To the Editor: Individual-level statistical analyses are paramount for obtaining accurate estimates of an exposure-outcome relation in population groups. However, data privacy and confidentiality concerns have led to a conflict between ethico-legal restrictions to access microdata and scientific accuracy achievable through analyses of individual-level data.1,2 Disclosure of results is also restricted by rules protecting subjects’ confidentiality, such as small cell data suppression, rounding, and collapsing.3 Data aggregation is a major way to share data publicly while protecting the confidentiality of the subjects, but regression models for summary continuous data are lacking. We have developed a method to fill this gap. To perform linear regression analyses on a continuous aggregate outcome we need only 2 parameters: the mean and the standard deviation (SD), within strata of a set of categorical predictors. The frequencies (counts) of the combinations of the predictors can be used as weights. The SD of the raw data for all combinations of the predictors can be used to calculate the pooled variance that subsequently can be used to correct the standard errors (SE) of the estimated parameters using aggregate data. To illustrate the application of the method, we focused on the association between receipt of WIC (The Special Supplemental Nutrition Program for Women, Infants, and Children) food for the mother during this pregnancy (http://www.fns.usda.gov/wic/about-wic) and gestational weight gain. We used a subset of the 2012 Natality Public Use Births File of the National Center for Health Statistics (NCHS).4 The subset is drawn from the 2003 revision of the U.S. Standard Certificate of Live Birth and includes singleton term pregnancies (37–41 weeks gestation) of underweight (body mass index is less than 18.5) women aged 20 to 25 years, who did not complete high school, and were Medicaid recipients. Records with unknown prenatal care initiation information and “other” race/ethnicity were excluded. The final sample contains 5,270 observations and 4 variables. The aggregate dataset has 12 observations and includes the number, mean, and SD for each combination of the levels of the predictors. A technical description of the method, the SAS (SAS Institute, Cary, NC) program to analyze the data and the aggregate, and individual-level datasets are in the eAppendix (https://links.lww.com/EDE/A830; URL). The point estimates of the regression model based on the aggregate data are identical to those based on the microdata (Table). The SEs based on aggregate regression do not differ by more than 1% from those based on the individual-level regression. Including covariates, even product terms, improved the estimation.TABLE: Linear Regression Models Based on Individual-level and Aggregate DataThe main limitation is that continuous covariates cannot be accommodated. However, continuous covariates can be collapsed into categories. This method has several potential applications. It can be used to perform meta-analyses and pooled analyses of multicenter, international, or comparative studies. Perhaps more important, it can be used to analyze summary information from datasets that otherwise cannot be accessed due to data confidentiality concerns, at least until open data initiatives and validated mechanisms to share microdata are in place.5,6 The inclusion of this method in our analytic toolkit challenges us to revisit the practice of categorizing continuous endpoints. Most data repository reporting systems make summary statistics publicly available in the form of counts and proportions, even when measures are originally of continuous nature, for variables such as birthweight, body mass index, and lab tests. Categorization of continuous outcomes has some shortcomings,7,8 such as assuming risk homogeneity within groups, multiple testing, and loss of power. Categorizing continuous data also creates dissent regarding the choice of the categories and appropriateness of the cutoff points, which hampers comparison of results across studies. Using means and standard deviations across groups may avoid such shortcomings, if properly analyzed. Rahim Moineddin Department of Family and Community Medicine University of Toronto Toronto, Ontario, Canada Marcelo Luis Urquia St. Michael’s Hospital Toronto, Ontario, Canada [email protected]
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.011 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.003 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.004 | 0.005 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".