MétaCan
Menu
Retour à la cohorte
Enregistrement W2314719657 · doi:10.1097/ede.0000000000000172

Regression Analysis of Aggregate Continuous Data

2014· letter· en· W2314719657 sur OpenAlexafffundabout
Rahim Moineddin, Marcelo L. Urquía

Notice bibliographique

RevueEpidemiology · 2014
Typeletter
Langueen
DomaineHealth Professions
ThématiqueFood Security and Health in Diverse Populations
Établissements canadiensUniversity of TorontoSt. Michael's Hospital
Organismes subventionnairesCanadian Institutes of Health Research
Mots-clésCategorical variableAggregate dataMicrodata (statistics)StatisticsLinear regressionPopulationData setConfidentialityRegression analysisEconometricsAggregate (composite)Standard errorMedicineComputer scienceMathematicsEnvironmental healthCensus

Résumé

récupéré en direct d'OpenAlex

To the Editor: Individual-level statistical analyses are paramount for obtaining accurate estimates of an exposure-outcome relation in population groups. However, data privacy and confidentiality concerns have led to a conflict between ethico-legal restrictions to access microdata and scientific accuracy achievable through analyses of individual-level data.1,2 Disclosure of results is also restricted by rules protecting subjects’ confidentiality, such as small cell data suppression, rounding, and collapsing.3 Data aggregation is a major way to share data publicly while protecting the confidentiality of the subjects, but regression models for summary continuous data are lacking. We have developed a method to fill this gap. To perform linear regression analyses on a continuous aggregate outcome we need only 2 parameters: the mean and the standard deviation (SD), within strata of a set of categorical predictors. The frequencies (counts) of the combinations of the predictors can be used as weights. The SD of the raw data for all combinations of the predictors can be used to calculate the pooled variance that subsequently can be used to correct the standard errors (SE) of the estimated parameters using aggregate data. To illustrate the application of the method, we focused on the association between receipt of WIC (The Special Supplemental Nutrition Program for Women, Infants, and Children) food for the mother during this pregnancy (http://www.fns.usda.gov/wic/about-wic) and gestational weight gain. We used a subset of the 2012 Natality Public Use Births File of the National Center for Health Statistics (NCHS).4 The subset is drawn from the 2003 revision of the U.S. Standard Certificate of Live Birth and includes singleton term pregnancies (37–41 weeks gestation) of underweight (body mass index is less than 18.5) women aged 20 to 25 years, who did not complete high school, and were Medicaid recipients. Records with unknown prenatal care initiation information and “other” race/ethnicity were excluded. The final sample contains 5,270 observations and 4 variables. The aggregate dataset has 12 observations and includes the number, mean, and SD for each combination of the levels of the predictors. A technical description of the method, the SAS (SAS Institute, Cary, NC) program to analyze the data and the aggregate, and individual-level datasets are in the eAppendix (https://links.lww.com/EDE/A830; URL). The point estimates of the regression model based on the aggregate data are identical to those based on the microdata (Table). The SEs based on aggregate regression do not differ by more than 1% from those based on the individual-level regression. Including covariates, even product terms, improved the estimation.TABLE: Linear Regression Models Based on Individual-level and Aggregate DataThe main limitation is that continuous covariates cannot be accommodated. However, continuous covariates can be collapsed into categories. This method has several potential applications. It can be used to perform meta-analyses and pooled analyses of multicenter, international, or comparative studies. Perhaps more important, it can be used to analyze summary information from datasets that otherwise cannot be accessed due to data confidentiality concerns, at least until open data initiatives and validated mechanisms to share microdata are in place.5,6 The inclusion of this method in our analytic toolkit challenges us to revisit the practice of categorizing continuous endpoints. Most data repository reporting systems make summary statistics publicly available in the form of counts and proportions, even when measures are originally of continuous nature, for variables such as birthweight, body mass index, and lab tests. Categorization of continuous outcomes has some shortcomings,7,8 such as assuming risk homogeneity within groups, multiple testing, and loss of power. Categorizing continuous data also creates dissent regarding the choice of the categories and appropriateness of the cutoff points, which hampers comparison of results across studies. Using means and standard deviations across groups may avoid such shortcomings, if properly analyzed. Rahim Moineddin Department of Family and Community Medicine University of Toronto Toronto, Ontario, Canada Marcelo Luis Urquia St. Michael’s Hospital Toronto, Ontario, Canada [email protected]

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,007
score de la tête « metaresearch » (Gemma)0,011
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesMétarecherche, Méta-épidémiologie (sens strict), Intégrité de la recherche, Charge utile insuffisante (le modèle a refusé de juger)
Catégories consensuellesIntégrité de la recherche
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: Sans objet
GenreSignal candidat: Commentaire · Signal consensuel: Commentaire
Score de désaccord entre enseignants0,040
Score d'incertitude au seuil1,000

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0070,011
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0030,000
Bibliométrie0,0010,001
Études des sciences et des technologies0,0010,000
Communication savante0,0000,000
Science ouverte0,0010,001
Intégrité de la recherche0,0040,005
Charge utile insuffisante (le modèle a refusé de juger)0,0020,001

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,524
Tête enseignante GPT0,554
Écart entre enseignants0,030 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; les deux têtes enseignantes s’accordent sur ce qui est montré ici.

Devis d'étudeSans objet
Domainenon disponible
GenreCommentaire

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations13
Publié2014
Routes d'admission3
Résumé présentoui

Explorer davantage

Même revueEpidemiologyMême sujetFood Security and Health in Diverse PopulationsTravaux en français237 207