Estimating spatial patterning of dietary behaviors using grocery transaction data
Notice bibliographique
Résumé
ObjectiveTo demonstrate a method for estimating neighborhood foodselection with secondary use of digital marketing data; grocerytransaction records and retail business registry.IntroductionUnhealthy diet is becoming the most important preventablecause of chronic disease burden (1). Dietary patterns vary acrossneighborhoods as a function of policy, marketing, social support,economy, and the commercial food environment (2). Assessmentof community-specific response to these socio-ecological factorsis critical for the development and evaluation policy interventionsand identification of nutrition inequality. Mass administration ofdietary surveys is impractical and prohibitory expensive, and surveystypically fail to address variation of food selection at high geographicresolution. Marketing companies such as the Nielsen cooperationcontinuously collect and centralize scanned grocery transactionrecords from a geographically representative sample of retail foodoutlets to guide product promotions. These data can be harnessed todevelop a model for the demand of specific foods using store andneighborhood attributes, providing a rich and detailed picture of the“foodscape” in an urban environment. In this study, we generated aspatial profile of food selection from estimated sales in food outletsin the Census Metropolitan Area (CMA) of Montreal, Canada,using regular carbonated soft drinks (i.e. non-diet soda) as an initialexample.MethodsFrom the Nielsen cooperation, we obtained weekly grocerytransaction data generated by a sample of 86 grocery stores and 42pharmacies in the Montreal CMA in 2012. Extracted store-specificsoda sales were standardized to a single serving size (240ml) andaveraged across 52 weeks, resulting in 128 data points. Using linearregression, natural log-transformed soda sales were modelled as afunction of store type (grocery vs. pharmacies), chain identificationcode and socio-demographic attributes of store neighborhood, whichare median family income, proportion of individuals who receivedpost-secondary diplomas, and population density as measured by the2011 Canadian Household Survey. Selection of the predictors andfirst-order interaction terms was guided by the minimization of themean squared error using 10-fold cross-validation. The final modelwas applied to all operating chain grocery stores and pharmacies in2012 (n=980) recorded in a comprehensive and commonly availablebusiness establishment database. The resulting predicted store-specific weekly average soda sales was spatially interpolated toprovide a graphical representation of the soda sales (representing anunhealthy foodscape) across the Montreal CMA.ResultsFigure 2 demonstrates the spatial distribution of the predicted sodasales in the Montreal CMA.ConclusionsThe current lack of neighborhood-level dietary surveillanceimpedes effective public health actions aimed at encouraging healthyfood selection and subsequent reduction of chronic illness. Ourmethod leverages existing grocery transaction data and store locationinformation to address the gap in population monitoring of nutritionstatus and urban foodscapes. Future applications of our methodologyto other store types (e.g. convenience stores) and food productsacross multiple time points (e.g. mouths and years) will permit acomprehensive, timely and automated assessment of dietary trends,identification of neighborhoods in special dietary needs, developmentof tailored community health promotions, and the measurement ofneighbourhood-specific response to nutrition policies and unhealthyfood advertising.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,002 | 0,001 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,003 |
| Science ouverte | 0,001 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».