A Mobile App to Rapidly Appraise the In-Store Food Environment: Reliability, Utility, and Construct Validity Study
Notice bibliographique
Résumé
BACKGROUND: Consumer food environments are increasingly being recognized as influential determinants of food purchasing and subsequent intake and health. We developed a tool to enable efficient, but relatively comprehensive, appraisal of the in-store food environment. The Store Scout mobile app facilitates the evaluation of product (availability and range), placement (visibility, accessibility, proximity to high-traffic areas, and location relative to other products), price (price promotion), and promotion (displays and advertising) across 7 categories of food products, with appraisal given immediately as scores (0-100, where a higher score is more in line with best practice). Primary end users are public health nutritionists and nutritionists employed by store organizations; however, store managers and staff are also potential end users. OBJECTIVE: This study aims to evaluate the reliability (interrater reliability and internal consistency), utility (distribution of scores), and construct validity (score by store type) of measurements using the Store Scout mobile app. METHODS: The Store Scout mobile app was used independently by 2 surveyors to evaluate the store environment in 54 stores: 34 metropolitan stores (9 small and 11 large supermarkets, 10 convenience stores, and 4 petrol stations) in Brisbane, Australia, and 20 remote stores (19 small supermarkets and 1 petrol station) in Indigenous Australian communities in Northern Australia. The agreement between surveyors in the overall and category scores was evaluated using intraclass correlation coefficients (ICCs). Interrater reliability of measurement items was assessed using percentage agreement and the Gwet agreement coefficient (AC). Internal consistency was assessed by comparing the responses of items measuring similar aspects of the store environment. We examined the distribution of score values using boxplots and differences by store type using the Kruskal-Wallis test. RESULTS: The median difference in the overall score between surveyors was 4.4 (range 0.0-11.1), with an ICC of 0.954 (95% CI 0.914-0.975). Most measurement items had very good (n=74/196, 37.8%) or good (n=81/196, 41.3%) interrater reliability using the Gwet AC. A minimal inconsistency of measurement was found. Overall scores ranged from 19.2 to 81.6. There was a significant difference in score by store type (P<.001). Large Brisbane supermarkets scored highest (median 77.4, range 53.2-81.6), whereas small Brisbane supermarkets (median 63.9, range 41.0-71.3) and small remote supermarkets (median 63.8, range 56.5-74.9) scored significantly higher than Brisbane petrol stations (median 33.1, range 19.2-37.8) and convenience stores (median 39.0, range 22.4-63.8). CONCLUSIONS: These findings suggest good reliability and internal consistency of food environment measurements using the Store Scout mobile app. We identified specific aspects that can be improved to further increase the reliability of this tool. We found a good distribution of score values and evidence that scoring could capture differences by store type in line with previous evidence, which gives an indication of construct validity. The Store Scout mobile app shows promise in its capability to measure and track the health-enabling characteristics of store environments.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,017 | 0,038 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,002 | 0,001 |
| Études des sciences et des technologies | 0,001 | 0,001 |
| Communication savante | 0,001 | 0,001 |
| Science ouverte | 0,001 | 0,001 |
| Intégrité de la recherche | 0,001 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,001 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».