Automation of the Updated Food Label Information Program (FLIP 2020): A Comprehensive Canadian Branded Grocery and Restaurant Food Composition Database
Notice bibliographique
Résumé
Traditional methods for creating food composition databases struggle to cope with the large number of products and the rapid pace of turnover in the food supply. The objective is to overview the updated Food Label Information Program (FLIP2020), a big data approach for the evaluation of the Canadian food supply and present the latest methods used in the development of this database. The University of Toronto's Food Label Information Program (FLIP) is a database of Canadian prepackaged and chain restaurant foods and beverages collected since 2010. FLIP 2020 was developed using website “scraping” and machine learning (ML) coupled with artificial intelligence-enhanced optical character recognition (AI-OCR) to collect and manage food labelling information (e.g., nutritional composition, price, product images, ingredients, brand, etc.) on all foods and beverages available on seven major Canadian e-grocery retailer websites and 201 Canadian chain restaurants between May 2020 and February 2021. FLIP 2020 is comprised of 74,445 prepackaged food products and 21,225 menu items available on websites of seven retailers, 2 location-specific duplicate retailers and 141 chain restaurants. Food products were classified under multiple national and international categorization systems, in order to analyse similar foods under different systems. Of 57,006 food and beverage products available on seven retailers’ websites, nutritional composition data were available for about 60% of the products and ingredients were available for about 45%. Data for energy, protein, carbohydrate, fat, sugar, sodium and saturated fat were present for 54–65% of the products, while fibre information was available for 37%. Of the 201 eligible chain restaurants with ≥ 20 national outlets, 70% provided nutritional information. All provided energy, 84% provided saturated fat, total sugar and sodium, and 50% provided all 13 required nutrients listed on the Nutrition Facts table. FLIP, with its comprehensive sampling and granularity and use of ML/AI-OCR, is a powerful tool for evaluating and monitoring the Canadian food supply environment. This research was supported by funds from a Canadian Institutes of Health Research Project Grant.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,000 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,001 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».