MétaCan
Menu
Retour à la cohorte
Enregistrement W2807984017 · doi:10.5210/ojphi.v10i1.8908

Evaluation of approaches that adjust for biases in participatory surveillance systems

2018· article· en· W2807984017 sur OpenAlexaboutno aff
Kristin Baltrusaitis, Kathleen Noddin, Colleen Nguyen, Adam W. Crawley, John S. Brownstein, Laura F. White

Notice bibliographique

RevueOnline Journal of Public Health Informatics · 2018
Typearticle
Langueen
DomaineMedicine
ThématiqueInfluenza Virus Research Studies
Établissements canadiensnon disponible
Organismes subventionnairesnon disponible
Mots-clésDisease surveillancePsychological interventionMedicinePopulationCitizen journalismHealth careMissing dataDiseaseActuarial scienceEnvironmental healthComputer scienceBusinessNursingPolitical science

Résumé

récupéré en direct d'OpenAlex

ObjectiveTo estimate and compare influenza attack rates (AR) in the United States (US) using different approaches to adjust for reporting biases in participatory syndromic surveillance data.IntroductionBecause the dynamics and severity of influenza in the US vary each season, yearly estimates of disease burden in the population are essential to evaluate interventions and allocate resources. The CDC uses data from a national health-care based surveillance system and mathematical models to estimate the overall burden of disease in the general population. Over the past decade, crowd-sourced syndromic surveillance systems have emerged as a digital data source that collects health-related information in near real-time. These systems complement traditional surveillance systems by capturing individuals who do not seek medical care and allowing for a longitudinal view of illness burden. However, because not all participants report every week and participants are more likely to report when ill, the number of weekly reports is temporally and spatially inconsistent and the estimates of disease burden and incidence may be biased. In this study, we use data from Flu Near You (FNY), a participatory surveillance system based in the US and Canada1, to estimate and compare Influenza-like Illness (ILI) ARs using different approaches to adjust for reporting biases in participatory surveillance data.MethodsThis analysis uses FNY data from the 2015-16 influenza season. Four different approaches of bias adjustment were assessed. The first approach includes all FNY participants, defined as users and household members, who submitted at least one symptom report, whereas the second approach only includes FNY participants who submitted at least 10 symptom reports. The third approach includes all FNY participants who submitted at least one symptom report, but drops the first symptom report for all participants. For the first three approaches, all missing reports were assumed to be non-ILI when estimating attack rates. Finally, the fourth approach includes FNY participants who submitted at least 10 symptom reports and uses multiple imputation to account for missing reports. Age-stratified and overall estimates of ILI ARs were calculated for each of the four approaches to bias adjustment by dividing the sum of the weekly incident cases of ILI, defined as the first report of fever with cough and/or sore throat, by the population at risk at the beginning of the period.ResultsDuring the 2016-2017 influenza season, FNY received an average of 10,723 unique symptom reports per week from 46,390 registered users and their household members. For FNY, the youngest age group assessed, 5-17, had the largest ILI AR, and the ILI ARs decreased as the age group increased for all approaches. Overall, the approach that drops all first reports had the smallest ARs, whereas the approach that selects a cohort of users who submit at least 10 reports during the season and imputes the missing reports had the largest ARs. Although the influenza ARs estimated by the CDC were less than the ILI ARs estimated using FNY data for all age-groups, a similar pattern was observed across age groups, except for the 50-64 age group, which had the largest influenza AR.ConclusionsAs expected, the ARs estimated using FNY data were greater than the CDC’s influenza ARs because FNY estimates ARs of ILI and does not adjust for the probability of reporting ILI when experiencing non-flu illness. The approach of dropping the first report had the smallest ARs because during the 2015-16 influenza season the weekly percent of ILI cases that were first time reports ranged from 18-59%. This approach was developed to adjust for the potential correlation between symptom presence and willingness to join the platform. However, important information about the dynamics of disease may be lost when using this approach. The multiple imputation method was used only for individuals who submitted at least 10 reports to maintain a missing data rate below 30%. The imputation model also assumed that data were missing at random, which may not be appropriate in this case, because approximately 30% of FNY users have reported that they are more likely to report when ill. As shown in Table 1, the AR estimate depends on the bias adjustment approach. Simulation-based studies should be performed to further evaluate these methods.References1. Smolinski MS, Crawley AW, Baltrusaitis K, Chunara R, Olsen JM, Wójcik O, et al. Flu Near You: Crowdsourced Symptom Reporting Spanning 2 Influenza Seasons. Am J Public Health. 20152. Rolfes MA, Foppa IM, Garg S, Flannery B, Brammer L, Singleton JA, et al. Estimated Influenza Illnesses, Medical Visits, Hospitalizations, and Deaths Averted by Vaccination in the United States. 2016 Dec 9 [2017 Sept 25];https://www.cdc.gov/flu/about/disease/2015-16.htm

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,346
score de la tête « metaresearch » (Gemma)0,550
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesMétarecherche
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Observationnel · Signal consensuel: aucune
GenreSignal candidat: Empirique · Signal consensuel: aucune
Score de désaccord entre enseignants0,346
Score d'incertitude au seuil0,806

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,3460,550
Méta-épidémiologie (sens strict)0,0040,002
Méta-épidémiologie (sens large)0,0020,006
Bibliométrie0,0040,003
Études des sciences et des technologies0,0020,003
Communication savante0,0040,007
Science ouverte0,0050,010
Intégrité de la recherche0,0040,003
Charge utile insuffisante (le modèle a refusé de juger)0,0030,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,827
Tête enseignante GPT0,540
Écart entre enseignants0,287 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Devis d'étudeObservationnel
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations4
Publié2018
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueOnline Journal of Public Health InformaticsMême sujetInfluenza Virus Research StudiesTravaux en français237 207