Detecting Influenza Epidemics Using Self-reported Data Through Mobile App (FeverCoach)
Bibliographic record
Abstract
Background: Timely forecast of influenza activity is critical for a public health system to prepare for an influenza epidemic and mitigate its burden. Currently, influenza surveillance relies on traditional data sources such as reports from health care providers, which lag behind real-time by several days to weeks. In an effort to reduce the time lag, internet search information, voluntary web-based records, and electronic health records have been suggested as the alternative data sources for influenza surveillance. However, low specificity, low rate of report, or privacy concerns limits the use of such data. Objective: FeverCoach mobile application provides tailored information to help caregivers manage a febrile child. Using the self-reported diagnosis data submitted to the app, we developed a new algorithm that accurately predicted the influenza trend in South Korea. Methods: Users of FeverCoach agreed to the use of de-identified data for research purposes. The app shows information about use of antipyretics and adjuvant way to relieve fever when users enter the child’s age, sex, body temperature, and the duration of fever. Users can choose from the list of 21 candidate diseases including Influenza after a physician office visit. Additional information about the disease was provided following submission of the diagnosis. Public influenza-like illness (ILI) data was obtained from the Korea Centers for Disease Control and Prevention (KCDC) website. The data was collected from September 2016 to March 2017. Ordinary least squares linear regression was used to build a model using the data from the app to predict the influenza trend. To perform linear regression, we calculate logit(Pcdc) and logit(Papp) where logit(p) is natural log of p/(1-p), Pcdc is (ILI visit counts)/(total patient visit counts) and Papp is (Influenza report on FeverCoach)/(total diagnosis report on FeverCoach). Results: We collected 13,014 self-reported diagnoses. Of all users, 81% of the children were under 5 years of age. The animated visualization of spatiotemporal diagnosis report is available online at https://www.youtube.com/watch?v=-8kDXz43gO8. Ordinary least square regression showed significant association between logit(Pcdc) and logit(Papp) (R2=0.860, P<.001). Using this regression model, we could detect an influenza epidemic 5 days before the 2016-2017 season’s influenza epidemic alert by KCDC. Conclusions: We found that it is possible to predict influenza epidemics earlier than KCDC with a relatively small amoount data. Collection of specific and accurate data was made possible by targeting a well-defined population.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".