Detecting Influenza Epidemics Using Self-reported Data Through Mobile App (FeverCoach)
Bibliographic record
Abstract
Background: Timely forecast of influenza activity is critical for a public health system to prepare for an influenza epidemic and mitigate its burden. Currently, influenza surveillance relies on traditional data sources such as reports from health care providers, which lag behind real-time by several days to weeks. In an effort to reduce the time lag, internet search information, voluntary web-based records, and electronic health records have been suggested as the alternative data sources for influenza surveillance. However, low specificity, low rate of report, or privacy concerns limits the use of such data. Objective: FeverCoach mobile application provides tailored information to help caregivers manage a febrile child. Using the self-reported diagnosis data submitted to the app, we developed a new algorithm that accurately predicted the influenza trend in South Korea. Methods: Users of FeverCoach agreed to the use of de-identified data for research purposes. The app shows information about use of antipyretics and adjuvant way to relieve fever when users enter the child’s age, sex, body temperature, and the duration of fever. Users can choose from the list of 21 candidate diseases including Influenza after a physician office visit. Additional information about the disease was provided following submission of the diagnosis. Public influenza-like illness (ILI) data was obtained from the Korea Centers for Disease Control and Prevention (KCDC) website. The data was collected from September 2016 to March 2017. Ordinary least squares linear regression was used to build a model using the data from the app to predict the influenza trend. To perform linear regression, we calculate logit(Pcdc) and logit(Papp) where logit(p) is natural log of p/(1-p), Pcdc is (ILI visit counts)/(total patient visit counts) and Papp is (Influenza report on FeverCoach)/(total diagnosis report on FeverCoach). Results: We collected 13,014 self-reported diagnoses. Of all users, 81% of the children were under 5 years of age. The animated visualization of spatiotemporal diagnosis report is available online at https://www.youtube.com/watch?v=-8kDXz43gO8. Ordinary least square regression showed significant association between logit(Pcdc) and logit(Papp) (R2=0.860, P<.001). Using this regression model, we could detect an influenza epidemic 5 days before the 2016-2017 season’s influenza epidemic alert by KCDC. Conclusions: We found that it is possible to predict influenza epidemics earlier than KCDC with a relatively small amoount data. Collection of specific and accurate data was made possible by targeting a well-defined population.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.002 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".