Applying a New Model for Sharing Population Health Data to National Syndromic Influenza Surveillance: DiSTRIBuTE Project Proof of Concept, 2006 to 2009
Bibliographic record
Abstract
The Distributed Surveillance Taskforce for Real-time Influenza Burden Tracking and Evaluation (DiSTRIBuTE) project began as a pilot effort initiated by the International Society for Disease Surveillance (ISDS) in autumn 2006 to create a collaborative electronic emergency department (ED) syndromic influenza-like illness (ILI) surveillance network based on existing state and local systems and expertise. DiSTRIBuTE brought together health departments that were interested in: 1) sharing aggregate level data; 2) maintaining jurisdictional control; 3) minimizing barriers to participation; and 4) leveraging the flexibility of local systems to create a dynamic and collaborative surveillance network. This approach was in contrast to the prevailing paradigm for surveillance where record level information was collected, stored and analyzed centrally. The DiSTRIBuTE project was created with a distributed design, where individual level data remained local and only summarized, stratified counts were reported centrally, thus minimizing privacy risks. The project was responsive to federal mandates to improve integration of federal, state, and local biosurveillance capabilities. During the proof of concept phase, 2006 to 2009, ten jurisdictions from across North America sent ISDS on a daily to weekly basis year-round, aggregated data by day, stratified by local ILI syndrome, age-group and region. During this period, data from participating U.S. state or local health departments captured over 13% of all ED visits nationwide. The initiative focused on state and local health department trust, expertise, and control. Morbidity trends observed in DiSTRIBuTE were highly correlated with other influenza surveillance measures. With the emergence of novel A/H1N1 influenza in the spring of 2009, the project was used to support information sharing and ad hoc querying at the state and local level. In the fall of 2009, through a broadly collaborative effort, the project was expanded to enhance electronic ED surveillance nationwide.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.021 | 0.032 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.003 | 0.005 |
| Open science | 0.005 | 0.004 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".