Study protocol for a scoping review of Lyme disease prediction methodologies
Bibliographic record
Abstract
INTRODUCTION: In the temperate world, Lyme disease (LD) is the most common vector-borne disease affecting humans. In North America, LD surveillance and research have revealed an increasing territorial expansion of hosts, bacteria and vectors that has accompanied an increasing incidence of the disease in humans. To better understand the factors driving disease spread, predictive models can use current and historical data to predict disease occurrence in populations across time and space. Various prediction methods have been used, including approaches to evaluate prediction accuracy and/or performance and a range of predictors in LD risk prediction research. With this scoping review, we aim to document the different modelling approaches including types of forecasting and/or prediction methods, predictors and approaches to evaluating model performance (eg, accuracy). METHODS AND ANALYSIS: This scoping review will follow the Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Review guidelines. Electronic databases will be searched via keywords and subject headings (eg, Medical Subject Heading terms). The search will be performed in the following databases: PubMed/MEDLINE, EMBASE, CAB Abstracts, Global Health and SCOPUS. Studies reported in English or French investigating the risk of LD in humans through spatial prediction and temporal forecasting methodologies will be identified and screened. Eligibility criteria will be applied to the list of articles to identify which to retain. Two reviewers will screen titles and abstracts, followed by a full-text screening of the articles' content. Data will be extracted and charted into a standard form, synthesised and interpreted. ETHICS AND DISSEMINATION: This scoping review is based on published literature and does not require ethics approval. Findings will be published in peer-reviewed journals and presented at scientific conferences.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.107 | 0.157 |
| Meta-epidemiology (narrow) | 0.005 | 0.005 |
| Meta-epidemiology (broad) | 0.014 | 0.013 |
| Bibliometrics | 0.017 | 0.013 |
| Science and technology studies | 0.005 | 0.004 |
| Scholarly communication | 0.008 | 0.010 |
| Open science | 0.005 | 0.006 |
| Research integrity | 0.011 | 0.008 |
| Insufficient payload (model declined to judge) | 0.163 | 0.031 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".