AREP CONUS-CA-HI: Annotated Registry of Established Forest Pathogens
Bibliographic record
Abstract
General understanding of emergent diseases in forest systems, as well as the ability to assess risk of future introductions within a process-barrier framework of biological invasion, are impeded by a lack of systematically compiled information on origins and functional traits of pathogens that have become established outside their range. To address these gaps, substantially update previous registries, and provide critical biosecurity information, from 2021 to 2024 we assembled a list of established forest phytopathogens that are present, but not thought to be native in the continental United States (CONUS), Canada (CA), and the Hawaiian islands (HI) under working hypotheses of their origins. We restricted the present list to phytopathogens that cause disease on native tree or woody shrub species, excluding the larger number of species of non-native phytopathogens of trees that are exclusive to agricultural and horticultural species and landscapes which are already well-represented in pest databases (e.g., CABI, EPPO, APHIS, etc.). We used previous databases as a scaffold and supplemented those lists with additional taxa by cross-referencing with lists of forest pathogens and hypothetical origins for other regions (Australia and Europe) as well as by reviewing taxonomic, host, and distribution history of pathogen species that have been recorded in both the study area (CONUS-CA-HI) and at least one other continent. For each of the 93 species in our database, we provide a relational database of a) taxonomic information, b) invasion status in each region, c) first year on record in each region, d) working hypotheses of original range (where possible), e) traits including disease type and name, dispersal mode, and organs and host life stages infected, f) major hosts, g) and > 7,000 chronological records of potential location-year and host-location-year combinations for each pathogen. We also provide a reference-annotated classification system for types of evidence for first years (c), original ranges (d), and major hosts (f). A draft version of this database was reviewed by forest health experts via targeted online surveys and at a workshop of forest pathologists in January 2024 and by the Forest Pathology Committee ahead of the 2025 national meeting of the American Phytopathological Society in Hawaii. This represents a significant expansion compared to previous registries of non-native infectious microorganisms of forest trees in terms of its comprehensiveness, precision, and accuracy of information that includes a way to assess and compare the level of uncertainty associated with key information.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".