Documenting the occurrence through space & time of aquatic non-indigenous fish, mollusks, algae, & plants threatening North America's Great Lakes utilizing herbaria & zoological museum specimens
Bibliographic record
Abstract
North America’s Great Lakes contain 21% of the planet’s fresh water, and their protection is a matter of national security to both the USA & Canada. One of the greatest threats to the health of this unparalleled natural resource is invasion by non-indigenous species, several of which already have had catastrophic impacts on property values, the fisheries, shipping, and tourism industries, and continue to threaten the survival of native species and wetland ecosystems. The Great Lakes Invasives Network is a consortium (20 institutions) of herbaria and zoology museums from among the Great Lakes states of Minnesota, Wisconsin, Illinois, Indiana, Michigan, Ohio, and New York created to better document the occurrence of selected non-indigenous species and their congeners in space and time by imaging and providing online access to the information on the specimens of the critical organisms. The list of non-indigenous species (1 alga, 42 vascular plants, 22 fish, and 13 mollusks) to be digitized was generated by conducting a query of all fish, plants, algae, and mollusks present in the database of GLANSIS – the Great Lakes Aquatic Nonindigenous Species Information System – maintained by the National Oceanic and Atmospheric Administration (NOAA). The network consists of collections at 20 institutions, including 4 of the 10 largest herbaria in North America, each of which curates 1-7 million specimens (NY, F, MICH, and WIS). Eight of the nation’s largest zoology museums are also represented, several of which (e.g., Ohio State and U of Minnesota) are internationally recognized for their fish and mollusk collections. Each genus includes at least one species that is considered a Great Lakes non-indigenous taxon – several have many, whereas others have congeners on “watchlists”, meaning that they have not arrived in the Great Lakes Basin yet, but have the potential to do so, especially in light of human activity and climate change. Because the introduction and spread of these species, their close relatives, and hybrids into the region is known to have occurred almost entirely from areas in North America outside of the Basin, our effort will include non-indigenous specimens collected from throughout North America. Digitized specimens of Great Lakes non-indigenous species and their congeners will allow for more accurate identification of invasive species and hybrids from their non-invasive relatives by a wider audience of end users. The metadata derived from digitized specimens of Great Lakes non-indigenous species and their congeners will help biologists to track, monitor, and predict the spread of invasive species through space and time, especially in the face of a more rapidly changing climate in the upper Midwest. All together consortium members will digitize >2 million individual specimens from >860,000 sheets/lots of non-indigenous species and their congeneric taxa. Data and metadata are uploaded to the Great Lakes Invasives Network, a Symbiota portal (GreatLakesInvasvies.org), and ingested by the National Resource for Advancing Digitization of Biodiversity Collections (ADBC) (iDigBio.org) national resource. Several initiatives are already in place to alert citizens to the dangers of spreading aquatic invasive species among our nation's waterways, but this project is developing complementary scientific and educational tools for scientists, students, wildlife officers, teachers, and the public who have had little access to images or data derived directly from preserved specimens of invasive species collected over the past three centuries.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.003 | 0.004 |
| Scholarly communication | 0.000 | 0.003 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".