Assessment of crAssphage as a human fecal source tracking marker in the lower Great Lakes
Bibliographic record
Abstract
CrAssphage or crAss-like phage ranks as the most abundant phage in the human gut and is present in human feces-contaminated environments. Due to its high human specificity and sensitivity, crAssphage is a potentially robust source tracking indicator that can distinguish human fecal contamination from agricultural or wildlife sources. Its suitability in the Great Lakes area, one of the world's most important water systems, has not been well tested. In this study, we tested a qPCR-based quantification method using two crAssphage marker genes (ORF18-mod and CPQ_064) at Toronto recreational beaches along with their adjacent river mouths. Our results showed a 71.4 % (CPQ_064) and 100 % (ORF18-mod) human sensitivity for CPQ_064 and ORF18-mod, and a 100 % human specificity for both marker genes. CrAssphage was present in 57.7 % or 71.2 % of environmental water samples, with concentrations ranging from 1.45 to 5.14 log10 gene copies per 100 mL water. Though concentrations of the two marker genes were strongly correlated, ORF18-mod features a higher human sensitivity and higher positive detection rates in environmental samples. Quantifiable crAssphage was mostly present in samples collected in June and July 2021 associated with higher rainfall. In addition, rivers had more frequent crAssphage presence and higher concentrations than their associated beaches, indicating more frequent and greater human fecal contamination in the rivers. However, crAssphage was more correlated with E. coli and Enterococcus at the beaches than in the rivers, suggesting human fecal sources may be more predominant in driving the increases in E. coli and Enterococcus at the beaches when impacted by river plumes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".