Deciphering a cryptic minefield: A guide to Cryptosporidium gp60 subtyping
Bibliographic record
Abstract
For 25 years, analysis of the gp60 gene has been the cornerstone of Cryptosporidium subtyping, particularly for Cryptosporidium hominis and Cryptosporidium parvum , during population-based and epidemiological studies. This gene, which encodes a 60 kDa glycoprotein, is highly polymorphic with several variable features that make it particularly useful for differentiating within Cryptosporidium species. However, while this variability has proven useful for subtyping, it has on occasion resulted in alternative interpretations, and descriptions of novel and unusual features have been added to the nomenclature system, resulting in inconsistency and confusion. The components of the gp60 gene sequence used in the nomenclature that are discussed here include “R” repeats, “r” repeats, alphabetical suffixes, “variant” designations, and the use of the Greek alphabet as a family designation. As the subtyping scheme has expanded over the years, its application to different Cryptosporidium species has also made the scheme more complex. For example, key features may be absent, such as the typical TCA/TCG/TCT serine microsatellite that forms a major part of the nomenclature in C. hominis and C. parvum . As is to be expected in such a variable gene, different primer sets have been developed for the amplification of the gp60 in various species and these have been collated. Here we bring together all the current components of gp60 , including a guide to the nomenclature in various species, software to assist in analysing sequences, and links to useful reference resources with an aim to promote standardisation of this subtyping tool. • Provision of recommended, standardised rules for gp60 nomenclature. • Description of features underlying the nomenclature in different Cryptosporidium spp. • Pitfalls and historical alternative interpretations are highlighted. • Resources to assist the community in gp60 subtyping.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.008 |
| Meta-epidemiology (narrow) | 0.003 | 0.003 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.006 | 0.004 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.003 | 0.005 |
| Open science | 0.004 | 0.003 |
| Research integrity | 0.003 | 0.006 |
| Insufficient payload (model declined to judge) | 0.064 | 0.094 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".