Deciphering a cryptic minefield: A guide to Cryptosporidium gp60 subtyping
Bibliographic record
Abstract
For 25 years, analysis of the gp60 gene has been the cornerstone of Cryptosporidium subtyping, particularly for Cryptosporidium hominis and Cryptosporidium parvum , during population-based and epidemiological studies. This gene, which encodes a 60 kDa glycoprotein, is highly polymorphic with several variable features that make it particularly useful for differentiating within Cryptosporidium species. However, while this variability has proven useful for subtyping, it has on occasion resulted in alternative interpretations, and descriptions of novel and unusual features have been added to the nomenclature system, resulting in inconsistency and confusion. The components of the gp60 gene sequence used in the nomenclature that are discussed here include “R” repeats, “r” repeats, alphabetical suffixes, “variant” designations, and the use of the Greek alphabet as a family designation. As the subtyping scheme has expanded over the years, its application to different Cryptosporidium species has also made the scheme more complex. For example, key features may be absent, such as the typical TCA/TCG/TCT serine microsatellite that forms a major part of the nomenclature in C. hominis and C. parvum . As is to be expected in such a variable gene, different primer sets have been developed for the amplification of the gp60 in various species and these have been collated. Here we bring together all the current components of gp60 , including a guide to the nomenclature in various species, software to assist in analysing sequences, and links to useful reference resources with an aim to promote standardisation of this subtyping tool. • Provision of recommended, standardised rules for gp60 nomenclature. • Description of features underlying the nomenclature in different Cryptosporidium spp. • Pitfalls and historical alternative interpretations are highlighted. • Resources to assist the community in gp60 subtyping.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".