Insight into Ligand Diversity and Novel Biological Roles for Family 32 Carbohydrate-Binding Modules
Bibliographic record
Abstract
Family 32 carbohydrate-binding modules (CBM32s) are found in a diverse group of microorganisms, including archea, eubacteria, and fungi. Significantly, many members of this family belong to plant and animal pathogens where they are likely to play a key role in enzyme toxin targeting and function. Indeed, ligand targets have been shown to range from insoluble plant cell wall polysaccharides to complex eukaryotic glycans. Besides a potential direct involvement in microbial pathogenesis, CBM32s also represent an important family for the study of CBM evolution due to the wide variety of complex protein architectures that they are associated with. This complexity ranges from independent lectin-like proteins through to large multimodular enzyme toxins where they can be present in multiple copies (multimodularity). Presented here is a rigorous analysis of the evolutionary relationships between available polypeptide sequences for family 32 CBMs within the carbohydrate active enzyme database. This approach is especially helpful for determining the roles of CBM32s that are present in multiple copies within an enzyme as each module tends to cluster into groups that are associated with distinct enzyme classes. For enzymes that contain multiple copies of CBM32s, however, there are differential clustering patterns as modules can either cluster together or in very distant sections of the tree. These data suggest that enzymes containing multiple copies possess complex mechanisms of ligand recognition. By applying this well-developed approach to the specific analysis of CBM relatedness, we have generated here a new platform for the prediction of CBM binding specificity and highlight significant new targets for biochemical and structural characterization.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".