α-Glucan Recognition by a New Family of Carbohydrate-Binding Modules Found Primarily in Bacterial Pathogens
Bibliographic record
Abstract
TmPul13, a family 13 glycoside hydrolase from Thermotoga maritima, is a four-module protein having pullulanase activity; the three N-terminal modules are of unknown function while the large C-terminal module is likely the catalytic module. Dissection of the functions of the three unknown modules revealed that the 100 amino acid module at the extreme N-terminus of TmPul13 comprises a new family of carbohydrate-binding modules (CBM) that a bioinformatic analysis shows are most frequently found in pullulanase-like sequences from bacterial pathogens. Detailed binding studies of this isolated CBM, here called TmCBM41, reveals a preference for alpha-(1,4)-linked glucans, but occasional alpha-(1,6)-linked glucose residues, such as those found in pullulan, are tolerated. UV difference, isothermal titration calorimetry, and analytical ultracentrifugation binding studies suggest that maltooligosaccharides longer than four glucose residues are able to bind two TmCBM41 molecules per oligosaccharide when sugar concentrations are below the CBM concentration. This is explained in terms of an equilibrium expression involving the formation of both a 1 to 1 sugar to CBM complex and a 1 to 2 sugar to CBM complex (i.e., a CBM dimer ligated by an oligosaccharide). The presence of an alpha-(1-6) linkage in the oligosaccharide appears to prevent this phenomenon.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".