Mining and rational design of psychrophilic catalases using metagenomics and deep learning models
Bibliographic record
Abstract
A complete catalase-encoding gene, designated soiCat1, was obtained from soil samples via metagenomic sequencing, assembly, and gene prediction. soiCat1 showed 73% identity to a catalase-encoding gene of Mucilaginibacter rubeus strain P1, and the amino acid sequence of soiCAT1 showed 99% similarity to the catalase of a psychrophilic bacterium, Pedobacter cryoconitis. soiCAT1 was identified as a psychrophilic enzyme due to the low optimum temperature predicted by the deep learning model Preoptem, which was subsequently validated through analysis of enzymatic properties. Experimental results showed that soiCAT1 has a very narrow range of optimum temperature, with maximal specific activity occurring at the lowest test temperature (4 °C) and decreasing with increasing reaction temperature from 4 to 50 °C. To rationally design soiCAT1 with an improved temperature range, soiCAT1 was engineered through site-directed mutagenesis based on molecular evolution data analyzed through position-specific amino acid possibility calculation. Compared with the wild type, one mutant, soiCAT1S205K, exhibited an extended range of optimum temperature ranging from 4 to 20 °C. The strategies used in this study may shed light on the mining of genes of interest and rational design of desirable proteins. • Numerous putative catalases were mined from soil samples via metagenomics. • A complete sequence encoding a psychrophilic catalase was obtained. • A mutant psychrophilic catalase with an extended range of optimum temperature was engineered through site-directed mutagenesis.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".