Characterization of novel lignocellulose-degrading enzymes from the porcupine microbiome using synthetic metagenomics
Bibliographic record
Abstract
Plant cell walls are composed of cellulose, hemicellulose, and lignin, collectively known as lignocellulose. Microorganisms degrade lignocellulose to liberate sugars to meet metabolic demands. Using a metagenomic sequencing approach, we previously demonstrated that the microbiome of the North American porcupine (Erethizon dorsatum) is replete with genes that could encode lignocellulose-degrading enzymes. Here, we report the identification, synthesis and partial characterization of four novel genes from the porcupine microbiome encoding putative lignocellulose-degrading enzymes: β-glucosidase, α-L-arabinofuranosidase, β-xylosidase, and endo-1,4-β-xylanase. These genes were identified via conserved catalytic domains associated with cellulose- and hemicellulose-degradation. Phylogenetic trees were created for each of these putative enzymes to depict genetic relatedness to known enzymes. Candidate genes were synthesized and cloned into plasmid expression vectors for inducible protein expression and secretion. The putative β-glucosidase fusion protein was efficiently secreted but did not permit Escherichia coli (E. coli) to use cellobiose as a sole carbon source, nor did the affinity purified enzyme cleave p-Nitrophenyl β-D-glucopyranoside (p-NPG) substrate in vitro over a range of physiological pH levels (pH 5-7). The putative hemicellulose-degrading β-xylosidase and α-L-arabinofuranosidase enzymes also lacked in vitro enzyme activity, but the affinity purified endo-1,4-β-xylanase protein cleaved a 6-chloro-4-methylumbelliferyl xylobioside substrate in acidic and neutral conditions, with maximal activity at pH 7. At this optimal pH, KM, Vmax, and kcat were determined to be 32.005 ± 4.72 μM, 1.16x10-5 ± 3.55x10-7 M/s, and 94.72 s-1, respectively. Thus, our pipeline enabled successful identification and characterization of a novel hemicellulose-degrading enzyme from the porcupine microbiome. Progress towards the goal of introducing a complete lignocellulose-degradation pathway into E. coli will be accelerated by combining synthetic metagenomic approaches with functional metagenomic library screening, which can identify novel enzymes unrelated to those found in available databases.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".