Polysaccharide utilization loci-driven enzyme discovery reveals BD-FAE: a bifunctional feruloyl and acetyl xylan esterase active on complex natural xylans
Bibliographic record
Abstract
BACKGROUND: Nowadays there is a strong trend towards a circular economy using lignocellulosic biowaste for the production of biofuels and other bio-based products. The use of enzymes at several stages of the production process (e.g., saccharification) can offer a sustainable route due to avoidance of harsh chemicals and high temperatures. For novel enzyme discovery, physically linked gene clusters targeting carbohydrate degradation in bacteria, polysaccharide utilization loci (PULs), are recognized 'treasure troves' in the era of exponentially growing numbers of sequenced genomes. RESULTS: We determined the biochemical properties and structure of a protein of unknown function (PUF) encoded within PULs of metagenomes from beaver droppings and moose rumen enriched on poplar hydrolysate. The corresponding novel bifunctional carbohydrate esterase (CE), now named BD-FAE, displayed feruloyl esterase (FAE) and acetyl esterase activity on simple, synthetic substrates. Whereas acetyl xylan esterase (AcXE) activity was detected on acetylated glucuronoxylan from birchwood, only FAE activity was observed on acetylated and feruloylated xylooligosaccharides from corn fiber. The genomic contexts of 200 homologs of BD-FAE revealed that the 33 closest homologs appear in PULs likely involved in xylan breakdown, while the more distant homologs were found either in alginate-targeting PULs or else outside PUL contexts. Although the BD-FAE structure adopts a typical α/β-hydrolase fold with a catalytic triad (Ser-Asp-His), it is distinct from other biochemically characterized CEs. CONCLUSIONS: The bifunctional CE, BD-FAE, represents a new candidate for biomass processing given its capacity to remove ferulic acid and acetic acid from natural corn and birchwood xylan substrates, respectively. Its detailed biochemical characterization and solved crystal structure add to the toolbox of enzymes for biomass valorization as well as structural information to inform the classification of new CEs.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".