The Integration of Genome Mining, Comparative Genomics, and Functional Genetics for Biosynthetic Gene Cluster Identification
Bibliographic record
Abstract
Antimicrobial resistance is a worldwide health crisis for which new antibiotics are needed. One strategy for antibiotic discovery is identifying unique antibiotic biosynthetic gene clusters that may produce novel compounds. The aim of this study was to demonstrate how an integrated approach that combines genome mining, comparative genomics, and functional genetics can be used to successfully identify novel biosynthetic gene clusters that produce antimicrobial natural products. Secondary metabolite clusters of an antibiotic producer are first predicted using genome mining tools, generating a list of candidates. Comparative genomic approaches are then used to identify gene suites present in the antibiotic producer that are absent in closely related non-producers. Gene sets that are common to the two lists represent leading candidates, which can then be confirmed using functional genetics approaches. To validate this strategy, we identified the genes responsible for antibiotic production in Pantoea agglomerans B025670, a strain identified in a large-scale bioactivity survey. The genome of B025670 was first mined with antiSMASH, which identified 24 candidate regions. We then used the comparative genomics platform, EDGAR, to identify genes unique to B025670 that were not present in closely related strains with contrasting antibiotic production profiles. The candidate lists generated by antiSMASH and EDGAR were compared with standalone BLAST. Among the common regions was a 14 kb cluster consisting of 14 genes with predicted enzymatic, transport, and unknown functions. Site-directed mutagenesis of the gene cluster resulted in a reduction in antimicrobial activity, suggesting involvement in antibiotic production. An integrated approach that combines genome mining, comparative genomics, and functional genetics yields a powerful, yet simple strategy for identifying potentially novel antibiotics.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".