Bibliographic record
Abstract
Natural products (NPs) are compounds that are produced by living things, like bacteria. These compounds are encoded by what are known as biosynthetic gene clusters (BGCs). A BGC is a set of genes that when all expressed, work together to form a particular NP. NPs are of extreme importance as they are used for many different medical applications, particularly as antibiotics. Since bacteria must evolve quickly to compete against one another for resources, it makes sense to turn to these organisms to find a solution for bacterial infections. The problem is that pathogenic bacteria have been evolving methods to fight antibiotics faster than we have been able to discover new ones. There is a great need to improve the natural product drug discovery pipeline. One way to do this is to change the conditions microbes grow in to have them express previously silent BGCs. In order to do this effectively, it is beneficial to have a strong baseline for molecules produced under standard conditions. Pseudoalteromonas is a relatively underexplored genus of marine bacteria. Thus, my goal is to catalogue the natural products produced by Pseudoalteromonas under standard culture conditions using Mass Spectrometry analysis. The starting point for this so-called Library Project has involved identifying existing natural product databases. The databases that have been explored can be classified into four categories that are also reflective of the order in which they will be used. These categories for the databases are: databases for culturing information, databases to catalog known natural products, mass spectrometry (MS) analysis tool-based databases, and databases that will associate novel natural products with BGCs. Now that the appropriate databases have been identified, the next step is to begin culturing and extracting from these microbes, to then run and analyze mass spectrums of these extracts.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.005 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.018 | 0.019 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".