Dereplication, Residual Complexity, and Rational Naming: The Case of the <i>Actaea</i> Triterpenes
Bibliographic record
Abstract
The genus Actaea (including Cimicifuga) has been the source of ∼200 cycloartane triterpenes. While they are major bioactive constituents of complementary and alternative medicines, their structural similarity is a major dereplication problem. Moreover, their trivial names seldom indicate the actual structure. This project develops two new tools for Actaea triterpenes that enable rapid dereplication of more than 170 known triterpenes and facilitates elucidation of new compounds. A predictive computational model based on classification binary trees (CBTs) allows in silico determination of the aglycone type. This tool utilizes the Me (1)H NMR chemical shifts and has potential to be applicable to other natural products. Actaea triterpene dereplication is supported by a new systematic naming scheme. A combination of CBTs, (1)H NMR deconvolution, characteristic (1)H NMR signals, and quantitative (1)H NMR (qHNMR) led to the unambiguous identification of minor constituents in residually complex triterpene samples. Utilizing a 1.7 mm cryo-microprobe at 700 MHz, qHNMR enabled characterization of residual complexity at the 10-20 μg level in a 1-5 mg sample. The identification of five co-occurring minor constituents, belonging to four different triterpene skeleton types, in a repeatedly purified natural product emphasizes the critical need for the evaluation of residual complexity of reference materials, especially when used for biological assessment.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".