Single Cell Multi‐Omics Reveals Rare Biosynthetic Cell Types in the Medicinal Tree <scp> <i>Camptotheca acuminata</i> </scp>
Bibliographic record
Abstract
Camptotheca acuminata Decne is a woody medicinal tree that produces over a hundred bioactive compounds, including camptothecin, which has been used as the starting material to semi-synthesise many leading anticancer drugs (Lorence and Nessler 2004). Camptothecin and its derivatives are potent inhibitors of DNA topoisomerase I and are widely used for the treatment of lung, cervical, ovarian and colon cancers. Camptothecin biosynthesis in C. acuminata involves complex catalytic steps, most of which remain undeciphered. In this pathway, tryptamine and secologanic acid are coupled, leading to strictosidinic acid. The formation of strictosidinic acid is catalysed by strictosidine/strictosidine acid synthase enzymes (STR) (Figure 1a). While a biosynthetic route for the conversion of the indole ring to the quinoline ring has been proposed, most of the underlying biosynthetic genes have yet to be identified (Figure 1a) (Sadre et al. 2016). In addition, the cell type specificity of this pathway also remains undescribed. Here, we generated a single cell multiome (RNA-seq and Assay for Transposase Accessible Chromatin by sequencing [ATAC-seq] from the same nuclei) to probe the cell type specificity of camptothecin biosynthetic genes. We performed organ-level metabolite profiling on key biosynthetic intermediates (Figure S1) across multiple C. acuminata organs and found that camptothecin was detected across all organs tested (Figure S1). Young leaf was chosen for single cell omics experiments (Table S1) due to the relative ease of nuclei isolation (Figure S2a–c), which permits the simultaneous profiling of gene expression and chromatin accessibility from the same nucleus using the 10× Genomics Multiome Kit. To aid downstream bioinformatic analyses, we produced a new version of the C. acuminata genome assembly and annotation (Tables S2–S5, Figure S2d), in which we used Nanopore full-length cDNAs to refine gene model predictions. For the gene expression assay of the multi-omics experiment, we obtained gene expression and accessible chromatin profiles for 4012 high quality nuclei and 26,074 expressed genes (Table S6, Figure S3), which includes previously characterised biosynthetic genes such as Secologanic Acid Synthase (SLAS), Tryptophan Decarboxylase (TDC) and Strictosidine/Strictosidinic Acid Synthase (STR) (Table S7) and previously reported camptothecin decorating enzymes, Camptothecin-11-Hydroxylase (CPT11H) and Camptothecin-10-Hydroxy-Methyltransferase (CPT10-OMT) (Nguyen et al. 2021; Salim et al. 2018). Using unsupervised clustering and previously established leaf cell type marker genes in Arabidopsis (Table S8) (Kim et al. 2021), we identified major cell types of the leaf (i.e., mesophyll, epidermis and vasculature) (Figure 1b and Figure S4a). Amongst the expressed biosynthetic genes (Table S7), SLAS and TDC were expressed in the vasculature, while STR genes were highly specific to a rare cell type that accounts for only 4.4% of all leaf cells, which we termed “STR+ cells” (Figure 1c). Expression of key biosynthetic step(s) in a rare cell type is reminiscent of the restriction of vinca alkaloid biosynthetic genes to the rare idioblast cells in Catharanthus roseus (Li et al. 2023). We next performed a joint RNA-ATAC analysis by matching the cell barcodes from both assays (Figure S4b) and thus transferring the cell type identity from the RNA-seq assay to the ATAC-seq assay. The ATAC-seq signals were strongly enriched at transcriptional start and end sites (Figure S5a,b) and highly enriched at ATAC-seq peaks (Figure S5b,c, Table S9). Amongst 48,456 ATAC-seq peaks, 199 (0.41%) were detected as STR+ marker peaks (ATAC-seq peaks that were specifically enriched in STR+ cells) (Figure 1d). We found that amongst the 121 genes that were within 2-kb of a STR+ marker peak, 34 (28%) of them were most highly expressed in STR+ cells, representing 9.3-fold enrichment over the expected background ratio (Figure 1e, p < 2.2 × 10−16, Chi-squared test). The enrichment of STR+ -expressed genes suggests that STR+ marker peaks could act as cell type-specific enhancers. De novo motif enrichment on STR+ marker peaks revealed an overrepresentation of MYB binding motifs (Fornes et al. 2019) amongst these accessible chromatin regions (Figure 1f), suggesting MYB family TFs may be involved in cell type-specific expression of STR genes in C. acuminata. Taken together, our results showcased that single-cell omics techniques are powerful in detecting rare biosynthetic cell types or cell populations and the regulatory landscape within. The datasets generated by this study are valuable resources for mining new biosynthetic genes and cis-regulatory elements for camptothecin biosynthesis. C.L., C.R.B. and T.-T.T.D. designed the study. V.-H.B. performed metabolite profiling. C.L. generated single nuclei multiome libraries. L.H.C. performed cDNA sequencing. J.C.W. performed single cell library construction and quality control. B.V. performed genome assembly and data management. J.P.H. performed genome annotation. C.L. and V.-H.B. wrote the manuscript with input from all authors. This project was supported by the University of Georgia, Georgia Research Alliance (C.R.B.) and Georgia Seed Development (C.R.B.), National Science Foundation MCB-2309665 (C.R.B. and C.L.). T.-T.T.D. receives funding from NSERC Alliance Collaboration (ALLRP 571673-21) and Catalyst (ALLRP 579871-22). T.-T.T. Dang is grateful for the Michael Smith Health BC Scholar Award (SCH-2020-0401). Sequencing was performed at the University of Minnesota Genomics Center (UMGC) and UC Davis DNA Technologies and Expression Analysis Core. We thank Kathrine Mailloux for high molecular weight DNA isolation and Sophia Jones for plant care. All sequencing data associated with this study are available at the National Center for Biotechnology Institute Sequence Read Archive BioProject PRJNA1243444. Genome assembly, annotation, and Seurat objects for single cell multiome experiments are available via the digital repository figshare (https://figshare.com/s/039f397aeeae23d320e4). Custom codes can be found at https://github.com/cxli233/Camptotheca_single_cell. Figure S1: Key monoterpene indole alkaloids (MIA) of Camptotheca acuminata. Figure S2: Camptotheca acuminata v4 genome assembly and annotation and nuclei preparation for single nuclei omics. Figure S3: Characterisation of leaf single nuclei RNA-seq libraries. Figure S4: Maker genes and multimodal integration for leaf single nuclei omics datasets. Figure S5: Characterisation of leaf single nuclei ATAC-seq libraries. Table S1: List of libraries and sequencing data generated by this study. Table S2: Assembly statistics. Table S3: Benchmarking universal single-copy orthologs (BUSCO) results on the Camptotheca acuminata v4 assembly and annotation. Table S4: Repetitive sequence content in the Camptotheca acuminata v4 assembly. Table S5: Camptotheca acuminata v4 gene annotation summary. Table S6: Details of leaf single nuclei RNA-seq libraries. Table S7: Gene names of known monoterpene indole alkaloid biosynthetic genes in Camptotheca acuminata. Table S9: Leaf cell type marker genes. Table S10: Details of leaf single nuclei ATAC-seq libraries. Data S1: pbi70386-sup-0003-Supinfo1.docx. Please note: The publisher is not responsible for the content or functionality of any supporting information supplied by the authors. Any queries (other than missing content) should be directed to the corresponding author for the article.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".