Additional file 1 of Ultra-sensitive isotope probing to quantify activity and substrate assimilation in microbiomes
Bibliographic record
Abstract
Additional file 1. Supplementary Results and Discussion. Differences when using Calis-p for quantification of natural carbon isotope ratios (Protein-SIF) versus labeling with heavy isotopes (Protein-SIP). Figure S1. Comparison of the previously published version of Calis-p (v0.0) [1] with the new version (v2.1) in regards to their accuracy for quantifying natural carbon isotope abundances (Protein-SIF) of species in microbial community samples. The absolute difference between δ13C values of individual species in mock communities determined with the Fast Fourier Transforms algorithm (i.e. default model) and isotope ratio mass spectrometry (IRMS) of the corresponding pure cultures is shown (method details in [1]). Five mock community datasets with a total of 32 species and strains were analyzed. For 20 species, the δ13C values were known from IRMS performed on pure cultures. For these species, the δ13C values were determined. Each dataset contained different amounts of data. The absolute difference between δ13C values obtained via protein-SIF and IRMS was calculated and sorted according to how many peptides or PSMs were available for SIF calculation by Calis-p after filtering the peptides. The plots give the absolute differences for different ranges of peptide and PSM numbers used for SIF calculation. Additionally, for Calis-p v2.1 plots are also shown using the median as the center statistic for δ13C value calculations. Figure S2. Number of peptide spectral matches (PSMs) identified at different 13C label percentages using six different peptide identification strategies. For the spike-in samples E. coli cells labeled at different percentages with 13C6 glucose were mixed into a mock community consisting of 32 species of bacteria, archaea, eukaryota and bacteriophages (UNEVEN community from Kleiner et al. (2017) [6]), which also contained unlabeled E. coli cells. Labeled and unlabeled E. coli cells in the spike-in sample were at a 1:1 ratio.Three biological replicates were analyzed for each label percentage. Peptides were identified using the SEQUEST HT Node in Proteome Discoverer (version 2.2.) with six different strategies to account for the mass shifts caused by addition of heavy atoms. Standard search: no dynamic modifications to account for addition of label; Open search: the precursor mass tolerance was set to 20 Da allowing for the potential addition of 20 neutrons (e.g. 13C atoms) in a peptide; Dynamic modifications: allowing for up to three dynamic modifications each of two custom peptide modifications adding a 1 neutron mass shift and a 2 neutron mass shift (up to 9 neutrons in total per peptide); Modifications on termini: six dynamic modifications were set up that were restricted to either the C or the N-terminus of the peptide. The modifications account for mass shifts of 1 to 6 neutrons and depending on the search strategy the low mass shifts (1, 2 and 3 neutrons) were set up as modifications on the C or the N-Terminus or low and high mass shift modifications were distributed between both termini. Each modification can only be added to a terminus once. This strategy allows for a total of 21 neutron additions to a peptide. Figure S3. Assimilation of 13C in a mock labeling experiment created by mixing labeled and unlabeled cells of E. coli in ratios corresponding to 1/100 to 1 generations of growth. Cells were labeled with 1% 13C glucose (top row) and 10% 13C glucose (bottom row). Cells were labeled with CC2-labeled glucose (left column) and fully (C1-6) labeled glucose. Peptide identification-based detection of excess neutron masses was most sensitive, able to reproducibly detect growth for 1/100 generation, but not quantitative. Use of 1% 13C glucose with a single atom labeled resulted in the best quantitation (R2 >0.996) and label recovery (92%). The detailed data for this figure can be found in Suppl. table S4. Figure S4. Comparison of existing Protein-SIP approaches with Calis-p. Three Protein-SIP approaches (SIPPER, MetaProSIP, and Calis-p) were compared using optimal parameters chosen by an expert operator for each tool. We used the datasets from the mock community with added 1%, 5%, or 10% 13C labeled E.coli and as a control the mock community with only unlabeled E. coli. The resulting 13C atom % output values from SIPPER, MetaProSIP, and Calis-p for each dataset were filtered such that organisms for which 13C content was quantified for 9 or more peptides were considered for plotting. The black line in each plot represents the expected 13C atom % value for all unlabeled peptides. The red line in each plot represents the expected 13C atom % value for the E. coli peptides from each respective dataset after accounting for the expected mixing of labeled and unlabeled E. coli peptides in the sample. The number of peptides used for each box and other details can be found in tables S7-S10. Figure S5. Strong differences in heavy water incorporation in intestinal microbiota in response to different diet conditions. 63 species isolated from human intestinal microbiota were grown together in triplicates in either a high fiber or high protein medium in the presence of unlabeled water or water with either 25% 2H or 18O [7]. Calis-p based stable isotope ratios are shown for all species combined per replicate. Each box shows the data for all peptides for one replicate. The ‘n’ gives the number of peptides that passed the Calis-p quality filters. The solid red lines indicate the average quartiles for the three replicates, the dashed red line the average median for the three replicates. Statistically significant differences are indicated with ‘*’ based on Student’s t-test on the medians of replicates at p < 0.05. The script (Suppl. file S1, also available here https://github.com/yihualiud/Data-analysis-of-Calis-p-protein-SIP-results/blob/main/Heavy Water_SIP_Calisp_final.R ) and Calis-p output peptide datasets (Suppl. Datasets S1 and S2) used to generate this figure and Figure 7 have been provided as an example. Supplementary Methods. How Calis-p calculates center statistics for a species based on its peptides. Figure S6. Explanation of procedure how clumpiness of label is estimated.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.022 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.003 | 0.005 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.807 | 0.212 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".