MétaCan
Menu
Back to cohort
Record W4238173086 · doi:10.32920/14638233

Plasma Peptidome

2021· preprint· en· W4238173086 on OpenAlexafffund
Jaimie Dufresne, Pete Bowden, Thanusi Thavarajah, Angelique Florentinus-Mefailoski, Zhuo Zhen Chen, Monika Tucholska, Tenzin Norzin, Margaret T. Ho, Morla Phan, Nargiz Mohamed, Amir Ravandi, Eric J. Stanton, Arthur S. Slutsky, Claúdia C. dos Santos, Alexander D. Romaschin, John C. Marshall, Christina Addison, Shawn Malone, Daren K. Heyland, Philip Scheltens, Joep Killestein, Charlotte E. Teunissen, Eleftherios P. Diamandis, K. W. Michael Siu, John Marshall

Bibliographic record

Venuenot available
Typepreprint
Languageen
FieldChemistry
TopicAdvanced Proteomics Techniques and Applications
Canadian institutionsUniversity of WindsorMount Sinai HospitalSt. Michael's HospitalMcMaster UniversityKingston General HospitalOttawa HospitalUniversity of ManitobaSt. Boniface HospitalClinical Evaluation Research UnitUniversity of TorontoToronto Metropolitan University
FundersOntario Institute for Cancer Research
KeywordsChemistryChromatographyElectrospray ionizationFormic acidMass spectrometryEx vivoTandem mass spectrometryHuman plasmaProtein mass spectrometryPeptideElectrosprayBiochemistryIn vitro

Abstract

fetched live from OpenAlex

Background It may be possible to discover new diagnostic or therapeutic peptides or proteins from blood plasma using LC–ESI–MS/MS to identify, quantify and compare the statistical distributions of peptides cleaved ex vivo from plasma samples from different clinical populations. Methods A systematic method for the organic fractionation of plasma peptides was applied to identify and quantify the endogenous tryptic peptides from human plasma from multiple institutions by C18 HPLC followed nano electrospray ionization and tandem mass spectrometry (LC–ESI–MS/MS) with a linear quadrupole ion trap. The endogenous tryptic peptides, or tryptic phospho peptides (i.e. without exogenous digestion), were extracted in a mixture of organic solvent and water, dried and collected by preparative C18. The tryptic peptides from 6 institutions with 12 different disease and normal EDTA plasma populations, alongside ice cold controls for pre-analytical variation, were characterized by mass spectrometry. Each patient plasma was precipitated in 90% acetonitrile and the endogenous tryptic peptides extracted by a stepwise gradient of increasing water and then formic acid resulting in 10 sub-fractions. The fractionated peptides were manually collected over preparative C18 and injected for 1508 LC–ESI–MS/MS experiments analyzed in SQL Server R. Results Peptides that were cleaved in human plasma by a tryptic activity ex vivo provided convenient and sensitive access to most human proteins in plasma that show differences in the frequency or intensity of proteins observed across populations that may have clinical significance. Combination of step wise organic extraction of 200 μL of plasma with nano electrospray resulted in the confident identification and quantification ~ 14,000 gene symbols by X!TANDEM that is the largest number of blood proteins identified to date and shows that you can monitor the ex vivo proteolysis of most human proteins, including interleukins, from blood. A total of 15,968,550 MS/MS spectra ≥ E4 intensity counts were correlated by the SEQUEST and X!TANDEM algorithms to a federated library of 157,478 protein sequences that were filtered for best charge state (2+ or 3+) and peptide sequence in SQL Server resulting in 1,916,672 distinct best-fit peptide correlations for analysis with the R statistical system. SEQUEST identified some 140,054 protein accessions, or some ~ 26,000 gene symbols, proteins or loci, with at least 5 independent correlations. The X!TANDEM algorithm made at least 5 best fit correlations to more than 14,000 protein gene symbols with p-values and FDR corrected q-values of ~ 0.001 or less. Log10 peptide intensity values showed a Gaussian distribution from E8 to E4 arbitrary counts by quantile plot, and significant variation in average precursor intensity across the disease and controls treatments by ANOVA with means compared by the Tukey–Kramer test. STRING analysis of the top 2000 gene symbols showed a tight association of cellular proteins that were apparently present in the plasma as protein complexes with related cellular components, molecular functions and biological processes. Conclusions The random and independent sampling of pre-fractionated blood peptides by LC-ESI-MS/MS with SQL Server-R analysis revealed the largest plasma proteome to date and was a practical method to quantify and compare the frequency or log10 intensity of individual proteins cleaved ex vivo across populations of plasma samples from multiple clinical locations to discover treatment-specific variation using classical statistics suitable for clinical science. It was possible to identify and quantify nearly all human proteins from EDTA plasma and compare the results of thousands of LC–ESI–MS/MS experiments from multiple clinical populations using standard database methods in SQL Server and classical statistical strategies in the R data analysis system.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.001
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesInsufficient payload (model declined to judge)
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Other · Consensus signal: none
Teacher disagreement score0.990
Threshold uncertainty score0.032

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0000.001
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0010.002
Science and technology studies0.0000.000
Scholarly communication0.0010.001
Open science0.0000.001
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0100.008

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.018
GPT teacher head0.285
Teacher spread0.268 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designNot applicable
Domainnot available
GenreOther

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2021
Admission routes2
Has abstractyes

Explore more

Same topicAdvanced Proteomics Techniques and ApplicationsFrench-language works237,207