Integrating Next-Generation Genomic Sequencing and Mass Spectrometry To Estimate Allele-Specific Protein Abundance in Human Brain
Bibliographic record
Abstract
Gene expression contributes to phenotypic traits and human disease. To date, comparatively less is known about regulators of protein abundance, which is also under genetic control and likely influences clinical phenotypes. However, identifying and quantifying allele-specific protein abundance by bottom-up proteomics is challenging since single nucleotide variants (SNVs) that alter protein sequence are not considered in standard human protein databases. To address this, we developed the GenPro software and used it to create personalized protein databases (PPDs) to identify single amino acid variants (SAAVs) at the protein level from whole exome sequencing. In silico assessment of PPDs generated by GenPro revealed only a 1% increase in tryptic search space compared to a direct translation of all human transcripts and an equivalent search space compared to the UniProtKB reference database. To identify a large unbiased number of SAAV peptides, we performed high-resolution mass spectrometry-based proteomics for two human post-mortem brain samples and searched the collected MS/MS spectra against their respective PPD. We found an average of ∼117 000 unique peptides mapping to ∼9300 protein groups for each sample, and of these, 977 were unique variant peptides. We found that over 400 reference and SAAV peptide pairs were, on average, equally abundant in human brain by label-free ion intensity measurements and confirmed the absolute levels of three reference and SAAV peptide pairs using heavy labeled peptides standards coupled with parallel reaction monitoring (PRM). Our results highlight the utility of integrating genomic and proteomic sequencing data to identify sample-specific SAAV peptides and support the hypothesis that most alleles are equally expressed in human brain.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".