Proteomic and N-Terminomic TAILS Analyses of Human Alveolar Bone Proteins: Improved Protein Extraction Methodology and LysargiNase Digestion Strategies Increase Proteome Coverage and Missing Protein Identification
Bibliographic record
Abstract
With 2129 proteins still classified by the Human Proteome Organisation Human Proteome Project (HPP) as "missing" without compelling evidence of protein existence (PE) in humans, we hypothesized that in-depth proteomic characterization of tissues that are technically challenging to access and extract would yield evidence for tissue-specific missing proteins. Paradoxically, although the skeleton is the most massive tissue system in humans, as one of the poorest characterized by proteomics, bone falls under the HPP umbrella term as a "rare tissue". Therefore, we aimed to optimize mineralized tissue protein extraction methodology and workflows for proteomic and data analyses of small quantities of healthy young adult human alveolar bone. Osteoid was solubilized by GuHCl extraction, with hydroxyapatite-bound proteins then released by ethylenediaminetetraacetic acid demineralization. A subsequent GuHCl solubilization extraction was followed by solid-phase digestion of the remaining insoluble cross-linked protein using trypsin and then 6 M urea dissolution incorporating LysC digestion. Bone extracts were digested in parallel using trypsin, LysargiNase, AspN, or GluC prior to liquid chromatography-mass spectrometry analysis. Terminal Amine Isotopic Labeling of Substrates was used to purify semitryptic peptides, identifying natural and proteolytic-cleaved neo N-termini of bone proteins. Our strategy enabled complete solubilization of the organic bone matrix leading to extensive categorization of bone proteins in different bone matrix extracts, and hence matrix compartments, for the first time. Moreover, this led to the high confidence identification of pannexin-3, a "missing protein", found only in the insoluble collagenous matrix and revealed for the first time by trypsin solid-phase digestion. We also found a singleton proteotypic peptide of another missing protein, meiosis inhibitor protein 1. We also identified 17 proteins classified in neXtprot as PE1 based on evidence other than from MS, termed non-MS PE1 proteins, including ≥9-mer proteotypic peptides of four proteins.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".