Carrier-Guided Proteome Analysis in a High Protein Background: An Improved Approach to Host Cell Protein Identification
Bibliographic record
Abstract
Many shotgun proteomics experiments are negatively influenced by highly abundant proteins, such as those measuring residual host cell proteins (HCP) amidst highly abundant recombinant biotherapeutic or plasma proteins amidst albumin and immunoglobulins. While western blotting and ELISAs can reveal the presence of specific low abundance proteins from highly abundant background proteins, mass spectrometry approaches are required to define the low abundance protein composition in these scenarios. The challenge in detecting low abundance proteins in a high protein background by standard shotgun approaches is that spectra are often not triggered on their peptides in data dependent acquisition methods but rather on the highly abundant background peptides. Here, we use tandem mass tags (TMT) to introduce a carrier proteome approach to enhance the detection of proteins, such as from residual host cell proteomes amidst a highly abundant background. Using a mixture of bovine serum albumin (BSA) and E. coli as a mock high background/low abundance target protein formulation, we demonstrate proof-of-principle experiments allowing the improved detection of target proteins amidst a high protein background. While we observed significant coisolation interference, we mitigated it by using a spike-in interference detection TMT channel. Finally, we use the approach to identify 300 residual E. coli proteins from a protein A pulldown of a human IgG antibody, demonstrating that it may be applicable to analysis of HCPs in biotherapeutic protein formulations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.002 | 0.005 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".