DigestR an open-source software tool for visualizing LC-MS proteomics data resulting from natural protein catabolism
Bibliographic record
Abstract
Abstract Protein catabolism is an essential biological function supported by every living organism. Although liquid chromatography mass spectrometry proteomics has advanced considerably over the past decade, protein catabolism in natural systems is still difficult to study. One reason for this is the lack of software tools designed specifically for decoding the complex mixtures of peptides that result from in vivo protein digestion. To address this, we developed DigestR, an open-source software tool designed specifically for the analysis of LC-MS proteomics data. DigestR allows users to visualize naturally occurring peptides and align them to a reference proteome at display them at either a proteome-wide and protein-specific level. These visualization tools allow users to track the patters of peptides occurring in natural systems and map naturally-occurring proteolytic cut sites. To demonstrate these functions, we used DigestR to analyze a mixture of peptides resulting from the in vitro digestion of human hemoglobin and bovine albumin with a cocktail of well characterized proteases. As expected, DigestR correctly identified both the proteins involved and the proteolytic cut sites produced by our protease cocktail. We then used DigestR to analyze the complex semi-ordered hemoglobin digestion pathway used by the malaria parasite Plasmodium falciparum. We show that DigestR successfully identified the proteolytic cut sites linked to the Plasmepsins, a protease known to be involved in hemoglobin digestion by the parasite. Collectively, these findings show that DigestR can be used to help visualize and interpret the complex mixtures of peptides occurring through in vivo protein catabolism. DigestR can be downloaded from www.lewisresearchgroup.org/software .
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.005 |
| Meta-epidemiology (narrow) | 0.004 | 0.002 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.005 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.003 |
| Open science | 0.004 | 0.004 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.058 | 0.024 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".