DAPPLE 2: a Tool for the Homology-Based Prediction of Post-Translational Modification Sites
Bibliographic record
Abstract
The post-translational modification of proteins is critical for regulating their function. Although many post-translational modification sites have been experimentally determined, particularly in certain model organisms, experimental knowledge of these sites is severely lacking for many species. Thus, it is important to be able to predict sites of post-translational modification in such species. Previously, we described DAPPLE, a tool that facilitates the homology-based prediction of one particular post-translational modification, phosphorylation, in an organism of interest using known phosphorylation sites from other organisms. Here, we describe DAPPLE 2, which expands and improves upon DAPPLE in three major ways. First, it predicts sites for many post-translational modifications (20 different types) using data from several sources (15 online databases). Second, it has the ability to make predictions approximately 2-7 times faster than DAPPLE depending on the database size and the organism of interest. Third, it simplifies and accelerates the process of selecting predicted sites of interest by categorizing them based on gene ontology terms, keywords, and signaling pathways. We show that DAPPLE 2 can successfully predict known human post-translational modification sites using, as input, known sites from species that are either closely (e.g., mouse) or distantly (e.g., yeast) related to humans. DAPPLE 2 can be accessed at http://saphire.usask.ca/saphire/dapple2 .
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".