A network-based maximum link approach towards MS identifies potentially important roles for undetected ARRB1/2 and ACTB in liver cancer progression
Bibliographic record
Abstract
Hepatocellular Carcinoma (HCC) ranks among the deadliest of cancers and has a complex etiology. Proteomics analysis using iTRAQ provides a direct way to analyse perturbations in protein expression during HCC progression from early- to late-stage but suffers from consistency and coverage issues. Appropriate use of network-based analytical methods can help to overcome these issues. We built an integrated and comprehensive Protein-Protein Interaction Network (PPIN) by merging several major databases. Additionally, the network was filtered for GO coherent edges. Significantly differential genes (seeds) were selected from iTRAQ data and mapped onto this network. Undetected proteins linked to seeds (linked proteins) were identified and functionally characterised. The process of network cleaning provides a list of higher quality linked proteins, which are highly enriched for similar biological process gene ontology terms. Linked proteins are also enriched for known cancer genes and are linked to many well-established cancer processes such as apoptosis and immune response. We found that there is an increased propensity for known cancer genes to be found in highly linked proteins. Three highly-linked proteins were identified that may play an important role in driving HCC progression - the G-protein coupled receptor signalling proteins, ARRB1/2 and the structural protein beta-actin, ACTB. Interestingly, both ARRB proteins evaded detection in the iTRAQ screen. ACTB was not detected in the original dataset derived from Mascot but was found to be strongly supported when we re-ran analysis using another protein detection database (Paragon). Identification of linked proteins helps to partially overcome the coverage issue in shotgun proteomics analysis. The set of linked proteins are found to be enriched for cancer-specific processes, and more likely so if they are more highly linked. Additionally, a higher quality linked set is derived if network-cleaning is performed prior. This form of network-based analysis complements the cluster-based approach, and can provide a larger list of proteins on which to perform functional analysis, as well as for biomarker identification.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.005 | 0.003 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".