Synthetic plasma pool cohort correction for affinity-based proteomics datasets allows multiple study comparison
Bibliographic record
Abstract
Proteomics stands as the crucial link between genomics and human diseases. Quantitative proteomics provides detailed insights into protein levels, enabling differentiation between distinct phenotypes. OLINK, a biotechnology company from Uppsala, Sweden, offers a targeted, affinity-based protein measurement method called Target 96, which has become prominent in the field of proteomics. The SCALLOP consortium, for instance, contains data from over 70.000 individuals across 45 independent cohort studies, all sampled by OLINK. However, when independent cohorts want to collaborate and quantitatively compare their target 96 protein values, it is currently advised to include 'identical biological bridging' samples in each sampling run to perform a reference sample normalization, correcting technical variations across measurements. Such a 'biological bridging sample' approach requires each of the involved cohorts to resend their biological bridging samples to OLINK to run them all together, which is logistically challenging, costly and time-consuming. Hence alternatives are searched and an evaluation of the current state of the art exposes the need for a more robust method that allows all OLINK Target 96 studies to compare proteomics data accurately and cost-efficiently. To meet these goals we developed the Synthetic Plasma Pool Cohort Correction, the 'SPOC correction' approach, based on the use of an OLINK-composed synthetic plasma sample. The method can easily be implemented in a federated data-sharing context which is illustrated on a sepsis use case.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".