Methodological development of molecular endotype discovery from synovial fluid of individuals with knee osteoarthritis: the STEpUP OA Consortium
Bibliographic record
Abstract
ABSTRACT Objectives To develop and validate a pipeline for quality controlled (QC) protein data for largescale analysis of synovial fluid (SF), using SomaLogic technology. Design Knee SF and associated clinical data were from partner cohorts. SF samples were centrifuged, supernatants stored at −80 °C, then analysed by SomaScan Discovery Plex V4.1 (>7000 SOMAmers/proteins). Setting An international consortium of 9 academic and 8 commercial partners (STEpUP OA). Participants 1746 SF samples from 1650 individuals comprising OA, joint injury, healthy controls and inflammatory arthritis controls, divided into discovery (n=1045) and replication (n=701) datasets. Primary and secondary outcome measures An optimised approach to standardisation was developed iteratively, monitoring reliability and precision (comparing coefficient of variation [%CV] of ‘pooled’ SF samples between plates and correlation with prior immunoassay for 9 analytes). Pre-defined technical confounders were adjusted for (by Limma) and batch correction was by ComBat. Poorly performing SOMAmers and samples were filtered. Variance in the data was determined by principal component (PC) analysis. Data were visualised by Uniform Manifold Approximation and Projection (UMAP). Results Optimal SF standardisation aligned with that used for plasma, but without median normalisation. There was good reliability (<20 %CV for >80% of SOMAmers in pooled samples) and overall good correlation with immunoassay. PC1 accounted for 48% of variance and strongly correlated with individual SOMAmer signal intensities (median correlation coefficient 0.70). These could be adjusted using an ‘intracellular protein score’. PC2 (7% variance) was attributable to processing batch and was batch-corrected by ComBat. Lesser effects were attributed to other technical confounders. Data visualisation by UMAP revealed clustering of injury and OA cases in overlapping but distinguishable areas of high-dimensional proteomic space. Conclusions We define a standardised approach for SF analysis using the SOMAscan platform and identify likely ‘intracellular’ protein as being a major driver of variance in the data. Strengths and limitations This is the largest number of individual synovial fluid samples analysed by a high content proteomic platform (SomaLogic technology) SomaScan offers reliable, precise relative SF data following standardisation for over 6000 proteins Significant variance in the data was driven by a protein signal which is likely intracellular in origin: it is not yet clear whether this is due to technical considerations, normal cell turnover or relevant pathological processes Adjusting for confounding factors might conceal the true structure of the data and reduce the ability to detect ‘molecular endotypes’ within disease groups
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.110 | 0.156 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.003 |
| Bibliometrics | 0.005 | 0.004 |
| Science and technology studies | 0.003 | 0.002 |
| Scholarly communication | 0.006 | 0.001 |
| Open science | 0.004 | 0.010 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".