Patient-derived xenografts (PDX) versus patient-derived organoids (PDO) as predictors of clinical response to anti-cancer therapies.
Bibliographic record
Abstract
e15169 Background: PDX and PDO are two of the most frequently applied avatar model systems used to predict treatment response to anti-cancer therapies. Despite their frequent use, and the significant financial and ethical costs associated with developing these models, there has never been a systematic assessment of their ability to predict matched-patient treatment response. We sought to define and compare the efficacy of PDX and PDO in predicting matched-patient response to treatment. Methods: We performed a systematic review and meta-analysis in accordance with PRISMA guidelines. MEDLINE and EMBASE were queried. Inclusion criteria: PDX or PDO derived from adult solid cancer patients treated with identical systemic anti-cancer agents as the matched patient, with response assessment performed for both patient and model. Fisher’s exact test and Kaplan-Meir estimator with log-rank test were used for statistical comparisons. A 6 criteria quality assessment method based on Newcastle-Ottawa scale was applied to patient-model pairs. Results: 21565 abstracts were screened. 274 were eligible for data extraction, with 411 patient-model pairs included (N = 267 PDX, N = 144 PDO). The most common cancer types were colorectal (N = 102, 25%) and ovarian cancers (N = 77, 19%). Most common treatment modalities were chemotherapy (244, 59%) and targeted therapy (122, 30%). 55% of models were responsive to therapy (N = 227). Overall concordance in treatment response between patient and matched models was 70%, with no difference between PDX and PDO ( Table ). No significant differences in sensitivity, specificity, positive- and negative predictive value (PPV and NVP) were observed (Table). 196 pairs (48%) and significantly more PDX had high quality data reporting (56% of PDX vs 33% of PDO, P < 0.001). Of pairs with high quality data reporting, PDX had higher sensitivity, while PDO had higher specificity ( Table ). Patients whose matched PDO responded to therapy had longer median progression-free survival (mPFS; responders: 11.3 vs non-responders: 3.4 months P < 0.01). For PDX this only remained true for high quality data pairs (mPFS; responders: 9.5 vs non-responders: 6 months P < 0.01). Conclusions: This is the first study to systematically assess the utility of PDX and PDO as patient avatars. Together, these results suggest that PDO generally perform similarly to PDX as predictors of matched-patient response despite a lower ethical and financial burden. Performance metrics for PDX and PDO as predictors of matched-patient treatment response. All PDX (N = 267) PDO (N = 144) P-Value Concordance (%) 71 69 0.9 Sensitivity (%) 87 85 0.81 Specificity (%) 58 60 0.69 PPV (%) 64 52 0.1 NPV (%) 84 89 0.37 High Quality Data PDX (N = 149) PDO (N = 47) P-Value Concordance (%) 72 81 0.57 Sensitivity (%) 95 73 <0.01 Specificity (%) 54 88 <0.01 PPV (%) 61 84 0.07 NPV (%) 94 79 0.07
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.031 | 0.060 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.010 | 0.022 |
| Bibliometrics | 0.004 | 0.005 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".