Advanced Machine Learning Techniques for Survival Analysis in Medical Data
Bibliographic record
Abstract
Survival analysis is critical in many fields, particularly in healthcare, to guide medical decisions. Traditional methods like Kaplan-Meier (KM) and Cox proportional hazards (Cox) models, which generate survival curves, have limitations for long-term predictions due to their population-level assumptions. Advances in machine learning (ML) offer improved survival predictions. We apply ML models in survival prediction with three primary goals: 1) identify risk groups, 2) determine significant variables in patient survival, and 3) accurately predict long-term survival. Using binary classification to predict t-period survival, we identify risk groups and propose a novel integrated variable importance analysis using ensemble techniques and statistical analysis to highlight essential variables. Our framework is tested on bone marrow transplant (predicting 100-day to 5-year survival) and liver transplant (predicting 20-year to 30-year survival) datasets, showing improved AUCs compared to conventional models. While bone marrow transplant risk groups are statistically stratified (p-value < 0.001), liver transplant groups, though visually stratified, are not statistically significant (p-value > 0.05), indicating limited long-term survival applicability of ML classification models, even those designed for survival predictions like random survival forest (RSF). To enhance long-term survival prediction, we introduce classification-augmented survival estimation (CASE), a framework treating survival as a classification task that produces survival curves. Using dataset augmentation, CASE addresses class imbalance and offers accurate survival time predictions. Applied to a liver transplant case study, CASE improved AUCs from 0.69 to 0.88 and F1 scores from 0.32 to 0.73. Compared to KM, Cox, and RSF models, CASE showed better performance across survival metrics, including our novel metric, the mean of individual areas under the survival curve (mAUSC). We also developed temporal feature importance methods to reveal how feature significance varies over time, providing actionable insights for real-world survival problems. Sağkalım analizi, özellikle sağlık alanında tıbbi kararları yönlendirmek için birçok alanda kritik öneme sahiptir. Kaplan-Meier (KM) ve Cox orantılı risk (Cox) modelleri gibi geleneksel yöntemler, sağkalım eğrileri oluştururken popülasyon düzeyindeki varsayımlarına dayandığıdan uzun vadeli tahminlerde iyi performans gösterememektedir. Makine öğrenmesi (ML) alanındaki ilerlemeler, sağkalım tahminlerini iyileştirme imkanı sunmaktadır. Bu çalışmada, ML modellerini sağkalım tahmininde üç ana amaçla uyguluyoruz: 1) Risk gruplarını belirlemek, 2) Hasta sağkalımında önemli değişkenleri belirlemek ve 3) Uzun vadeli sağkalımı doğru bir şekilde tahmin etmek. T-periyot sağkalımını tahmin etmek için ikili sınıflandırma kullanarak risk gruplarını belirliyoruz. Temel değişkenleri vurgulamak için topluluk teknikleri ve istatistiksel analiz kullanılarak entegre bir değişken önem analizi öneriyoruz. Bu yaklaşım, kemik iliği nakli (100 günlükten 5 yıla kadar sağkalımı tahmin eden) ve karaciğer nakli (20 yıllık ve 30 yıllık sağkalımı tahmin eden) veri setlerinde test edilmiş ve geleneksel modellere kıyasla AUC’lerde önemli iyileşmeler göstermektedir. Kemik iliği nakli risk grupları istatistiksel olarak anlamlı bir şekilde ayrılmıştır (p-değeri < 0,001), ancak karaciğer nakli grupları görsel olarak ayrılmış olmasına rağmen istatistiksel olarak anlamlı ayrılmamaktadır (p-değeri > 0,05). Bu durum, random survival forest (RSF) gibi sağkalım tahminleri için tasarlanmış modeller de dahil olmak üzere, ML sınıflandırma modellerinin uzun vadeli sağkalım tahminlerindeki sınırlılığını göstermektedir. Uzun vadeli sağkalım tahminini geliştirmek için sağkalımın bir sınıflandırma görevi olarak ele alındığı ve sağkalım eğrileri üreten sınıflandırma destekli sağkalım tahmini (CASE) adlı bir çerçeve sunuyoruz. Veri seti artırımı kullanılarak CASE, sınıf dengesizliğini ele alır ve doğru sağkalım süresi tahminleri sunar. Karaciğer nakli örnek çalışmasında uygulanan CASE, AUC’leri 0,69’dan 0,88’e ve F1 skorlarını 0,32’den 0,73’e yükseltmiştir. KM, Cox ve RSF modelleri ile karşılaştırıldığında, CASE, bireysel sağkalım eğrisi altındaki alanların ortalaması (mAUSC) adlı yeni bir metriğimiz de dahil olmak üzere sağkalım metrikleri açısından daha iyi performans göstermiştir. Son olarak, gerçek dünyadaki sağkalım problemleri için uygulanabilir öngörüler sağlayan, zaman içinde değişken öneminin nasıl değiştiğini ortaya koyan geçici özellik önem yöntemleri de sunmaktayız.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.002 | 0.006 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.008 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".