Response to Letter: Global Immune Biomarkers and Donor Serostatus Can Predict Cytomegalovirus Infection Within Seropositive Lung Transplant Recipients
Bibliographic record
Abstract
We thank Deng et al1 for their interest in our study. We explored absolute lymphocyte count (ALC) and mitogen component of the Quantiferon-CMV assay as predictors of cytomegalovirus (CMV) infection in seropositive lung transplant recipients. After controlling for antiviral prophylaxis using Cox proportional hazards models, donor (D) seropositivity (adjusted hazard ratio [aHR], 2.33; 95% confidence interval [CI], 1.54-3.54; P < 0.001), lower ALC (aHR per unit decrease, 1.56; 95% CI, 1.19-2.08; P = 0.002), and lower mitogen values (aHR per unit decrease, 1.09; 95% CI, 1.03-1.14; P = 0.001) were all associated with CMV. Although adding ALC or mitogen to serostatus improved predictions, combining all 3 variables together was no better than using 2. This was unexpected as we had hypothesized that combining multiple predictors would improve model performance. We have included the variance inflation factors that support our conclusion that this finding was not due to collinearity, which were all low (Table 1). This was surprising, given the weak but statistically significant correlation between ALC and mitogen (Kendall’s tau 0.25, P < 0.01) and biologically plausible reasons for either collinearity or improved model performance. We found no evidence for interactions. To further assess the possibility of more complex, nonlinear relationships, we used previously calculated cutoffs to categorize patients into subgroups (Figure 1; Table 2). This resulted in separation into multiple risk categories, with patients CMV D–, ALC >1, and mitogen >3.6 at lowest risk (CMV in 4/25; 16%) and CMV D+, ALC ≤1, and mitogen ≤3.6 at highest risk (CMV in 25/32; 78%, aHR 12.38, 95% CI, 4.26-35.93; P < 0.001). Model performance was only marginally better (C-statistic 0.70 versus 0.68 without mitogen, similar Akaike and Bayesian information criterion values). Although this may enhance risk stratification for the minority of patients in the highest and lowest deciles, utility for most patients at moderate risk appears limited. TABLE 1. - Multivariable modeling results with VIFs Predictor Adjusted HR(95% CI) P VIF Adjusteda HR(95% CI) P VIF Lymphocyte count, ×1000 cells/μL (n = 189) 1.56 (1.19–2.08)b 0.002 1.0005 1.42 (1.07–1.89) 0.016 1.102 Mitogen value, IU/mL 1.09 (1.03–1.14)b 0.001 1.030 1.06 (1.004–1.12) 0.03 1.129 Donor CMV seropositive 2.33 (1.54–3.54)c <0.001 1.001 2.53 (1.63–3.92) <0.001 1.007 Valganciclovir prophylaxis duration, mo – – 1.24 (1.10–1.40) <0.001 1.021 aAdjusted for donor serostatus, valganciclovir prophylaxis duration, lymphocyte count, and mitogen value.bAdjusted for donor serostatus and valganciclovir prophylaxis duration.cAdjusted for valganciclovir prophylaxis duration only.CI, confidence interval; CMV, cytomegalovirus; HR, hazard ratio; VIF, variance inflation factor. TABLE 2. - CMV risk stratification based on donor serostatus, ALC (×1000 cells/μL), and mitogen values (IU/mL) Donor serostatus Biomarker results N CMV infection, n (%) Adjusted HRa(95% CI) P Negative ALC >1, mitogen >3.6 25 4 (16%) Ref – ALC >1, mitogen ≤3.6 21 8 (38%) 2.67 (0.80–8.89) 0.11 ALC ≤1, mitogen >3.6 10 6 (60%) 5.70 (1.60–20.25) 0.007 ALC ≤1, mitogen ≤3.6 18 10 (56%) 7.66 (2.36–24.89) <0.001 Positive ALC >1 mitogen >3.6 49 26 (53%) 5.15 (1.79–14.80) 0.002 ALC >1, mitogen ≤3.6 20 14 (70%) 7.66 (2.51–23.42) <0.001 ALC ≤1, mitogen >3.6 14 10 (71%) 7.10 (2.22–22.68) <0.001 ALC ≤1 mitogen ≤3.6 32 25 (78%) 12.38 (4.26–35.93) <0.001 aAdjusted for duration of valganciclovir prophylaxis.ALC, absolute lymphocyte count; CI, confidence interval; CMV, cytomegalovirus; HR, hazard ratio. FIGURE 1.: Unadjusted Kaplan-Meier estimates of freedom from cytomegalovirus (CMV) infection overall stratified by donor serostatus, ALC (× 1000 cells/μL), and mitogen values (IU/mL). P values refer to log-rank test results. ALC, absolute lymphocyte count.We are excited that complex machine learning/artificial intelligence algorithms may better model these complex relationships, improving predictions.2 However, important limitations exist. These techniques are complex, with relevant expertise less widespread than standard regression. Training machine learning models to accurately predict complex outcomes requires large, well-curated data sets with clearly defined predictors. Careful consideration of temporal relationships is crucial to avoid “data leakage” where spurious associations arise from variables associated with the outcome.3 Some methods do not incorporate prior knowledge beyond the data set and cannot help explore causality or contributions of individual factors, complicating interpretation. Overfitting is a major risk, impacting generalizability and highlighting the need for rigorous external validation to verify findings. Given its simplicity, availability, and low cost, ALC is an appealing biomarker. Growing evidence consistently shows a relationship between low ALC and CMV infection, summarized in the updated CMV guidelines.4 Given additional costs and complexity of performing tests such as the mitogen assay, predictive utility would need to be substantially better to justify use. Overall, at the current time, simple predictors such as donor serostatus and ALC offer a practical approach to CMV risk prediction which could be easily translated into clinical practice. This would be an important step toward more individualized CMV risk prediction, informing patient management decisions and improving clinical outcomes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.019 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.012 | 0.011 |
| Insufficient payload (model declined to judge) | 0.031 | 0.019 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".