A Machine Learning Explanation of the Pathogen-Immune Relationship of SARS-CoV-2 (COVID-19), and a Model to Predict Immunity and Therapeutic Opportunity: A Comparative Effectiveness Research Study
Bibliographic record
Abstract
BACKGROUND: Approximately 80% of those infected with COVID-19 are immune. They are asymptomatic unknown carriers who can still infect those with whom they come into contact. Understanding what makes them immune could inform public health policies as to who needs to be protected and why, and possibly lead to a novel treatment for those who cannot, or will not, be vaccinated once a vaccine is available. OBJECTIVE: The primary objectives of this study were to learn if machine learning could identify patterns in the pathogen-host immune relationship that differentiate or predict COVID-19 symptom immunity and, if so, which ones and at what levels. The secondary objective was to learn if machine learning could take such differentiators to build a model that could predict COVID-19 immunity with clinical accuracy. The tertiary purpose was to learn about the relevance of other immune factors. METHODS: This was a comparative effectiveness research study on 53 common immunological factors using machine learning on clinical data from 74 similarly grouped Chinese COVID-19-positive patients, 37 of whom were symptomatic and 37 asymptomatic. The setting was a single-center primary care hospital in the Wanzhou District of China. Immunological factors were measured in patients who were diagnosed as SARS-CoV-2 positive by reverse transcriptase-polymerase chain reaction (RT-PCR) in the 14 days before observations were recorded. The median age of the 37 asymptomatic patients was 41 years (range 8-75 years); 22 were female, 15 were male. For comparison, 37 RT-PCR test-positive patients were selected and matched to the asymptomatic group by age, comorbidities, and sex. Machine learning models were trained and compared to understand the pathogen-immune relationship and predict who was immune to COVID-19 and why, using the statistical programming language R. RESULTS: When stem cell growth factor-beta (SCGF-β) was included in the machine learning analysis, a decision tree and extreme gradient boosting algorithms classified and predicted COVID-19 symptom immunity with 100% accuracy. When SCGF-β was excluded, a random-forest algorithm classified and predicted asymptomatic and symptomatic cases of COVID-19 with 94.8% AUROC (area under the receiver operating characteristic) curve accuracy (95% CI 90.17%-100%). In total, 34 common immune factors have statistically significant associations with COVID-19 symptoms (all c<.05), and 19 immune factors appear to have no statistically significant association. CONCLUSIONS: The primary outcome was that asymptomatic patients with COVID-19 could be identified by three distinct immunological factors and levels: SCGF-β (>127,637), interleukin-16 (IL-16) (>45), and macrophage colony-stimulating factor (M-CSF) (>57). The secondary study outcome was the suggestion that stem-cell therapy with SCGF-β may be a novel treatment for COVID-19. Individuals with an SCGF-β level >127,637, or an IL-16 level >45 and an M-CSF level >57, appear to be predictively immune to COVID-19 100% and 94.8% (AUROC) of the time, respectively. Testing levels of these three immunological factors may be a valuable tool at the point of care for managing and preventing outbreaks. Further, stem-cell therapy via SCGF-β and M-CSF appear to be promising novel therapeutics for patients with COVID-19.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".