Serial Clinical and Biomarker Monitoring during Treatment Can Stratify Patients with Low Risk Gvhd
Bibliographic record
Abstract
Background: The majority of patients respond to steroid treatment for graft-vs-host disease (GVHD), but the typically high doses and prolonged courses can cause substantial morbidity. GVHD can be classified at onset into two groups with significant differences in treatment response and non-relapse mortality (NRM) by clinical criteria (Minnesota risk) but the more favorable standard risk group still experiences high rates of treatment failure and NRM [MacMillan, BBMT 2015]. We previously validated a serum biomarker-based (ST2+REG3α) risk stratification system, MAGIC algorithm probabilities (MAPs), that can stratify patients both at initiation and during treatment into risk groups for treatment response and NRM [Hartwell, JCI Insight 2017, Srinagesh, Blood Adv 2019]. GVHD that is Minnesota standard risk and has a low MAP (i.e., Ann Arbor 1) at start of treatment represents a low risk group with >80% response rate to steroids and ~10% NRM [Etra, Blood 2023]. We hypothesized that serial monitoring of GVHD symptom severity and MAPs twice (days (d) 7 and 14) in patients with clinical and biomarker defined low risk GVHD could further stratify patients into clinically meaningful groups. Methods: We retrospectively determined MAPs at d7 and 14 in 450 patients with low risk GVHD who participated in a prospective, multi-center, observational trial and were treated with systemic steroids between 2015-2023. MAPs <0.291 at days 7 and 14 were defined as low risk based on a previously validated threshold [Major-Monfried Blood 2018]. Clinical responses were scored as complete (CR), partial (PR), or non-response (NR) by standard criteria. Initiation of second line therapy was considered NR. Patients were categorized into three groups (A-C) based on clinical responses and MAPs at days d7 and 14. A: CR/PR by d14 and low MAP at both d7+14; B: stable or worse GVHD (NR) by d14 with low MAP at both d7 and 14; C: any response with a high MAP at d7 or 14. The endpoints were clinical response at day 28 (overall response rate, ORR=CR or PR), regardless of earlier response status, and 6-month NRM. Results: The median age of the study population was 54 years (range: 0-79); 19% of patients received PT-Cy for GVHD prophylaxis. Target organ involvement was skin ± upper GI (66%), lower GI (23%), isolated UGI (11%), and skin+liver (1%), and GVHD severity was grade I (35%), II (61%), and III (4%) at start of treatment. Patients in group A (n=310, 69%), clinical responders with low MAPs through day 14, had excellent d28 ORR (93%) and very low NRM (4%). GVHD in this group was considered ultra-low risk (ULR) and subset analyses revealed highly consistent d28 ORR and NRM, regardless of target organ severity, age (<18, 18-59, >60), HCT-CI score, conditioning intensity, and GVHD prophylaxis. Group B patients comprised clinical non-responders with low MAPs through day 14 (n=112, 25%); half of them responded by d28 and their NRM was low (8%), albeit double that of group A (p=0.079). Group C was small (n=28, 6%) and included both clinical responders (n=20, 4%) and clinical non-responders (n=8, 2%) who had at least one high MAP through day 14. Their d28 ORR was 54% but NRM increased four-fold (32%) and was significantly worse than patients who maintained low MAPs regardless of clinical response (p<0.001). Conclusion: Serial monitoring by clinical response and biomarkers during the first two weeks of treatment identifies a small number of high risk patients whose GVHD is initially low risk (group C). This group experiences high NRM regardless of clinical response to steroid treatment and may benefit from treatment escalation. Patients with low risk GVHD who do not have an early clinical response to steroids but maintain low MAPs, (group B) have reasonably good long-term outcomes with standard treatment. Finally, serial monitoring identifies a very large subset of patients with ULR GVHD (group A) who respond exceptionally well to standard treatment and experience very low rates of 6-month NRM regardless of pre-transplant characteristics and GVHD prophylaxis. Such patients may benefit from treatment de-escalation strategies such as shorter courses of lower dose steroids, a strategy we are currently testing in a clinical trial (NCT05090384).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".