Natural Language Processing to Assess Documentation of Features of Critical Illness in Discharge Documents of Acute Respiratory Distress Syndrome Survivors
Bibliographic record
Abstract
RATIONALE: Transitions to outpatient care are crucial after critical illness, but the documentation practices in discharge documents after critical illness are unknown. OBJECTIVES: To characterize the rates of documentation of various features of critical illness in discharge documents of patients diagnosed with acute respiratory distress syndrome (ARDS) during their hospital stay. METHODS: We used natural language processing tools to build a keyword-based classifier that categorizes discharge documents by presence of terms from four groups of keywords related to critical illness. We used a multivariable modified Poisson regression model to infer patient- and hospital-level characteristics associated with documentation of relevant keywords. A manual chart review was used to validate the accuracy of the keyword-based classifier, and to assess for ARDS documentation during the hospital stay. MEASUREMENTS AND MAIN RESULTS: Of 815 discharge documents, ARDS was identified in only 111 (13%). Mechanical ventilation was identified in 770 (92%) and intensive care unit (ICU) admission in 693 (83%) of discharge documents. Symptoms or recommendations related to post-intensive care syndrome were included in 306 (38%) of discharge documents. Patient age (older; relative risk [RR] = 0.97/yr, 95% confidence interval [CI] = 0.96-0.98) and higher PaO2:FiO2 (decreasing illness severity; RR = 0.96/10-unit increment, 95% CI = 0.93-0.98) were associated with decreased documentation of ARDS. Being discharged from a surgical (RR = 0.33, 95% CI = 0.22-0.50) compared with a medicine service was also associated with decreased rates of ARDS documentation. The manual chart review revealed 98% concordance between ARDS documentation in the discharge summary and during the hospital stay. Accuracy of the document classifier was 100% for ARDS and mechanical ventilation, 98% for ICU admission, and 95% for symptoms of post-intensive care syndrome. CONCLUSIONS: In the discharge documents of survivors of ARDS, ARDS itself is rarely mentioned, but mechanical ventilation and ICU stay frequently are. The low rates of documentation of ARDS appear to be concordant with low rates of documentation during the hospital stay, consistent with known underrecognition in the ICU. Natural language processing tools can be used to effectively analyze large numbers of discharge documents of patients with critical illness.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".