Elementary, My Dear Watson—The Era of Natural Language Processing in Transplantation
Bibliographic record
Abstract
A novel pilot study by Srinivas et al (page 671) demonstrates the feasibility of automated data extraction from unstructured data within electronic health records, which may enhance available data for model building to predict graft loss following kidney transplantation. A novel pilot study by Srinivas et al (page 671) demonstrates the feasibility of automated data extraction from unstructured data within electronic health records, which may enhance available data for model building to predict graft loss following kidney transplantation. The use of modern data acquisition and analytics commonly used in the technology sector by Google, IBM, and others has long been the envy of the health services researcher. In this issue, Srinivas et al demonstrated the feasibility of using data mining and natural language processing (NLP) in data abstraction (1Srinivas TR Taber DJ Su Z et al.Big data, predictive analytics and quality improvement in kidney transplantation- a proof of concept.Am J Transplant. 2016; (doi: 10.1111/ajt.14099.)PubMed Google Scholar). In their study, they developed predictive models for graft loss and patient survival in kidney transplantation using single-center retrospective data. Although the data cannot be characterized as “big data,” the elements abstracted represent nondiscrete fields from textual sources and thus represent proof of concept for application to larger unstructured data sources. The main clinical application of the study was to predict 1- and 3-year graft loss and patient survival. Interest is mounting in novel methods to develop prediction models, specifically, by the addition of electronic health record (EHR) data. EHR data include both structured and unstructured data elements. Structured elements, such as laboratory values, are often collected repeatedly and at irregular intervals. Unstructured elements, such as biopsy reports or encounter notes, require processing into discrete concepts. The study used a layered approach to model building. Data included for model development, for instance, was iterative and included the United Network for Organ Sharing, a manually curated transplant database, EHR comorbidity, and posttransplant trajectory, as well as unstructured data via NLP. This allows investigation of the incremental gain in adding information, as increasing cost is associated with each additional layer. The data chosen are assumption free. Data mining sought to avoid any inherent bias in variable selection. Models may be overfitted or center-specific, limiting generalizability. This study is a proof of concept rather than a definitive prediction model. Previous work demonstrated that manually extracted EHR data improve predictive accuracy over administrative data for 30-day readmissions following kidney transplants (2Taber DJ Palanisamy AP Srinivas TR et al.Inclusion of dynamic clinical data improves the predictive performance of a 30-day readmission risk model in kidney transplantation.Transplantation. 2015; 99: 324-330Crossref PubMed Scopus (24) Google Scholar); however, there are practical constraints for manual, real-time capture of EHR data. To our knowledge, this study provides the first application of NLP tools via IBM Watson to extract unstructured data. There are, however, some limitations. The innovative use of NLP relied on a proprietary solution that effectively functions as a black box, and exploration of its accuracy and performance were not reported. The sparseness of this reporting will limit the ability to conduct further research on NLP approaches to enhance unstructured data abstraction. The statistical approach raises a number of questions, and any conclusions drawn from the final model must absolutely be viewed through a cautious but optimistic lens. The strengths of the statistical analysis include combining clinical adjudication and statistical significance for the variable selection in the multivariate model building and the use of bootstrapping methodology for model internal validation. Nevertheless, there are limitations. First, the primary outcome measure is time to event (i.e. graft loss and patient survival), is usually subject to censoring. To deal with this issue, for example, the authors simply excluded a significant number of transplants that did not have 3-year follow-up and no graft loss. This approach, however, will create bias, and the resulting predictive model may not perform well for the study population (3Hastie T Tibshirani R Friedman J The Elements of Statistical Learning: Data Mining, Inference, and Prediction..Second Edition. Springer Science & Business Media, New York2009: 758Google Scholar). Second, the authors used baseline and follow-up data for the posttransplant exposure period up to 90 days for the 1-year graft loss model and up to 365 days for the 3-year graft loss and patient survival models. However, if a participant had a graft loss within the first year of transplant (e.g. at 200 days), it is not clear whether the authors used posttransplant exposure up to 365 days or 200 days in the model building for this participant, which could lead to measuring exposures that occurred after the graft loss. This calls into question any causal inference (3Hastie T Tibshirani R Friedman J The Elements of Statistical Learning: Data Mining, Inference, and Prediction..Second Edition. Springer Science & Business Media, New York2009: 758Google Scholar). Perhaps the two greatest statistical questions raised by the authors’ approach is the handling of censored and missing data. The authors used arbitrary methods to account for missing data, particularly for the trajectory variables, which are likely to introduce bias in model building. The use of logistic regression, which is not well suited to modeling time-to-event data such as graft loss or death, is defended by the authors by citing preliminary survival analysis by Cox proportional hazards; however, the Cox and logistic regression models are quite different. For missing data, multiple imputation for some of the missing variables may improve the validity and external validation of the resulting predictive model (3Hastie T Tibshirani R Friedman J The Elements of Statistical Learning: Data Mining, Inference, and Prediction..Second Edition. Springer Science & Business Media, New York2009: 758Google Scholar). Ultimately, this study is presented as a proof of concept and not as a methodological paper for the analysis of retrospective cohorts. Nevertheless, the utility of using NLP to abstract data that may substantially enhance the ability to predict graft loss and patient survival is a leap forward. Further parallel research on NLP algorithm performance to enhance abstraction accuracy and better statistical approaches in larger multicenter data sets will be necessary to achieve the goal of predicting short-, intermediate-, and long-term graft and patient survival. The authors of this manuscript have no conflicts of interest to disclose as described by the American Journal of Transplantation.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.004 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".