Elementary, My Dear Watson—The Era of Natural Language Processing in Transplantation
Notice bibliographique
Résumé
A novel pilot study by Srinivas et al (page 671) demonstrates the feasibility of automated data extraction from unstructured data within electronic health records, which may enhance available data for model building to predict graft loss following kidney transplantation. A novel pilot study by Srinivas et al (page 671) demonstrates the feasibility of automated data extraction from unstructured data within electronic health records, which may enhance available data for model building to predict graft loss following kidney transplantation. The use of modern data acquisition and analytics commonly used in the technology sector by Google, IBM, and others has long been the envy of the health services researcher. In this issue, Srinivas et al demonstrated the feasibility of using data mining and natural language processing (NLP) in data abstraction (1Srinivas TR Taber DJ Su Z et al.Big data, predictive analytics and quality improvement in kidney transplantation- a proof of concept.Am J Transplant. 2016; (doi: 10.1111/ajt.14099.)PubMed Google Scholar). In their study, they developed predictive models for graft loss and patient survival in kidney transplantation using single-center retrospective data. Although the data cannot be characterized as “big data,” the elements abstracted represent nondiscrete fields from textual sources and thus represent proof of concept for application to larger unstructured data sources. The main clinical application of the study was to predict 1- and 3-year graft loss and patient survival. Interest is mounting in novel methods to develop prediction models, specifically, by the addition of electronic health record (EHR) data. EHR data include both structured and unstructured data elements. Structured elements, such as laboratory values, are often collected repeatedly and at irregular intervals. Unstructured elements, such as biopsy reports or encounter notes, require processing into discrete concepts. The study used a layered approach to model building. Data included for model development, for instance, was iterative and included the United Network for Organ Sharing, a manually curated transplant database, EHR comorbidity, and posttransplant trajectory, as well as unstructured data via NLP. This allows investigation of the incremental gain in adding information, as increasing cost is associated with each additional layer. The data chosen are assumption free. Data mining sought to avoid any inherent bias in variable selection. Models may be overfitted or center-specific, limiting generalizability. This study is a proof of concept rather than a definitive prediction model. Previous work demonstrated that manually extracted EHR data improve predictive accuracy over administrative data for 30-day readmissions following kidney transplants (2Taber DJ Palanisamy AP Srinivas TR et al.Inclusion of dynamic clinical data improves the predictive performance of a 30-day readmission risk model in kidney transplantation.Transplantation. 2015; 99: 324-330Crossref PubMed Scopus (24) Google Scholar); however, there are practical constraints for manual, real-time capture of EHR data. To our knowledge, this study provides the first application of NLP tools via IBM Watson to extract unstructured data. There are, however, some limitations. The innovative use of NLP relied on a proprietary solution that effectively functions as a black box, and exploration of its accuracy and performance were not reported. The sparseness of this reporting will limit the ability to conduct further research on NLP approaches to enhance unstructured data abstraction. The statistical approach raises a number of questions, and any conclusions drawn from the final model must absolutely be viewed through a cautious but optimistic lens. The strengths of the statistical analysis include combining clinical adjudication and statistical significance for the variable selection in the multivariate model building and the use of bootstrapping methodology for model internal validation. Nevertheless, there are limitations. First, the primary outcome measure is time to event (i.e. graft loss and patient survival), is usually subject to censoring. To deal with this issue, for example, the authors simply excluded a significant number of transplants that did not have 3-year follow-up and no graft loss. This approach, however, will create bias, and the resulting predictive model may not perform well for the study population (3Hastie T Tibshirani R Friedman J The Elements of Statistical Learning: Data Mining, Inference, and Prediction..Second Edition. Springer Science & Business Media, New York2009: 758Google Scholar). Second, the authors used baseline and follow-up data for the posttransplant exposure period up to 90 days for the 1-year graft loss model and up to 365 days for the 3-year graft loss and patient survival models. However, if a participant had a graft loss within the first year of transplant (e.g. at 200 days), it is not clear whether the authors used posttransplant exposure up to 365 days or 200 days in the model building for this participant, which could lead to measuring exposures that occurred after the graft loss. This calls into question any causal inference (3Hastie T Tibshirani R Friedman J The Elements of Statistical Learning: Data Mining, Inference, and Prediction..Second Edition. Springer Science & Business Media, New York2009: 758Google Scholar). Perhaps the two greatest statistical questions raised by the authors’ approach is the handling of censored and missing data. The authors used arbitrary methods to account for missing data, particularly for the trajectory variables, which are likely to introduce bias in model building. The use of logistic regression, which is not well suited to modeling time-to-event data such as graft loss or death, is defended by the authors by citing preliminary survival analysis by Cox proportional hazards; however, the Cox and logistic regression models are quite different. For missing data, multiple imputation for some of the missing variables may improve the validity and external validation of the resulting predictive model (3Hastie T Tibshirani R Friedman J The Elements of Statistical Learning: Data Mining, Inference, and Prediction..Second Edition. Springer Science & Business Media, New York2009: 758Google Scholar). Ultimately, this study is presented as a proof of concept and not as a methodological paper for the analysis of retrospective cohorts. Nevertheless, the utility of using NLP to abstract data that may substantially enhance the ability to predict graft loss and patient survival is a leap forward. Further parallel research on NLP algorithm performance to enhance abstraction accuracy and better statistical approaches in larger multicenter data sets will be necessary to achieve the goal of predicting short-, intermediate-, and long-term graft and patient survival. The authors of this manuscript have no conflicts of interest to disclose as described by the American Journal of Transplantation.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,004 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».