The use of radiomics and natural language processing to detect pain in the simulation-CT images of patients undergoing radiotherapy for bone metastasis
Bibliographic record
Abstract
Cancer develops when cells lose their ability to control division and form a tumor.Malignant tumor cells can invade nearby tissues or spread (metastasize) to other parts of the body.Bone is one of the most common sites for cancer to metastasize to.Bone Metastases (BM) can result in inflammation, structural damage, and severe pain.70 to 90% of patients with BM suffer from severe pain.Therefore, detecting and controlling BM-associated pain has the potential to improve the quality of life of BM patients.This thesis project aimed to develop and evaluate an Artificial Intelligence (AI) pipeline for detecting pain in cancer patients with BM by combining information from clinical texts and radiographic images.The project fits within an ultimate research goal of enabling early prediction and management of BM pain before it becomes distressing.It addressed three specific objectives in three studies: 1) Construction of a Natural Language Processing (NLP) pipeline to extract pain scores from consultation notes, 2) Construction of a radiomics pipeline to extract BM lesion features from radiographic images, and 3) Development of a radiomics-based machine-learning model of pain in patients with BM by combining NLPquantified pain scores with radiomic features.In the first study, we trained and tested an NLP pipeline using publicly-available hospital discharge notes and achieved a precision and recall of 0.86 and 0.83 in detecting sentence level pain scores.The pipeline was then used to automatically extract and classify note-level pain from clinical notes at our institution with 0.925 F1 score.In the second study, a radiomics model was generated based on a novel lesion-centerpoint-based geometric Regions Of Interest (ROIs).The geometric ROIs were automatically delineated around lesion centerpoints that were manually pinpointed by radiation oncologists on CT images.This allowed us to greatly simplify the data preparation process.We demonstrated that, the introduced pipeline was successful in Abstract iii differentiating BM from healthy bones.In the third study, a Machine Learning (ML) pipeline was developed to detect pain in cancer patients with thoracic spinal BM.The study used data from 176 patients treated at Cedar Cancer Center between January 2016 and September 2019.Our NLP pipeline was used to extract pain scores from radiation oncology consultation notes.Radiomics features were extracted from each ROI, and various ML classifiers were evaluated using precision, recall, F1-score, and Area Under the Receiver Operating Characteristic Curve (ROC-AUC).The results showed that the pipeline was successful in differentiating between painful and painless BM lesions with an accuracy, specificity, and ROC-AUC of 0.82, 0.85, and 0.83, respectively.Overall in this thesis, we developed a robust radiomics pipeline to identify painful BM lesions in CT images.Our pipeline is fast and scalable as it is trained using NLP-extracted pain scores from clinical notes and it requires just centerpoints to identify BM lesions in CT images.This work represents the first step in building a clinically-practical pain detection pipeline and is consistent with the ultimate goal of better managing pain in patients with BM. iv Abrégé Le cancer se développe lorsque les cellules perdent leur capacité à contrôler la division et forment une tumeur.Les cellules de tumeur maligne peuvent envahir les tissus voisins ou se propager (métastaser) à d'autres parties du corps.Les os sont l'un des sites les plus courants pour les métastases cancéreuses.Les Métastases Osseuses (MO) peuvent entraîner une inflammation, des dommages structurels et une douleur intense.De 70 à 90% des patients atteints de MO souffrent de douleur intense.Par conséquent, détecter et contrôler la douleur associée à la MO a le potentiel d'améliorer la qualité de vie des patients atteints de MO.Ce projet de thèse visait à développer et à évaluer un pipeline IA pour détecter la douleur chez les patients atteints de MO en combinant des informations provenant de notes cliniques et d'images radiographiques.Le projet s'inscrit dans un objectif de recherche qui permettra ultimement la prédiction et la gestion précoce de la douleur MO avant qu'elle ne devienne angoissante.Il comporte trois objectifs spécifiques dans trois études: 1) Construction d'un pipeline du traitement du langage naturel (NLP) pour extraire les scores de douleur à partir des notes de consultation, 2) Construction d'un pipeline radiomique pour extraire les caractéristiques de lésion MO à partir d'images radiographiques, et 3) Développement d'un modèle d'apprentissage machine basé sur la radiomique pour prédire la douleur chez les patients atteints de MO en combinant les scores de douleur quantifiés par NLP et les caractéristiques radiomiques.Dans la première étude, nous avons formé et testé un pipeline NLP à l'aide de notes de congé d'hôpital disponibles publiquement.Globalement, nous avons obtenu une précision et un rappel de 0.86 est 0.83 dans la détection de la douleur.Le pipeline a ensuite été utilisé pour extraire et classer automatiquement la douleur à partir des notes cliniques de notre institution.Dans la seconde étude, un modèle radiomique a été généré sur la base de régions Abrégé v d'intérêt (ROIs) géométriques originales, basées sur le centre de la lésion.Les ROIs géométriques ont été délimitées automatiquement autour des centres des lésions ayant été repérées manuellement par des radiothérapeutes à partir des images CT.Cela nous a permis de simplifier considérablement le processus de préparation des données.Nous avons démontré que le pipeline a pu différencier le MO des os sains avec succès.Dans la troisième étude, un pipeline d'apprentissage automatique a été développé pour détecter la douleur chez les patients atteints de cancer avec MO de la colonne thoracique.L'étude a utilisé les données de 176 patients traités au centre de cancer des Cèdres entre janvier 2016 et septembre 2019.Notre pipeline NLP a été utilisé pour extraire les scores de douleur des notes de consultation en radiothérapie.Des caractéristiques radiomiques ont été extraites de chaque ROI, et utilisées par divers classificateurs d'apprentissage automatique qui ont été évalués en utilisant la précision, le rappel, le score F1 et l'aire sous la courbe (AUC) des courbes d'efficacité du récepteur.Les résultats ont montré que le pipeline a réussi à différencier les lésions BM douloureuses et indolores avec une précision, une spécificité et une AUC de 0.82, 0.85 et 0.83, respectivement.Dans l'ensemble, nous avons développé dans cette thèse un pipeline radiomique robuste pour identifier les lésions MO douloureuses à partir d'images CT.Notre pipeline est rapide et évolutif, car il est construit en utilisant des scores de douleur extraits du NLP des notes cliniques et qu'il nécessite simplement des points centraux pour identifier les lésions MO dans les images CT.Ce travail représente la première étape dans la construction d'un pipeline de détection de la douleur pratique en clinique et est cohérent avec le but ultime qui est de mieux gérer la douleur chez les patients avec MO. vi and mentor.I deeply appreciate his support and work in helping me achieve my Ph.D.; without him, this thesis would not have been completed or written.One simply could not wish for a better or friendlier supervisor.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.008 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".