Machine Learning Classification of Spatial Patterns of Malignant Cells Reveals Implications in Prognosis and Tumor Microenvironment Composition in Lymphoma
Notice bibliographique
Résumé
Background Malignancies exhibit variable cellular distribution patterns and the relationship between these topographic variations, underlying biological processes, and clinical outcomes remain poorly understood. Point process analyses, widely used in ecology, can elucidate the spatial distribution of points in complex systems but have rarely been applied to tumor heterogeneity. We recently demonstrated (Hoppe et al, Cancer Discovery 2023), that cells co-expressing high MYC and BCL2 but lacking BCL6 (M+2+6-), in Diffuse Large B Cell Lymphoma (DLBCL) are consistently correlated with poor survival compared to other MYC/BCL2/BCL6 combinations. Machine learning approaches can be applied to understand nuances of cellular point patterns and help with correlating with clinicopathological variables. Here, we developed a code frame that can be generalized across tissue regions accounting for heterogeneity to quantitatively study spatial patterns of M+2+6- cells in DLBCL. Methods We developed a scalable automated workflow for generating spatial point patterns. The individual steps of the pipeline have been consolidated into a standalone package that can be executed in a facile manner for any image type, without dependencies on any proprietary software. Using multiplexed fluorescent immunohistochemistry (mfIHC) in four cohorts of DLBCL (n=449), spatial point patterns were derived, and Geyer's point process model was applied. Machine learning classification models were benchmarked for spatial statistics derived from the point process model. A multi-omic analysis, including single-cell transcriptomic analyses of 22 DLBCL samples, was then conducted. Spatial transcriptomic technique, Stereoseq, was also conducted on 2 DLBCL samples. Results The workflow consists of the following parts: 1) The python script using the OpenCV package to manipulate the kernel size and intensity of the spatial coordinates overlaid on the images; 2) A QuPath groovy script to automate the import and export of the images and the parameter thresholds for the pixel classifier; 3) An R script to build the spatial point patterns from coordinates and save different oncogene co-expression as marks within the accurate geojson annotations; and 4) an R script to obtain measures of quality in terms of minimizing the number of points excluded while generating accurate spatial point pattern windows. After applying the pipeline, we see that patients could be divided into two: one group showed “clustered” spatial organization, while the other displayed a “dispersed” M+2+6- cell distribution. We achieved an accuracy of 98% in classifying patients as “dispersed” and “clustered” across four cohorts, through the random forest model. Cases with “dispersed” M+2+6- cells had shorter overall survival across all analyzed cohorts (P < 0.05 in 4/4 cohorts). Patients enriched in the “dispersed” phenotype, predominantly belonged to the ABC cell of origin subtype. We derived a “dispersed” pattern gene signature through multi-omic analyses which expressed genes implicated in cell migration and adhesion. Validation of the dispersed signature was conducted using Stereoseq, where M+2+6- cells enriched in the signature displayed greater values of L function across distances. Patients enriched in the dispersed phenotype displayed lower infiltration of immune subtypes though deconvolution hinting at a possible immunologically cold microenvironment. Conclusion This study demonstrates the clinical relevance studying the spatial distribution of malignant cell subpopulations through point pattern analysis. We anticipate that this machine learning pipeline can be developed for clinical use, enabling the classification of spatial phenotypes in DLBCL biopsies for patient stratification.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,003 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,001 |
| Bibliométrie | 0,001 | 0,001 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,001 | 0,001 |
| Science ouverte | 0,001 | 0,001 |
| Intégrité de la recherche | 0,001 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,001 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».