Machine learning algorithms for accurate differential diagnosis of lymphocytosis based on cell population data
Bibliographic record
Abstract
The morphological identification of lymphoproliferative disorders with peripheral expression is of utmost importance for the accurate diagnosis and subsequent clinical decisions (Bain, 2015). According to quality assurance control surveys from Canada and the UK, neoplastic misclassification of reactive specimens occurs in 10–26% of cases (Brereton et al, 2015; Johnston et al, 2016) while up to 21% of neoplastic cases are assigned as reactive (Johnston et al, 2016), highlighting the need for improvement. Evaluation of the leucocyte subpopulation using DXH800 (Beckman Coulter, Miami, FL) analysers yields cell population data (CPD) with information on volume, conductivity (mean population values and SD) and 10 different scatter measures (Briggs, 2009; Bain, 2015). Machine learning algorithms (MLA) (Zini, 2005) could objectively assist the cytologist in the critical task of evaluating such mass data and would enhance accuracy, increase speed and reduce subjectivity but has, to our knowledge, not been applied for lymphoid classification using CPD parameters. In this study, the aforementioned 14 CPD, together with absolute lymphoid count, age, and gender were used to design, develop and implement novel MLAs as a support tool for lymphocyte-related diagnosis. A total of 400 samples were included, corresponding to (i) 136 healthy controls (HC), (ii) 139 virus-infected (VI) (from cytomegalovirus and Epstein–Barr infected subjects) and (iii) 125 chronic lymphocytic leukaemia (CLL) patients. Classification was according to International Council for Standardization in Haematology guidelines (Palmer et al, 2015) and inclusion criteria were based on flow cytometry or serological data (Data S1; Swerdlow et al, 2017). Well aware of the natural diversity, the neoplastic category was deliberately restricted to CLL because of its incidence (Eichhorst et al, 2015) in this first approach of MLA for the lymphocyte categorisation. All specimens diagnosed with CLL presented lymphocytosis but this condition occurred only in 70·5% of the VI cases. The “atypical lymphoid” analyser flag was rendered in more than 98% of the CLL and VI cases (Table SI). Raw CPD were analysed and modelled using open-source software: Python (v.3.6; https://www.python.org/downloads/release/python-360/) and the Sckit-Learn (v.0.18.1) package (https://pypi.python.org/pypi/scikit-learn/0.18.1). Three hundred samples (100 from each category) were used to build the model. The first stage included feature analysis by different approaches: (i) Analysis of variance (ANOVA) of 14 CPD parameters, which revealed that all of them displayed significant differences among the three categories (Table SII), (ii) Unsupervised K-Means clustering was performed to evaluate the potential of intrinsic group-induced differences in the CPD parameters. The combination of absolute lymphoid count and CPD parameters allowed good separation in two groups (for K = 2 clusters – control and pathological-, χ2 = 75, P-value = 5·16 × 10−17) or three diagnostic categories (for K = 3, χ2 = 267·02 and P-value = 1·40 × 10−56) also revealing associations between the obtained clusters (for the unlabelled cases) and the specific category. Age and gender were discarded as they did not add to efficient cluster separation (Table SIII), (iii) Principal component analysis (PCA) enabled cluster validity to be assessed. The second stage included building supervised MLA (classifiers) for true classification purposes (Alpaydin, 2016). The minimal discriminant variables corresponded to absolute lymphoid count and CPD. With these, the model-building dataset was split in two; 225 samples were used to train the model and the remaining 75 were used exclusively to test the model. The following classification algorithms were tested: decision trees (DT), random forests (RF), naive Bayes classifier (NBC), k-nearest neighbour (KNN), neural networks (NN) and support vector machines (SVM). Model evaluation and selection were based on the accuracies obtained after a 10-fold cross validation. Table 1 displays the obtained classification accuracies for the training, testing and 10-fold cross validation of each particular model. The best model corresponded to NN classifier, with an accuracy of 98·7%, followed by SVM (98·0%) and KNN (98·0%) classifiers. Finally, each classifier was validated employing the remaining and unique 100 new samples (validation data set). The validation accuracies are also shown in Table 1. Again, the best accuracy was obtained with NN classifier (98·0%). The other models rendered slightly lower accuracies, between 96·0% and 98·0%, except for DT that yielded 89·0% validation accuracy and was disregarded henceforth. As NN displayed the optimal discrimination potential, it was targeted for further analysis and performance assessment evaluation by means of a confusion matrix (Table 2) and through precision, recall and F-1 scores (Table SIV). The analysis of the confusion matrix along with the performance measures allowed a better understanding of the obtained results. CLL and VI were both classified with 100% precision, and CLL and HC with 100% recall, whereas the lowest values for either parameter were still 95·0%. The F-1 score was 97·0% for HC and VI, and 100% for CLL. A precision of 100% for the CLL and VI groups means that all the cases assigned to these categories were correct and no false-positive assignments were made, whereas the 95·0% observed for the HC group was due to two false positives (2 VI) in this category. Concerning the recall measurements, both HC and CLL displayed 100%, which means that all HC and CLL cases were correctly classified. Recall measurements were not affected by the false positive cases in the HC group. The 95·0% recall for the VI group implies 2 false negative assignments for this category. Overall, with supervised MLA, the obtained models were able to classify new cases with an overall accuracy that ranged from 96·0% to 98·0% (Table 1), indicating that different MLA have impressive discriminatory potential. As such, with the best performing MLA, NN, the true positive rate for the viral infection classification in our study corresponded to 94·90% and 100% true positive rates for the neoplastic assignment (Table 2). In conclusion, the NN algorithm, based on CPD and absolute lymphoid counts, appears to be simplest and most efficient ancillary aid for the clinical laboratory diagnostic approach of lymphocytosis. This standardised tool can be easily implemented globally but despite the tremendous aid that MLA offer for medical diagnosis, one should never ease into the pitfall of completely ruling out human intervention. We thank technicians of the Haematology department and the Biochemistry department of Synlab Global Diagnostics S.A. for providing the serological confirmation of the virally infected samples. Finally, special thanks to the Scientific and Technical direction of Synlab Global Diagnostics S.A. in the Esplugas de Llobregat laboratory for supporting this project. LB, IL and RGG designed the study. LB selected the samples, collected, analysed and interpreted the data. LB, IL and RGG wrote the manuscript and critically revised the paper. Data S1. Supplementary methods. Table SI. Clinical characteristics of healthy controls (HC), virus-infected (VI) and chronic lymphocytic leukaemia (CLL). Table SII. Mean and SD values for the CPD parameters for the three different categories [Healthy controls (HC), viral-infected (VI) specimens and chronic lymphocytic leukaemia (CLL) samples]. Table SIII. (A) Chi-square and P-values obtained for different features groups (K = 2 and K = 3) in clustering analyses. (B) Distribution of HC, VI and CLL samples within the different clusters (B for K = 2 and C for K = 3) according CPD and absolute lymphoid counts. Table SIV. Performance of the neural network classifier. Please note: The publisher is not responsible for the content or functionality of any supporting information supplied by the authors. Any queries (other than missing content) should be directed to the corresponding author for the article.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".