Combining a forward supervised filter learning with a sparse NMF for breast cancer histopathological image classification
Bibliographic record
Abstract
Histopathological images play a important role in clinical diagnosis, particularly in identifying and assessing the severity of abnormal conditions like benign lesions and malignant tumors. Traditional machine learning techniques for processing histopathology images involve the extraction of manual features from these images, which is typically done with the assistance of industry experts. Recent advancements in Deep Learning (DL), especially with Convolutional Neural Networks (CNN), have enabled the automatic extraction of multi-level abstract features directly from raw data. This capability significantly enhances the performance of complex computer vision tasks. Classic CNN models like AlexNet and VggNet employ back-propagation algorithms to learn filters in the training phase. However, these algorithms demand large labeled datasets, resulting in extensive computational processing. Additionally, they often face the vanishing gradient problem, which can negatively impact the quality of the learning process. Besides, in many domains, acquiring enough labeled images for conducting properly the training phase is a real challenge. To address these challenges, a feed-forward propagation approach was proposed using Non-Negative Matrix Factorization(NMF). The NMF technique factorizes the input data into two latent factors (non-negative matrices). It has been shown that by enforcing constraints such as sparsity on the latent factors, dominant features that are mostly correlated with tumors types can be extracted. In this work, a novel model combining sparse NMF and Support Vector Machine (SVM) was developed for classifying histopathological images. We have derived a mathematical model of a novel feed-forward filter learning approach that combines sparse NMF (SNMF) and Support Vector Machine technique (SVM). The model was used to design and implement a feed-forward CNN classifier to classify histopathology images. This model has been evaluated on the histopathology images from Sultan Qaboos University Hospital (SQUH dataset) and the public BreaKHis dataset. The experiments we have conducted demonstrate the efficiency of the proposed model, especially on small-sized SQUH datasets achieving an AUC of 0.90, 0.89, 0.85, and 0.86 on 4x,10x, 20x, and 40x magnifications, respectively, and achieving an AUC of 0.95 BreaKHis dataset. • Proposing a novel feed-forward CNN for classifying histopathology images. The filters are learned using Sparse Non-negative Matrix Factorization with Support Vector Machine. • Conducted experiments that demonstrate the advantage of using feed-forward methods over back-propagation one on small size datasets. • Achieves high AUC on small datasets: 0.90 on SQUH and 0.95 on BreaKHis. Efficient in training time, outperforming traditional models like VggNet-16 and ResNet-50.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".