Systematical Analysis and Pathological Classification of Breast Cancer from Mammographic Images with Using Specific Machine Learning Methods
Bibliographic record
Abstract
For years, breast cancer has been a serious problem and malignant tumor case primarily causes death of women all around the world. In this paper, a computer based breast tumor analysis and pathological case classification system has been achieved and some novelties are included to the image processing methods, especially in segmentation and base frequency distribution acquisition of the processed image and classification part. First, the possible noises and artifacts are eliminated by using common filtering. Second, the filtered images are segmented with integrating gray level Image Processing methods. Then, these images (ROIs) are converted to the base frequency distribution images with using Fast Fourier Transform (FFT) and Lab&HSV color spaces. The most important key for these images is frequency distribution can be obtained with specific color tones and totally 100 images (50 benign-50 malignant) are accumulated to fed the two different Machine Learning models in literature such as Probabilistic Neural Network as Learning Vector Quantization (LVQ) and Support Vector Regression (SVR) for classification of Benign and Malignant cases without the need for additional medical data. Then the performance of the proposed system is analyzed with 30 different test images (15 benign-15 malignant) according to the metrics like accuracy, sensitivity, specificity, precision, F-score and area under the ROC curve (AUC score). The experimental results on the open access mammogram image set show that discriminating between Benign and Malignant cases can be achieved with an important success rate as 91.38% with LVQ and %.92 with SVR.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.003 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".