MétaCan
Menu
Back to cohort
Record W4412138693 · doi:10.2196/62996

Deep Learning for the Early Detection of Invasive Ductal Carcinoma in Histopathological Images: Convolutional Neural Network Approach With Transfer Learning

2025· article· en· W4412138693 on OpenAlexvenueno aff
Yawo Ezunkpe

Bibliographic record

VenueJMIR Formative Research · 2025
Typearticle
Languageen
FieldComputer Science
TopicAI in cancer detection
Canadian institutionsnot available
Fundersnot available
KeywordsConvolutional neural networkTransfer of learningPreprintArtificial intelligenceDeep learningComputer science

Abstract

fetched live from OpenAlex

Background Invasive ductal carcinoma (IDC) is considered the most common form of breast cancer, accounting for a significant percentage of mortality worldwide. Therefore, its early detection is vital to further improve patients’ outcomes and survival rates. However, conventional diagnostic methods in the form of manual histopathological examinations are time-consuming, subjective, and prone to errors. Therefore, there is an urgent need to develop automated solutions for accurate IDC detection in histopathology images to assist pathologists in clinical decision-making. Objective We aim to develop and validate a convolutional neural network (CNN) model for early detection of IDC by analyzing histopathological images. The specific objectives are designing a deep learning–based technique for automated detection of IDC, assessing its performance compared to traditional diagnostic methods, and evaluating its utility in a clinical setup for early breast cancer diagnosis. These methods will be available to practitioners in underdeveloped countries via an open-source application. Methods The dataset for the research included 277,524 publicly available histopathological images from Kaggle, comprising both IDC-positive and IDC-negative images. About 71.6% of images were IDC-positive (class 0), while 28.4% were IDC-negative (class 1). Since our data are unbalanced, we created a weighted loss function to overcome the class imbalance problem. Further development was based on a CNN using the approach of transfer learning with a pretrained architecture called Visual Geometry Group to uplift feature extraction so that performance may improve; hence, images were preprocessed and normalized to perform augmentation with robustness. The model was developed using a split of 80% for training and 20% for testing. Model performance was measured for accuracy, sensitivity, specificity, precision, recall, and F1-score in the confusion matrix and classification report. Results From our CNN base model, we obtained an accuracy of 89% on the test set. Later, the base model was used with a weighted loss function to balance the class weights, giving a lower accuracy of 86% on the test set. Data augmentation was performed but did not improve the results. To deal with the class imbalance effectively, we performed transfer learning with a pretrained model, which gave an accuracy of 90% on the test set. Conclusions The CNN-based model thus showed accuracy and reliability for early detection of IDC from histopathological images. This technique will potentially act as an efficient and accurate assistant tool for pathologists, contributing to the early diagnosis of breast cancer and improving clinical outcomes. This paper provides an important contribution toward refining the performance of this model and widening its applications in a clinical setting by integrating it with other diagnostic techniques for better outcomes.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.002
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.010
Threshold uncertainty score0.019

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0010.002
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0010.001
Science and technology studies0.0000.000
Scholarly communication0.0010.001
Open science0.0010.001
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0010.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.038
GPT teacher head0.312
Teacher spread0.274 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designBench or experimental
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations4
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueJMIR Formative ResearchSame topicAI in cancer detectionFrench-language works237,207