Deep Learning For The Classification of Lung Diseases \nUsing Chest X-Rays
Bibliographic record
Abstract
The discovery of X-rays marked a significant milestone in the field of medicine. One of the most common types of X-rays, the chest X-ray (CXR), allows doctors to examine an individual’s internal structure without surgery. Over the years, deep learning methods and algorithms have been developed to automate lung disease detection and identification. This paper introduces RADIA, a project that combines multiple deep learning techniques to identify abnormal areas and abnormal- ities in chest X-rays. RADIA builds upon previous studies conducted by the Stanford ML group, such as ChexNet and ChexPert. Our team utilized the ConvNeXt-Large, a deep learning convo- lutional model, implemented with a pre-trained ConvNext algorithm on the ImageNet database to classify various pathologies from public datasets like ChestX-ray14, CheXpert, MIMIC-CXR, PadChest, and VinDr-CXR, as well as a private dataset obtained from the Picture Archiving Com- munication System (PACS) at Verdun and Notre Dame Hospitals in Montreal in the collaboration with CIUSSS (Centre Integre Universitaire de Sante et de Services Sociaux du Centre-Sud-de-l’Ile- de-Montreal) and valuable consultants from the radiology team at Notre Dame Hospital contributed to the project’s success. Our team employed image enhancement and augmentation techniques to create various image versions. We used different and novel approaches to address the challenges, and the results were evaluated using metrics such as AUC, F1, and G-Means to analyze performance with imbalanced input data. It is essential to note that the project’s development extends beyond the creation of a web tool based on deep learning techniques. Our future plans involve building a decision helper that combines inference models and web tools to assist healthcare professionals.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".