Classification of collagen remodeling in asthma using second-harmonic generation imaging, supervised machine learning and texture-based analysis
Bibliographic record
Abstract
Airway remodeling is present in all stages of asthma severity and has been linked to reduced lung function, airway hyperresponsiveness and increased deposition of fibrillar collagens. Traditional histological staining methods used to visualize the fibrotic response are poorly suited to capture the morphological traits of extracellular matrix (ECM) proteins in their native state, hindering our understanding of disease pathology. Conversely, second harmonic generation (SHG), provides label-free, high-resolution visualization of fibrillar collagen; a primary ECM protein contributing to the loss of asthmatic lung elasticity. From a cohort of 13 human lung donors, SHG-imaged collagen belonging to non-asthmatic (control) and asthmatic donors was evaluated through a custom textural classification pipeline. Integrated with supervised machine learning, the pipeline enables the precise quantification and characterization of collagen, delineating amongst control and remodeled airways. Collagen distribution is quantified and characterized using 80 textural features belonging to the Gray Level Cooccurrence Matrix (GLCM), Gray Level Size Zone Matrix (GLSZM), Gray Level Run Length Matrix (GLRLM), Gray Level Dependence Matrix (GLDM) and Neighboring Gray Tone Difference Matrix (NGTDM). To denote an accurate subset of features reflective of fibrillar collagen formation; filter, wrapper, embedded and novel statistical methods were applied as feature refinement. Textural feature subsets of high predictor importance trained a support vector machine model, achieving an AUC-ROC of 94% ± 0.0001 in the classification of remodeled airway collagen vs. control lung tissue. Combined with detailed texture analysis and supervised ML, we demonstrate that morphological variation amongst remodeled SHG-imaged collagen in lung tissue can be successfully characterized.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".