A context-aware multi-stream attentive convolutional neural network for surface defect segmentation
Bibliographic record
Abstract
Surface defect segmentation (SDS) plays a crucial role in modern industrial inspection, driven by advancements in deep learning and computer vision. However, existing methods often struggle with defects that vary in size, shape, and appearance, especially when defects resemble the background. One key challenge is the effective fusion of low-level spatial and high-level semantic features, which conventional encoder-decoder networks handle inadequately, resulting in poor boundary precision and contextual inconsistency. To address these limitations, we propose MSAC-Net, a context-aware encoder-decoder architecture optimized for robust SDS. The encoder employs a dual-stream design: one stream captures semantic features via EfficientNetV2M, while the other enhances structural detail using a multi-scale convolutional cascade and channel-shuffle SEDNet module. A spatial attention-enhanced fusion module unifies these streams to improve feature integration. At the bottleneck, a spatially shifted MLP is combined with residual context features and a multi-scale aggregation module to capture long-range dependencies. The decoder incorporates cross-attention gates and ECA-enhanced residual convolution blocks to progressively refine segmentation. Evaluations on four benchmark datasets (DAGM2007, SD900, Magnetic Tile, and KolektorSDD2) show that MSAC-Net achieves an average mDSC of 95.15%, outperforming state-of-the-art methods due to superior multi-scale feature integration, attention-guided fusion, and a lightweight design, making it ideal for industrial deployment. • MSAC-Net is a multi-stream context aggregation network for segmenting defects with indistinct edges and variable shapes. • Employs channel shuffling, dilated convolutions, and attention to enhance defect segmentation with multi-contextual features. • Performs effective fusion of local contextual and global semantic features. • Outperforms state-of-the-art segmentation methods in four benchmark datasets. • Balances performance and computation, demonstrating MSAC-Net’s suitability for real-world defect inspection and localization.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".