TranSEF: Transformer Enhanced Self-Ensemble Framework for Damage Assessment in Canola Crops
Bibliographic record
Abstract
Crop health monitoring is crucial for implementing timely and effective interventions that ensure sustainability and maximize crop yield. Flea beetles (FB), Crucifer (Phyllotreta cruciferae) and Striped (Phyllotreta striolata), pose a significant threat to canola crop health and cause substantial damage if not addressed promptly. Accurate and timely damage quantification is crucial for implementing targeted pest management strategies if insecticidal seed treatments are overcome by FB feeding to minimize yield losses if the action threshold is exceeded. Traditional manual field monitoring for FB damage is time-consuming and error-prone due to reliance on human visual estimates of FB damage. This article proposes TranSEF, a novel self-ensemble semantic segmentation algorithm that utilizes a hybrid convolutional neural network-vision transformer (ViT) encoder–decoder framework. The encoder employs a modified cross-stage partial DenseNet (CSPDenseNet), MCSPDNet, which enhances attention to tiny regions by aggregating spatially aware features from shallow layers with deeper, more abstract features. ViTs effectively capture the global context in the decoder by modeling long-range dependencies and relationships across the image. Each decoder independently processes inputs from different stages of the MCSPDNet, acting as a weak learner within an ensemble-like approach. Unlike traditional ensemble learning approaches that train weak learners separately, TranSEF is trained end-to-end, making it a self-ensembling framework. TranSEF uses hybrid supervision with a composite loss function, where decoders generate independent predictions and simultaneously supervise each other. TranSEF achieves IoU scores of 0.831 for canola leaves and 0.807 for FB damage, and the overall mIoU improved by 2.29% and 1.56% over DeepLabv3+ and SegFormer, respectively, while utilizing only 35.42 M trainable parameters-significantly fewer than DeepLabv3+ (63 M) and SegFormer (61 M).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".