Designing optimal object detection networks for detecting damages to canola kernels
Bibliographic record
Abstract
Canola is an essential Canadian crop that generated a revenue of $14.4 B from its export in the 2022 fiscal year. To consistently export the best quality canola and set the correct pricing that truly reflects the grade, it is crucial to invest in reliable, high speed, and accurate grading technologies. Leveraging the current progress in Artificial Intelligence algorithms, this research proposes a comprehensive end-to-end system to detect damage to canola kernels and grade them. The proposed system comprises an accelerated sample preparation setup, custom-built specialized AI models, and an edge-AI microprocessor for in-field use. This thesis mainly focuses on the development of specialized neural networks to accurately detect damaged canola seeds as it is the most complex part of the system. Two foundational object detection networks, You Only Look Once (YOLO) version 5 and version 7 were optimized to be compatible with resource-constrained hardware environments. The aim of developing the two optimal networks was to mitigate trade-offs among speed, accuracy, cost, and model size, thereby creating a more balanced and optimized network. Several architectural design options were explored to compress the network structure of the two models in terms of size and cost yet retain or improve the metrics of accuracy and inference speed. The results from the design choices indicate that reconstructing YOLOv5 with ShuffleNet as the backbone reduces its size and cost and increases the inference speed but negatively affects the detection accuracy. Replacing the Convolutional Blocks present in the Spatial Pyramid Pooling Cross Stage Partial and all the Efficient Layer Aggregation Network modules of YOLOv7 with the Ghost Convolutional Network and adding two Convolutional Block Attention Modules improves both mean average precision metric compared to the parent model and reduces the size and cost. To prepare the samples faster, a semi-automatic option was adopted and its effect on dataset quality and the overall time required was also studied. This study is a first step toward developing a commercially viable canola grading system.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".