An efficient and lightweight banana detection and localization system based on deep CNNs for agricultural robots
Bibliographic record
Abstract
• The first implementation of banana detection and localization on edge devices. • Developed an efficient, low-complexity banana detection network called Slim-Banana. • Integrated RealSense depth sensor with TOF technology for precise 3D positioning in complex orchard environments. • Demonstrated high detection accuracy, efficient performance, and low resource consumption on Nvidia Orin NX. Accurate detection and localization of fruits in natural environments is a key step for fruit picking robots to achieve precise harvesting. However, existing banana detection and positioning methods have two main limitations in practical applications: a large number of model parameters that make deployment difficult, and a need for performance improvement. To tackle the above issues, a high-precision and lightweight banana bunch recognition and localization method was proposed and deployed on edge devices for application. First, a Slim-Banana model was proposed based on the improvement of YOLOv8l. In order to reduce the model calculation amount and maintain high performance, GSConv was introduced in the Slim-Banana model to replace the standard convolution, and combined with grouped convolution and spatial convolution. At the same time, the cross-stage local network (GSCSP) module was designed to reduce the computational complexity and the complexity of the network structure through a single-stage aggregation method. Then, the RealSense depth sensor is combined with TOF technology to perform image registration and 3D localization of the banana. Finally, the pipeline is deployed on the Nvidia Orin NX edge device and its performance and resource consumption in actual work are deeply analyzed. Experimental results show that the detection precision, recall, mAP and inference time of our method are 0.947, 0.948, 0.98 and 113.6 ms respectively, the network memory size required is 4449MiB, and the average localization errors in the X-axis, Y-axis and Z-axis directions are 13.47 mm, 12.87 mm and 13.87 mm respectively. To our knowledge, this is the first work that implements banana detection and localization on edge devices. Experimental results show that compared with existing methods, our method achieves better performance in complex orchard environments, achieving efficient and lightweight banana recognition and localization.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".