Video-Based Detection Infrastructure Enhancement for Automated Ship Recognition and Behavior Analysis
Bibliographic record
Abstract
Video-based detection infrastructure is crucial for promoting connected and autonomous shipping (CAS) development, which provides critical on-site traffic data for maritime participants. Ship behavior analysis, one of the fundamental tasks for fulfilling smart video-based detection infrastructure, has become an active topic in the CAS community. Previous studies focused on ship behavior analysis by exploring spatial-temporal information from automatic identification system (AIS) data, and less attention was paid to maritime surveillance videos. To bridge the gap, we proposed an ensemble you only look once (YOLO) framework for ship behavior analysis. First, we employed the convolutional neural network in the YOLO model to extract multi-scaled ship features from the input ship images. Second, the proposed framework generated many bounding boxes (i.e., potential ship positions) based on the object confidence level. Third, we suppressed the background bounding box interferences, and determined ship detection results with intersection over union (IOU) criterion, and thus obtained ship positions in each ship image. Fourth, we analyzed spatial-temporal ship behavior in consecutive maritime images based on kinematic ship information. The experimental results have shown that ships are accurately detected (i.e., both of the average recall and precision rate were higher than 90%) and the historical ship behaviors are successfully recognized. The proposed framework can be adaptively deployed in the connected and autonomous vehicle detection system in the automated terminal for the purpose of exploring the coupled interactions between traffic flow variation and heterogeneous detection infrastructures, and thus enhance terminal traffic network capacity and safety.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".