Falsification Discovery in AIS Messages ; Decision Support and Risk Assessment for Operational Effectiveness
Bibliographic record
Abstract
The SOLAS convention created an electronic system of message broadcasting between vessels: the Automatic Identification System (AIS). Albeit initially designed for security purposes, some people use the system another way, such as fleet surveillance, and the system suffers from errors, falsifications and spoofing. AIS data structure is complex, with 27 different messages, each one having a given number of data fields with various size, importance and type [1]. We propose a method based on the data quality dimensions and particularly the integrity of data within the message to assess the confidence in each message received, with a data processing done both on-the-fly with coming data and when data is stored in the database for comparison with archived data [2] and signal-based assessment we propose [3]. Each message will be assigned a confidence coefficient based on the processing of a checklist of ad-hoc items. The purpose is to assign a grade of alert and a level of associated risks (such as boarding, pollution, terrorism) to each message, group of messages sent by the same vessel or situation, suitable to be given for further studies to relevant authorities such as coast guards or MRCCs. References: [1] Tunaley, 2013. Utility of Various AIS Messages for Maritime Awareness. In proceedings of the 9th ASAR Workshop. Longueuil, Canada, October 2013. [2] Iphar, Napoli et Ray, 2016. Risk Analysis of falsified Automatic Identification System for the improvement of maritime traffic safety. In proceedings of the ESREL 2016 conference, Glasgow, United Kingdom, September 2016. [3] Alincourt, Ray, Ricordel, Dare-Emzivat et Boudraa, 2016. Methodology for AIS Signature Identification through Magnitude and Temporal Characterization. In proceedings of the OCEANS’16 SHANGHAI conference, Shanghai, China, April 2016.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".