Multiple imputation of multibeam angular response data for high resolution full coverage seabed mapping
Bibliographic record
Abstract
Abstract Acoustic data collected by multibeam echosounders (MBES) are increasingly used for high resolution seabed mapping. The relationships between substrate properties and the acoustic response of the seafloor depends on the acoustic angle of incidence and the operating frequency of the sonar, and these dependencies can be analysed for discrimination of benthic substrates or habitats. An outstanding challenge for angular MBES mapping at a high spatial resolution is discontinuity; acoustic data are seldom represented at a full range of incidence angles across an entire survey area, hindering continuous spatial mapping. Given quantifiable relationships between MBES data at various incidence angles and frequencies, we propose to use multiple imputation to achieve complete estimates of angular MBES data over full survey extents at a high spatial resolution for seabed mapping. The primary goals of this study are (i) to evaluate the effectiveness of multiple imputation for producing accurate estimates of angular backscatter intensity and substrate penetration information, and (ii) to evaluate the usefulness of imputed angular data for benthic habitat and substrate mapping at a high spatial resolution. Using a multi-frequency case study, acoustic soundings were first aggregated to homogenous seabed units at a high spatial resolution via image segmentation. The effectiveness and limitations of imputation were explored in this context by simulating various amounts of missing angular data, and results suggested that a substantial proportion of missing measurements (> 40%) could be imputed with little error using Multiple Imputation by Chained Equations (MICE). The usefulness of imputed angular data for seabed mapping was then evaluated empirically by using MICE to generate multiple stochastic versions of a dataset with missing angular measurements. The complete, imputed datasets were used to model the distribution of substrate properties observed from ground-truth samples using Random Forest and neural networks. Model results were pooled for continuous spatial prediction and estimates of confidence were derived to reflect uncertainty resulting from multiple imputations. In addition to enabling continuous spatial prediction, the high-resolution imputed angular models performed favourably compared to broader segmentations or non-angular data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.029 | 0.075 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.002 | 0.004 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".