Disentangling Internal Tides from Balanced Motions with Deep Learning and Surface Field Synergy
Bibliographic record
Abstract
A fundamental challenge in ocean dynamics is disentangling balanced motions and internal waves. Extracting internal tidal (IT) imprints from surface data is a central part of this challenge. Traditional harmonic analysis can fail under strong incoherence and poor temporal sampling, as in global satellite observations. New wide-swath satellites provide two-dimensional spatial coverage, allowing IT extraction to be reformulated as image translation. Building on our earlier deep-learning approach for extracting IT signatures from sea surface height (SSH) in an idealized turbulent simulation, we show that a simpler, computationally cheaper algorithm performs comparably in our experiments when the learning rate is annealed during training. Using this algorithm, we test different combinations of surface inputs: SSH, surface temperature, and surface velocity. All fields contribute synergistically to disentanglement in our deterministic benchmark, with surface velocity by far the most informative. These findings underscore the value of coordinated multi-platform observations and highlight the importance of surface velocity for separating balanced motions and internal waves. Additional analysis shows that both wave-signature information and scattering-medium information aid IT extraction. To exploit large-scale, mesoscale-reaching information in the scattering medium, the algorithm must be highly non-local. Residual errors concentrate at small spatial scales near mode-2 tidal wavelengths, likely reflecting incomplete input information, uncertainty in the simulation-derived reference fields, including possible Doppler-shift contamination, and limitations of the present deterministic architecture.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".