Project HF-EOLUS. Task 2. Obtaining Radial Metrics from INTECMAR's VILA and PRIO Stations' Spectra.
Bibliographic record
Abstract
Overview This repository is part of the HF-EOLUS project and constitutes the first step towards its Task 2: obtaining wind fields from oceanic HF-radar data. Task 2 aims to use, evaluate, and develop the extraction of wind data from HF-radar data. In the HF-EOLUS methodology, the first step in wind extraction is obtaining radial metrics from the spectra collected by HF-Radar stations. This repository applies a reproducible, containerized workflow to process high-frequency (HF) radar spectra from INTECMAR’s VILA and PRIO HF-Radar stations located on the Galician shelf (NW-Spain). The primary output is LLUV files containing radial metrics derived from CODAR SeaSonde spectral data. For more details, please visit the repository page on GitHub. The applied workflow features period-specific configurations, automatic manifest generation, and scalable AWS batch processing (Lambda & S3 Batch Operations) for each station, as detailed in [1]. The underlying analysis of the HF radar spectra is carried out by the SeaSondeR R package [2]. Finally, the processed LLUV datasets are publicly available on Zenodo (PRIO [3] VILA [4]). Warning: Running processes on AWS may incur costs. We do not take responsibility for any charges arising from the use of this software. Acknowledgements This work has been funded by the HF-EOLUS project (TED2021-129551B-I00), financed by MICIU/AEI /10.13039/501100011033 and by the European Union NextGenerationEU/PRTR - BDNS 598843 - Component 17 - Investment I3. Members of the Marine Research Centre (CIM) of the University of Vigo have participated in the development of this repository. Spectra from INTECMAR's VILA and PRIO HF-Radar stations, between 2011-08-04 and 2023-11-23 have been transferred free of charge by the Observatorio Costeiro da Xunta de Galicia ( ) for their use. This Observatory is not responsible for the use of these data nor is it linked to the conclusions drawn with them. The Costeiro da Xunta de Galicia Observatory is part of the RAIA Observatory ( ). We want to thank Dr. Pedro Montero from the INTECMAR for his help providing the spectra. Disclaimer This software is provided "as is", without warranty of any kind, express or implied, including but not limited to the warranties of merchantability, fitness for a particular purpose, and noninfringement. In no event shall the authors or copyright holders be liable for any claim, damages, or other liability, whether in an action of contract, tort, or otherwise, arising from, out of, or in connection with the software or the use or other dealings in the software. References Herrera Cortijo, J. L., Fernández-Baladrón, A., Rosón, G., Gil Coto, M., Dubert, J., & Varela Benvenuto, R. (2025). Complete Batch Processing of SeaSonde HF-Radar Spectra Files on AWS with SeaSondeR R Package (v1.0.0). Zenodo. https://doi.org/10.5281/zenodo.16453046 Herrera Cortijo, J. L., Fernández-Baladrón, A., Rosón, G., Gil Coto, M., Dubert, J., & Varela Benvenuto, R. (2025). SeaSondeR: Radial Metrics from SeaSonde HF-Radar Data (v0.2.9). Zenodo. https://doi.org/10.5281/zenodo.16455051 Herrera Cortijo, J. L., Fernández-Baladrón, A., Rosón, G., Gil Coto, M., Dubert, J., Montero, P., & Varela Benvenuto, R. (2025). PRIO LLUV Radial Metrics Dataset Computed Using SeaSondeR R Package [Data set]. Zenodo. https://doi.org/10.5281/zenodo.16528653 Herrera Cortijo, J. L., Fernández-Baladrón, A., Rosón, G., Gil Coto, M., Dubert, J., Montero, P., & Varela Benvenuto, R. (2025). VILA LLUV Radial Metrics Dataset Computed Using SeaSondeR R Package [Data set]. Zenodo. https://doi.org/10.5281/zenodo.16458694
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.012 |
| Meta-epidemiology (narrow) | 0.004 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.004 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.004 | 0.003 |
| Open science | 0.004 | 0.005 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.061 | 0.096 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".