Field Intercomparison of Radiometer Measurements for Ocean Colour Validation
Bibliographic record
Abstract
A field intercomparison was conducted at the Acqua Alta Oceanographic Tower (AAOT) in the northern Adriatic Sea, from 9 to 19 July 2018 to assess differences in the accuracy of in- and above-water radiometer measurements used for the validation of ocean colour products. Ten measurement systems were compared. Prior to the intercomparison, the absolute radiometric calibration of all sensors was carried out using the same standards and methods at the same reference laboratory. Measurements were performed under clear sky conditions, relatively low sun zenith angles, moderately low sea state and on the same deployment platform and frame (except in-water systems). The weighted average of five above-water measurements was used as baseline reference for comparisons. For downwelling irradiance ( E d ), there was generally good agreement between sensors with differences of <6% for most of the sensors over the spectral range 400 nm–665 nm. One sensor exhibited a systematic bias, of up to 11%, due to poor cosine response. For sky radiance ( L s k y ) the spectrally averaged difference between optical systems was <2.5% with a root mean square error (RMS) <0.01 mWm−2 nm−1 sr−1. For total above-water upwelling radiance ( L t ), the difference was <3.5% with an RMS <0.009 mWm−2 nm−1 sr−1. For remote-sensing reflectance ( R r s ), the differences between above-water TriOS RAMSES were <3.5% and <2.5% at 443 and 560 nm, respectively, and were <7.5% for some systems at 665 nm. Seabird-Hyperspectral Surface Acquisition System (HyperSAS) sensors were on average within 3.5% at 443 nm, 1% at 560 nm, and 3% at 665 nm. The differences between the weighted mean of the above-water and in-water systems was <15.8% across visible bands. A sensitivity analysis showed that E d accounted for the largest fraction of the variance in R r s , which suggests that minimizing the errors arising from this measurement is the most important variable in reducing the inter-group differences in R r s . The differences may also be due, in part, to using five of the above-water systems as a reference. To avoid this, in situ normalized water-leaving radiance ( L w n ) was therefore compared to AERONET-OC SeaPRiSM L w n as an alternative reference measurement. For the TriOS-RAMSES and Seabird-HyperSAS sensors the differences were similar across the visible spectra with 4.7% and 4.9%, respectively. The difference between SeaPRiSM L w n and two in-water systems at blue, green and red bands was 11.8%. This was partly due to temporal and spatial differences in sampling between the in-water and above-water systems and possibly due to uncertainties in instrument self-shading for one of the in-water measurements.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".