Comparing historical and modern methods of sea surface temperature measurement – Part 1: Review of methods, field comparisons and dataset adjustments
Bibliographic record
Abstract
Abstract. Sea surface temperature (SST) has been obtained from a variety of different platforms, instruments and depths over the past 150 yr. Modern-day platforms include ships, moored and drifting buoys and satellites. Shipboard methods include temperature measurement of seawater sampled by bucket and flowing through engine cooling water intakes. Here I review SST measurement methods, studies analysing shipboard methods by field or lab experiment and adjustments applied to historical SST datasets to account for variable methods. In general, bucket temperatures have been found to average a few tenths of a °C cooler than simultaneous engine intake temperatures. Field and lab experiments demonstrate that cooling of bucket samples prior to measurement provides a plausible explanation for negative average bucket-intake differences. These can also be credibly attributed to systematic errors in intake temperatures, which have been found to average overly-warm by >0.5 °C on some vessels. However, the precise origin of non-zero average bucket-intake differences reported in field studies is often unclear, given that additional temperatures to those from the buckets and intakes have rarely been obtained. Supplementary accurate in situ temperatures are required to reveal individual errors in bucket and intake temperatures, and the role of near-surface temperature gradients. There is a need for further field experiments of the type reported in Part 2 to address this and other limitations of previous studies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".