Over 10 million seawater temperature records for the United Kingdom Continental Shelf between 1880 and 2014 from 17 Cefas (United Kingdom government) marine data systems
Bibliographic record
Abstract
Abstract. The datasets described here bring together quality-controlled seawater temperature measurements from over 130 years of departmental government-funded marine science investigations in the UK (United Kingdom). Since before the foundation of a Marine Biological Association fisheries laboratory in 1902 and through subsequent evolutions as the Directorate of Fisheries Research and the current Centre for Environment Fisheries & Aquaculture Science, UK government marine scientists and observers have been collecting seawater temperature data as part of oceanographic, chemical, biological, radiological, and other policy-driven research and observation programmes in UK waters. These datasets start with a few tens of records per year, rise to hundreds from the early 1900s, thousands by 1959, and hundreds of thousands by the 1980s, peaking with > 1 million for some years from 2000 onwards. The data source systems vary from time series at coastal monitoring stations or offshore platforms (buoys), through repeated research cruises or opportunistic sampling from ferry routes, to temperature extracts from CTD (conductivity, temperature, depth) profiles, oceanographic, fishery and plankton tows, and data collected from recreational scuba divers or electronic devices attached to marine animals. The datasets described have not been included in previous seawater temperature collation exercises (e.g. International Comprehensive Ocean–Atmosphere Data Set, Met Office Hadley Centre sea surface temperature data set, the centennial in situ observation-based estimates of sea surface temperatures), although some summary data reside in the British Oceanographic Data Centre (BODC) archive, the Marine Environment Monitoring and Assessment National (MERMAN) database and the International Council for the Exploration of the Sea (ICES) data centre. We envisage the data primarily providing a biologically and ecosystem-relevant context for regional assessments of changing hydrological conditions around the British Isles, although cross-matching with satellite-derived data for surface temperatures at specific times and in specific areas is another area in which the data could be of value (see e.g. Smit et al., 2013). Maps are provided indicating geographical coverage, which is generally within and around the UK Continental Shelf area, but occasionally extends north from Labrador and Greenland to east of Svalbard and southward to the Bay of Biscay. Example potential uses of the data are described using plots of data in four selected groups of four ICES rectangles covering areas of particular fisheries interest. The full dataset enables extensive data synthesis, for example in the southern North Sea where issues of spatial and numerical bias from a data source are explored. The full dataset also facilitates the construction of long-term temperature time series and an examination of changes in the phenology (seasonal timing) of ecosystem processes. This is done for a wide geographic area with an exploration of the limitations of data coverage over long periods. Throughout, we highlight and explore potential issues around the simple combination of data from the diverse and disparate sources collated here. The datasets are available on the Cefas Data Hub (https://www.cefas.co.uk/cefas-data-hub/). The referenced data sources are listed in Sect. 5.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.005 | 0.009 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".