Coastal Ocean Data Analysis Product in North America (CODAP-NA) – An internally consistent data product for discrete inorganic carbon,oxygen, and nutrients on the U.S. North American ocean margins
Bibliographic record
Abstract
Abstract. Internally-consistent, quality-controlled data products play a very important role in promoting regional to global research efforts to understand societal vulnerabilities to ocean acidification (OA). However, there are currently no such data products for the coastal ocean where most of the OA-susceptible commercial and recreational fisheries and aquaculture industries are located. In this collaborative effort, we compiled, quality controlled (QC), and synthesized two decades of discrete measurements of inorganic carbon system parameters, oxygen, and nutrient chemistry data from the U.S. North American continental shelves, to generate a data product called the Coastal Ocean Data Analysis Product for North America (CODAP-NA). There are few deep-water (> 1500 m) sampling locations in the current data product. As a result, cross-over analyses, which rely on comparisons between measurements on different cruises in the stable deep ocean, could not form the basis for cruise-to-cruise adjustments. For this reason, care was taken in the selection of data sets to include in this initial release of CODAP-NA, and only data sets from laboratories with known quality assurance practices were included. New consistency checks and outlier detections were used to QC the data. Future releases of this CODAP-NA product will use this core data product as the basis for secondary QC. We worked closely with the investigators who collected and measured these data during the QC process. This version of the CODAP-NA is comprised of 3,292 oceanographic profiles from 61 research cruises covering all continental shelves of North America, from Alaska to Mexico in the west and from Canada to the Caribbean in the east. Data for 14 variables (temperature; salinity; dissolved oxygen concentration; dissolved inorganic carbon concentration; total alkalinity; pH on the Total Scale; carbonate ion concentration; fugacity of carbon dioxide; and concentrations of silicate, phosphate, nitrate, nitrite, nitrate plus nitrite, and ammonium) have been subjected to extensive QC. CODAP-NA is available as a merged data product (Excel, CSV, MATLAB, and NetCDF, https://doi.org/10.25921/531n-c230, https://www.ncei.noaa.gov/data/oceans/ncei/ocads/metadata/0219960.html) (Jiang et al., 2020). The original cruise data have also been updated with data providers' consent and summarized in a table with links to NOAA's National Centers for Environmental Information (NCEI) archives (https://www.ncei.noaa.gov/access/ocean-acidification-data-stewardship-oads/synthesis/NAcruises.html).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.002 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.003 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".