NMR Spectral bins, coralline algae metabolomics [dataset]
Bibliographic record
Abstract
Nuclear Magnetic Resonance (NMR) spectroscopy data of 14 species of crustose coralline red algae and one non-coralline calcareous red alga collected from the Great Barrier Reef, Australia. NMR data generated using the following methodology: Polar extracts were dried using a centrifugal evaporator, and metabolites were reconstituted in 200 µL deuterium oxide (D2O) buffered with phosphate-buffered saline (PBS) and including 0.05% sodium-3-(trimethylsilyl)-2,2,3,3-tetradeuteriopropionate (TSP) as an internal reference. Reconstituted samples were loaded into 3 mm NMR tubes using a glass syringe (Hamilton® Company, Reno, Nevada), and spectra were acquired with an 800 MHz Bruker® Avance III HDX spectrometer equipped with a Triple (TCI) Resonance 5 mm Cryoprobe. Spectra were acquired at 298 K, using the internal reference for field locking (TSP δ 0.00 ppm) and the zg30 pulse program with 0.8 relaxation delay, 8.20 pulse width and a spectral width of 16 kHz using 64 scans. Acquired NMR spectra for each sample were manually phase corrected, baseline adjusted using the ablative algorithm, referenced and normalised to TSP (1H δ 0.00), using MestReNova version 14.2.2 (Mestrelab Research S.L., Spain). Metabolites were identified based on 1H NMR spectra using Chenomx v11 software (Chenomx Inc., Edmonton, Canada) and, where possible, additional annotations were made by cross-referencing against standard compounds in the Human Metabolome Database (HMDB) and Yeast Metabolome Database (YMDB). Context: Data collected to examine the nature of algal metabolites to explore potential relationships with the settlement of coral larvae for applicability in reef restoration projects in the Great Barrier Reef.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.002 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.004 | 0.008 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.027 | 0.028 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".