Machine-learning models of δ <sup>13</sup> C and δ <sup>15</sup> N isoscapes in Amazonian wood
Bibliographic record
Abstract
Abstract. Illegal logging is one of the most prevalent environmental infractions in the Amazon, led by organized networks that cause substantial ecological and economic impacts. Official control mechanisms, such as Brazil’s Forest Origin Document (DOF), remain vulnerable to the fraudulent manipulation of virtual timber credits and inconsistencies in digital traceability. These deficiencies highlight the need for independent, scientifically based methodologies for timber traceability that can support law enforcement and ensure reliable provenance verification. Here, we tested whether the isotopic composition of carbon (δ13C) and nitrogen (δ15N) in wood can trace Amazonian timber origin. We developed basin-wide δ13C and δ15N isoscapes using machine-learning models to predict spatial variability. A total of 571 trees from 47 sites were analyzed for both isotopes. Tree disks or wedges were sampled from the basal trunk, sectioned transversely, and sub-sampled from heartwood to near the sapwood boundary to obtain whole-tree isotopic composition. The δ13C and, more strongly, the δ15N values exhibited substantial within-site heterogeneity, indicating individual-level physiological controls, interspecific differences, and/or small-scale environmental variation influencing isotope fractionation. Despite these sources of noise, isotopic values showed independent and predictable spatial patterns across the basin (R2 = 0.67 for δ15N and R2 = 0.60 for δ13C). Nitrogen isotopes were primarily controlled by edaphic factors, while carbon isotopes revealed a broad longitudinal gradient linked to climate. Together, these isotopic markers provide complementary information for basin-scale timber provenancing and form a robust, high-resolution framework for Amazon-wide traceability.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".