Bibliographic record
Abstract
The concept of geochemical background is discussed, defined and a distinction made between natural background and ambient background. A key issue is that background is a range, not a single value. The range of background should span the measurements likely to be encountered during sampling and analysis in situations devoid of major mineral occurrences and severe impacts by anthropogenic contamination. The acceptance of 'ambient background' as a quantifiable estimate implies that some anthropogenic impact is acknowledged, but it does not 'overwhelm' the natural patterns of variation due to geology, pedology, etc. The role of both spatial, map, and statistical data displays is discussed, and how appraisal of survey data with these tools informs as to whether the data should be divided into subsets before estimating background ranges. As an example, the data acquired by the US-EPA 3050B Aqua Regia (4:1 HCl-HNO3) variant for As and Pb in the <2 mm fraction of the 0-5 cm soil interval are discussed and background ranges estimated. It is demonstrated that the Maritimes 2007 data are poly-populational, and different background ranges should to be estimated for the three ecoprovinces present, the Appalachian and Acadian Highlands, the Northumberland Uplands and the Fundy Uplands. Several factors underlie the spatial definition one of which is geology, base metal ore occurrences, the Bathurst camp, lie in the Appalachian and Acadian Highlands ecoprovince; and widespread minor As occurrences associated with gold in mainland Nova Scotia occur in the Fundy Uplands ecoprovince. Both statistical numerical methods and graphical methods are demonstrated. It is shown that different numerical procedures and whether data are logarithmically transformed, a common practice in applied geochemistry, lead to different estimates. This begs the question, which is right, or at least the best? It is shown how the a combination of graphical inspection to remove outliers likely not representative of background processes, e.g., data related to the presence of major mineral occurrences or discernable anthropogenic contamination, and the use of percentiles leads to useful estimates of background range. In conclusion, some more complex multivariate approaches to background estimation and gaining an understanding of the data are briefly presented, with their constraints. For univariate, an element at a time, estimates it is recommended that the hybrid approach of map and statistical data displays, the removal of nonbackground data from the data set(s) and the use of percentiles be adopted.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.003 | 0.004 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.005 | 0.003 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.020 | 0.007 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".