MétaCan
Menu
Back to cohort
Record W4255094042 · doi:10.5194/egusphere-egu2020-537

Is there a 'right' spatial scale? Improving pedological multi- scale modelling by optimizing input data grain size: A Case Study using Average Local Variance.

2020· preprint· en· W4255094042 on OpenAlexaff
Christopher Scarpone, Anders Knudby, Stephanie Melles, Andrew A. Millward

Bibliographic record

Venuenot available
Typepreprint
Languageen
FieldEnvironmental Science
TopicSoil Geostatistics and Mapping
Canadian institutionsUniversity of OttawaToronto Metropolitan University
Fundersnot available
KeywordsVariance (accounting)Scale (ratio)Grain sizeVariable (mathematics)Computer scienceStatisticsSpatial analysisAggregate (composite)EconometricsData miningMathematicsCartographyGeology

Abstract

fetched live from OpenAlex

<p>Current soil mapping practitioners are faced with a plethora of choices of digital data for input to their modelling approaches; these data have local to global extents and are highly variable in their grain size. Deciding at what scale to represent individual covariates for a specific project, therefore, can be difficult and confusing. Moreover, a lack of accessible methodology and tools focused on determining an optimal input data scales (grain size) has led to the current status quo, which is to use data at the scale delivered by the data provider. Soil prediction models are typically applied using the grain size of the coarsest variable, scaling other data to match. In this study, average local variance was investigated as a method to determine optimal grain size(s) for input variables to a soil contaminant prediction model. The Meuse dataset was used, and heavy metal soil contamination was mapped using RandomForest. A Data Cube was employed to handle data inputs of varying grain size. Two scenarios were investigated for model prediction accuracy: (1) contaminant predictions made using data with optimized grain size, and (2) contaminant predictions made using input data where grain size was unchanged, “as received” from the data provider. Both model predictions were assessed using a cross-validation approach. Early results indicate that optimization of grain size based on average local variance can improve prediction accuracy and point toward the importance of understanding the spatial heterogeneity of an input variable and how it changes with different grain sizes prior to incorporation in a predictive model. This research lays a foundation for the creation of an automated approach practitioners can use to help untangle the relationship between the intrinsic spatial scale for a process of interest and how that process is represented in the scale of input data.</p>

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMeta-epidemiology (narrow), Open science, Insufficient payload (model declined to judge)
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: Simulation or modeling
GenreCandidate signal: Methods · Consensus signal: none
Teacher disagreement score0.482
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.000
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0000.000
Science and technology studies0.0010.000
Scholarly communication0.0000.000
Open science0.0010.010
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.080
GPT teacher head0.297
Teacher spread0.217 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designSimulation or modeling
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2020
Admission routes1
Has abstractyes

Explore more

Same topicSoil Geostatistics and MappingFrench-language works237,207