MétaCan
Menu
Back to cohort
Record W1967160455 · doi:10.1111/ddi.12062

The value of a datum – how little data do we need for a quantitative risk analysis?

2013· article· en· W1967160455 on OpenAlexaff
Brian Leung, Russell Steele

Bibliographic record

VenueDiversity and Distributions · 2013
Typearticle
Languageen
FieldEnvironmental Science
TopicSpecies Distribution and Climate Change
Canadian institutionsMcGill University
Fundersnot available
KeywordsGeodetic datumComputer scienceRelevance (law)EconometricsSample (material)StatisticsData scienceData miningGeographyMathematicsCartography

Abstract

fetched live from OpenAlex

Abstract Aim Conservation managers are typically faced with limited resources, time and information. The philosophy underlying risk assessment should be robust to these limitations. While there is a broad support for the concept of risk assessments, there is a tendency to rely on expert opinion and exclude formal data analysis, possibly because available information is often scarce. When data analyses are conducted, often much simplified models are advocated, even though this means excluding processes believed by experts to be important. In this manuscript, we ask: should statistical analyses be conducted and decisions modified based on a single datum? How many data points are needed before predictions are meaningful? Given limited data, how complex should models be? Location World‐wide. Methods We use simulation approaches with known ‘true’ values to assess which inferences are possible, given different amounts of information. We use two metrics of performance: the magnitude of uncertainty (using posterior mean squared error) and bias (using P – P plots). We assess six models of relevance to conservation ecologists. Results We show that the greatest reduction in uncertainty occurred at the smallest sample sizes for models examined, and much of parameter space could be excluded. Thus, analyses based on even a single datum potentially can be useful. Further, with only a few observations, the predicted distribution of outcomes matched the probabilities of actual occurrences, even for relatively complex state‐space models with multiple sources of stochasticity. Main conclusions We highlight the utility of quantitative analyses even with severely limited data, given existing practices and arguments in the conservation literature. The purpose of our manuscript is in part a philosophical discourse, as modifications are needed to how conservation ecologists are often trained to think about problems and data, and in part a demonstration via simulation analysis.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesScience and technology studies, Insufficient payload (model declined to judge)
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.303
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0020.000
Scholarly communication0.0000.000
Open science0.0000.001
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0020.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.060
GPT teacher head0.268
Teacher spread0.208 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations8
Published2013
Admission routes1
Has abstractyes

Explore more

Same venueDiversity and DistributionsSame topicSpecies Distribution and Climate ChangeFrench-language works237,207