MétaCan
Menu
Back to cohort
Record W4367285108 · doi:10.51744/cip2

Designing evaluations to provide evidence to inform action in new settings

2018· report· en· W4367285108 on OpenAlexaff
Calum Davey, Syreen Hassan, Nancy Cartwright, Macartan Humphreys, Edoardo Masset, Audrey Prost, David Gough, Sandy Oliver, Chris Bonell, James Hargreaves

Bibliographic record

Venuenot available
Typereport
Languageen
FieldDecision Sciences
TopicEvaluation and Performance Assessment
Canadian institutionsImpact
Fundersnot available
KeywordsPsychological interventionContext (archaeology)Action (physics)Public relationsPromotion (chess)PsychologyPolitical scienceGeography

Abstract

fetched live from OpenAlex

Policy and interventions should be informed by the best available evidence, but evaluations are not always optimally designed to inform decisions about policies and interventions in new contexts. Learning the most possible from evaluations is important; evaluating is expensive and policy makers should be confident about their decisions. Using evidence from previous studies can lead to better policy decisions, but there have been cases where doing so has led to interventions that have not worked. Learning from evaluations for decisions elsewhere has generally been more successful for interventions that are simple and are less context dependent (or context-dependent in a simple way, such as depending on the severity of the problem). With increasing focus on complex, context-dependent interventions, we need to ensure that evaluations can offer as much information as possible to guide decisions in other contexts. Consultation with DIFD to inform this paper underscored the points above. Examples where DIFD wants to learn more include: What has been learned from the recent outbreak of Ebola in West Africa that could inform future outbreaks, outbreaks of other diseases, or more generally about how health promotion can be reconciled quickly with cultural norms and expectations (such as to attend funerals and lay hands on deceased relatives)? What can be learned from the peace-process in Northern Ireland that could be applicable in South Sudan? What can be learned across evaluations of programmes that use mobile phone technology to change behaviours, both for future mobile-based interventions but also as a platform for understanding how habits can be changed efficiently? Large-scale, multi-component initiatives to improve the education system in a single country — what can the evaluation say about efforts to improve educational outcomes in other countries, and for engaging with public/private organisational cultures to affect change?The aim of this paper is to suggest possible ways to address the issue of learning more from evaluations and make recommendations for how CEDIL could advance this area in the programme of work. To achieve this aim, we conducted consultations with experts from a range of disciplines to identify key concepts and developed a framework for possible approaches. We summarised and contrasted the approaches and reflected on their potential to address DFID’s needs.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.029
metaresearch head score (Gemma)0.040
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Insufficient payload (model declined to judge)
Consensus categoriesMetaresearch, Insufficient payload (model declined to judge)
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Other · Consensus signal: Other
Teacher disagreement score0.394
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0290.040
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0020.003
Science and technology studies0.0000.000
Scholarly communication0.0010.002
Open science0.0010.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0100.009

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.709
GPT teacher head0.648
Teacher spread0.060 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designNot applicable
Domainnot available
GenreOther

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations21
Published2018
Admission routes1
Has abstractyes

Explore more

Same topicEvaluation and Performance AssessmentFrench-language works237,207