On the role and limitations of experimental and behavioural economics
Bibliographic record
Abstract
Ignorance is a formidable foe, and to have hope of even modest victories, we economists need to use every resource and every weapon we can muster, including thought experiments (theory), and the analysis of data from nonexperiments, accidental experiments, and designed experiments. We should be celebrating the small genuine victories of the economists who use their tools most effectively, and we should dial back our adoration of those who can carry the biggest and brightest and least-understood weapons. We would benefit from some serious humility, and from burning our ‘Mission Accomplished’ banners. It’s never gonna happen. Edward Leamer [2010, p. 44] Behavioural studies by and large do not necessarily assume that people always behave by the dictates of standard theory, and especially in the early days repeatedly showed deviations etc., but increasingly are concerned with looking at the source of the departures and with their implications, etc. Jack Knetsch [I]t would be useful for theory to identify behavior for which the theory cannot account, in the sense that the observations would force the theorist to reconsider. This would ensure that the theory is not performing well by ‘theorizing to the test’ … Similarly, it would be helpful to have the experimental design indicate which outcomes would be regarded as a failure as well as which would be considered a success. This question appears to be trivial in many cases, with success and failure riding on the statistical significance of an estimated parameter. However, one of the advantages of experimental work is the ability to control the environment and design the tests. This allows us to direct attention away from issues of statistical significance and toward issues of economic importance. The strength of the experiment will often be reflected in the content of this ‘failure’ category. … [I]t is important that both theoretical models and interpretations of experimental results be precise enough to apply beyond the experimental situation from which they emerge. This allows links to be made that multiply the power of single studies. Larry Samuelson [2005, pp. 100–1]
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.041 | 0.056 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.002 | 0.043 |
| Scholarly communication | 0.008 | 0.018 |
| Open science | 0.003 | 0.004 |
| Research integrity | 0.004 | 0.009 |
| Insufficient payload (model declined to judge) | 0.009 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".