Bibliographic record
Abstract
In policy discourse, the assessment of public happiness is at the heart of the happiness movement’s insistence that policy evaluators must look ‘beyond GDP’. But is this radical enough? Why substitute or complement one measurement exercise with another when the key weakness may have been the overemphasis on measurement itself, rather than the choice of measure? People often overrate numbers as the core tool in the process of political persuasion. They trot out claims like ‘if it isn’t counted, it doesn’t count’. But in reality, policies at all levels are much more influenced by stories, pictures, personal charisma, and arguments than they are by numbers. Towards the end of 2010, there was feverish debate in the UK media on the wisdom and sincerity of Prime Minister David Cameron’s decision, at a time of recession and massive cutbacks in public spending, to require the Office of National Statistics to conduct regular happiness surveys as part of national well-being assessment. What’s odd is that, after so many decades of happiness surveys, this should be seen as a controversially novel idea. But the Canadian economist John Helliwell, who has for many years tirelessly tried to persuade his government to give happiness assessment a higher profile in policy making, argues that ‘What is or could be dramatically different in the UK is for the government not just to undertake more widespread and thorough collection of subjective well-being data, but also to give them a central place in the choice and evaluation of public policies. That would be a global first’ (Stratton, 2010).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".