Convergence without diffusion? A comparative analysis of the choice of performance indicators in tax administration and social security
Bibliographic record
Abstract
This article cross-nationally compares the choice of performance indicators in two core fields of state activity, tax administration and social security. Exploring the selection of performance indicators in six countries (Australia, Canada, Netherlands, Sweden, the UK and the US), the article analyses the driving forces for the choice of particular indicators in the context of national administrative traditions and more recent reform agendas on the one hand and the trend towards international exchange and `benchmarking' on the other hand. The article explores the relative significance and interaction of different driving forces of choice and how this shapes the development and application of performance indicators. To that end, it combines instutionalist approaches with the literature on the mechanisms and effects of international exchange and policy diffusion. Our analysis suggests that existing broad similarities are linked to similarities in core activities and values underlying contemporary public service reforms. Variation in the choice of performance indicators (PIs) reflects domestic factors such as governance arrangements through which broad reform trends are filtered. These arrangements also mediate any direct international learning. Points for practitioners This article aims to contribute to the debate around how organizations could learn from the experience of others in designing performance indicators and management systems. Potential for cross-national and cross-sectional learning is particularly high in categories where a particular organization has not yet developed performance indicators but others have done so already. But any cross-reading from other countries' choices should take into account that the definition and use of performance indicators is to a substantial extent driven by domestic institutional traditions, governance arrangements and wider national approaches to performance management. The design of performance indicators should in particular take into account the accountability relations in which agencies are embedded.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.002 |
| Science and technology studies | 0.000 | 0.004 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".