When Do Punishment Institutions Work
Bibliographic record
Abstract
While peer punishment sometimes motivates increased cooperation, it sometimes reduces cooperation. We use a lab experiment to study why punishment sometimes fails. We begin with a gift exchange game with punishment as it has typically been implemented therein since punishment has often backfired in this game. We modify two features of punishment that could increase its efficacy: punishment's strength and its timing (whether the punisher publicly pre-commits to punishment or acts after the punishee). We replicate the result that peer punishment in gift exchange games can reduce cooperation, but show that this bad outcome disappears if punishment is more powerful. This does not seem primarily due to punishment's threat leading to spiteful behavior: we find little evidence of spite, and the same punishment does not perform better when it is chosen after the fact. We find two main reasons that punishment decreases cooperation: lower wages are offered (a stick is substituted for a carrot); and many punishers don't design punishment to properly incentivize high effort, particularly when punishment is weak in power. Punishment that is not publicly pre-committed is not effective in this game, even though this kind of punishment is similar to that used in public good games in the literature where punishment does seem to increase cooperation. The only punishment institution that increases cooperation is high-power punishment that is publicly pre-committed, which works through strong incentives rather than reciprocity. Finally, the existence of a punishment institution often decreases social surplus (when punishment-related losses are considered), although it may eventually increase social surplus if it is powerful and publicly pre-committed.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".