How Far Will Managers Go to Look Like a Good Steward? An Examination of Preferences for Trustworthiness and Honesty in Managerial Reporting†
Bibliographic record
Abstract
ABSTRACT Growing calls for expanded disclosure on managerial stewardship raise important questions about how finer (i.e., disaggregated) reporting, when paired with discretion over classification, will influence managerial behavior. To study this question, we develop an investment game in which, if the investor chooses to invest, the manager privately observes production costs, chooses their personal pay, and provides a cost report in one of three reporting regimes: aggregated, disaggregated without discretion, or disaggregated with discretion. In Experiment 1, as predicted, managers report lower personal pay under both disaggregated regimes than what they consume under the aggregated regime. Yet, when disaggregated reports allow for discretion, managers misclassify personal pay as production costs to such an extent that their actual consumption is no different than in the aggregated condition. In Experiment 2, we allow managers to choose either an aggregated report or a disaggregated report with discretion. We find that, rather than remaining silent, the vast majority of managers still prefer the opportunity to report on their pay explicitly so that they can use their reporting discretion to appear trustworthy, despite not actually being so. In summary, our evidence suggests a strong weight of preferences for appearing trustworthy in the managers' utility function, a much lower weight for actually being trustworthy, and little evidence that preferences for being honest are strong enough for discretionary disaggregated reporting to curb agency costs. In other words, whether disaggregation can reduce agency costs will depend on managers' reporting discretion. Our findings have important implications for control system designers, financial and sustainability accounting standard setters, and regulators.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".