Tax Risk as the Likelihood of an Unfavorable Settlement with Tax Authorities
Bibliographic record
Abstract
In this paper, we examine tax risk as the likelihood that a firm will face an unexpected and negative tax outcome. More specifically, we predict the probability that a firm will have an unfavorable settlement with tax authorities, what we define as a tax loss event. We further consider whether market participants anticipate this risk and consequently how they react to the tax loss event. We combine an industry-adjusted measure of high cash effective tax rate (Cash ETR) values with financial statement verification of a tax settlement to identify a sample of tax loss event firms. To predict the likelihood of these events, we develop a logistic model from a parsimonious set of factors that are associated with tax risk. Empirical tests show that our model provides strong predictions of a tax loss event. Additional tests show that our model performs well against a matched set of holdout control observations – that is, observations with unusually high Cash ETR values yet no tax settlement. In related valuation tests, the abnormal returns are negative in the event year for firms with a tax loss event and high ETR volatility, suggesting that investors may not recognize the risk beforehand. These results imply that the market punishes firms that appear to poorly manage the financial reporting of tax risk, yet only when they subsequently reveal a significant tax settlement. Our research adds to the emerging literature that examines the tax risk within firms. Specifically, we develop a prediction model useful for various firm stakeholders, as well as provide additional insight into the factors that create risk for the firm and how the market reacts to the revelation of such risks.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.017 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.002 | 0.003 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.008 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".