More power to you: properties of a more powerful event study methodology
Bibliographic record
Abstract
Purpose The purpose of this paper is to demonstrate with real data the enhanced statistical power of a GLS‐based event study methodology that requires the same input data as the traditional tests. Design/methodology/approach The paper uses full sample, subsample and simulated modified sample analyses to compare the statistical power of the GLS methodology with traditional methods. Findings The paper finds that it is often the case that traditional tests will not reject the null when a GLS‐based test may (strongly) reject the null. The power of the former is poor. Practical implications There are many published event studies where the null is not rejected. This may be because of the phenomenon being tested but it may also be because of the lack of power of traditional estimators. Hence, rerunning them with the authors' more powerful test is likely to reject some currently well‐accepted null hypotheses of no event effect, stimulating new research ideas. Moreover, as individual stocks have become more volatile, the additional power of the authors' methodology to detect abnormal performance for recent and future events becomes even more important. Originality/value There are more than 500 event studies in the top finance journals, which can broadly be split into two subgroups: contemporaneous shocks like changes in regulation and non‐contemporaneous events like mergers. GLS contemporaneous modeling of covariances in the former showed little efficiency gains. The paper's GLS modeling of variances for the latter demonstrates potentially huge effects. Practitioners should be skeptical of prior results accepting the null of no event effect and incorporate GLS to be confident of their future findings.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".