Bibliographic record
Abstract
Dynamic panel data models can suffer greatly from incidental parameter bias due to correlation between past realizations of the data and the unobserved heterogeneity, and this bias is a function of included regressors.This paper uses simulation-based methods that require explicit models and sets of assumptions to obtain consistent point estimates and exact confidence sets.A parametric discontinuous starting value is assumed for simulated series that jointly allows for stationary and unit root processes, where only the stationary case was considered in Gouriéroux, Phillips, and Yu (2010).This discontinuous assumption leads to least squared dummy variable (LSDV) estimator that are nuisance parameter free and location-scale invariant.These properties are conferred to the indirect inference objective function (IIOF) used to obtain bias-corrected estimates.Discontinuities are problematic for traditional asymptotic methods of constructing confidence sets.To account for this the indirect confidence set inference method is introduced, which uses a second round of Monte Carlo simulations [Dufour (2006)] to calibrate the distribution of the IIOF.The confidence set is constructed with test inversion, so the parameters are set to known values, the model is tested at that point, and all points that fail to reject the null hypothesis are in the confidence set.The confidence set is exact and level correct, since the IIOF is pivotal and both simulation rounds are exchangeable under the null.Adding regressors into panel data models can distort estimates, as this paper demonstrates with respect to the X-differencing method of Han, Phillips, and Sul (2014) with regressors.By introducing a model augmentation approach, the influence of regressors are corrected.The augmentation uses a projection of the regressors for A special thank you to my supervisor Lynda Khalaf, that without your patience, support, and knowledge I would likely not have been able to advance as quickly and cleanly as I have.I would also like to thank Russell Davidson (McGill) for pointing out that the initial observation could be random.I would like to thank Jeffery Wooldridge, for our brief discussion in Budapest, and for indicating that he had given a block-diagonal Mundlak device some consideration in response to query at a conference a few years prior.I would also like to thank my comrades in the PhD program, but specifically our microeconomics, macroeconomics and econometrics study groups: Sarah Mohan, Duangsuda (Neat) Sopchokchai, Anand Acharya, Bogdan Urban, and Nyamekye Asare.A thank you to some others who have helped along the way: Marie-
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.025 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.001 | 0.004 |
| Scholarly communication | 0.004 | 0.005 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.003 | 0.006 |
| Insufficient payload (model declined to judge) | 0.013 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".