Bibliographic record
Abstract
Introduction In this chapter we give a brief overview of the main theoretical results and approximations used in this book. These approximations are derived from the theory of higher order likelihood asymptotics. We present these fairly concisely, with few details on the derivations. There is a very large literature on theoretical aspects of higher order asymptotics, and the bibliographic notes give guidelines to the references we have found most helpful. The building blocks for the likelihood approximations are some basic approximation techniques: Edgeworth and saddlepoint approximations to the density and distribution of the sample mean, Laplace approximation to integrals, and some approximations related to the chi-squared distribution. These techniques are summarized in Appendix A, and the reader wishing to have a feeling for the mathematics of the approximations in this chapter may find it helpful to read that first. We provide background and notation for likelihood, exponential family models and transformation models in Section 8.2 and describe the limiting distributions of the main likelihood statistics in Section 8.3. Approximations to densities, including the very important p * approximation, are described in Section 8.4. Tail area approximations for inference about a scalar parameter are developed in Sections 8.5 and 8.6. These tail area approximations are illustrated in most of the examples in the earlier chapters. Approximations for Bayesian posterior distribution and density functions are described in Section 8.7. Inference for vector parameters, using adjustments to the likelihood ratio statistic, is described in Section 8.8.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.048 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.004 | 0.006 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.002 | 0.005 |
| Insufficient payload (model declined to judge) | 0.030 | 0.015 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".