Bias Control for M-Quantile-Based Small Area Estimators
Bibliographic record
Abstract
Projective outlier-robust M-quantile-based small area estimators can be substantially biased when the sample data contain representative outliers. In this article we propose two new predictive type bias corrected versions of these estimators for continuous and discrete outcomes. Given both area level and individual level outliers in the population, these new estimators are more efficient than the robust-predictive and robust-projective estimators that have been proposed in the small area estimation literature. We also propose two estimators of the prediction mean-squared error of these estimators: one based on Taylor linearization and the other based on a new semi-parametric bootstrap method. We summarize the empirical evidence for these theoretical results in this article, while in the supplementary material we describe in more detail how the properties of these M-quantile-based small area estimators have been assessed in model-based and design-based simulations, as well as in a realistic application focusing on estimation of average income and unemployment rates for local labor market areas in Italy. Supplementary materials for this article are available online, including a standardized description of the materials available for reproducing the work.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".