Bibliographic record
Abstract
In survey sampling, we are interested in inferring on a finite population of, for example, households, businesses or electricity users, based on a sample of only few hundred or few thousands units. The sampling procedure depends on our a priori knowledge of the population. In the case of a single sampling frame, the sample can be obtained using a direct sampling procedure. In some cases, one must recourse to multiple sampling frames in order to cover the whole population. A sample is then selected within each frame and the goal is to combine them to obtain an accurate estimate. When no sampling frame is available, indirect sampling procedures are typically used. Also, sampling methods offer an interesting alternative in the context of large volumes of data when it is required to reduce the dimension, which in turns permits data exploitation. Response rates have been steadily decreasing over time in household surveys. Efforts have been made for following up the nonrespondents in order to increase the response rates. Now, the objective consists of targeting the nonrespondents in order to balance the characteristics of the respondents at the end of the process, which may be useful for controlling the risks of bias. After data collection, nonresponse is treated at the estimation stage using some models.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.019 | 0.073 |
| Meta-epidemiology (narrow) | 0.004 | 0.002 |
| Meta-epidemiology (broad) | 0.004 | 0.004 |
| Bibliometrics | 0.004 | 0.002 |
| Science and technology studies | 0.003 | 0.003 |
| Scholarly communication | 0.007 | 0.005 |
| Open science | 0.005 | 0.002 |
| Research integrity | 0.012 | 0.017 |
| Insufficient payload (model declined to judge) | 0.034 | 0.022 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".