Comparing economic mobility with heterogeneity indices : an application to education in Peru
Bibliographic record
Abstract
The long literature on intergenerational transmission of well-being has largely been driven by concerns for inequality of opportunity and the persistence of low levels of wellbeing among certain social groups. A comparative strand of this literature seeks to compare indicators of these transmission mechanisms, i.e. mobility regimes, across societies, regions or time. In this paper I contribute to this literature by suggesting an additional way of comparing mobility regimes with indices of heterogeneity across distributions based on a traditional homogeneity test of multinomial distributions, which is helpful to compare discrete-time transition matrices. The indices measure the degree of dissimilarity between two or more transition matrices controlling for population size and the dimensions of the matrix. The indices provide a good alternative to between-group comparisons based on linear parametric models (chiefly OLS) in which either slope coefficients are compared directly or group dummy variables are interacted with parameters from the models. They also provide complementary information to comparisons based on summary indicators of transition matrices. An application to educational mobility in Peru shows that the transition matrices of males and females are more similar among the youngest cohorts of adults.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.061 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.006 | 0.011 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.001 | 0.004 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".