Bibliographic record
Abstract
My dissertation consists of three chapters in the area of labor economics. The first chapter, written jointly with Derek Stacey, develops a methodology to estimate occupation-specific multidimensional skill measures that exploit variation in both wages and multidimensional skill ratings. The resulting skill measures use a common interval scale (log wage units), which allows for straightforward comparisons across occupations, aggregation across skills, and interpretation of the various components of each occupation’s skill portfolio. The second chapter uses these skill measures to investigate the role multidimensional skills can play in explaining the gender wage gap. I find that different types of non-manual skill are remunerated to men and women at different rates. More specifically, women are paid more for their literacy skills, while men are paid more for their cognitive skills. I then show this skill-based wage gap is primarily borne by differences in the returns to skills as opposed to differences in occupation selection. The third chapter uses the multidimensional skill measures to analyze the relationship between worker skills and occupation skill requirements, with a focus on skill mismatch. Using data from the NLSY, I first document patterns on skill mismatch in workers’ initial jobs that suggest workers, on average, are not in occupations that primarily use their comparative skill advantage. As workers gain experience they, on average, do not correct this initial skill mismatch by transitioning occupations. Two possible explanations for this skill immobility is that workers are involuntarily constrained or, due to skill accumulation, are rationally staying in an initial mismatch. To explain observed skill immobility, I develop a model of occupational choice with heterogeneous worker skills, matching frictions, and skill accumulation where occupations differ in their skill intensities. The model is calibrated to match occupational skill requirements and is used to decompose worker mismatch into voluntary and involuntary mismatch. I find a quarter of end of career mismatches can be attributed to workers whose inherent skills have transformed to match their initial mismatch occupation.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.008 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.003 | 0.004 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.058 | 0.013 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".