Bibliographic record
Abstract
This dissertation provides a comprehensive analysis of the Canadian gender wage gap over the past two decades, employing modern methodologies and tools.\n\nIn the first chapter, selection bias and its impact on the entire earning distribution are examined. A selection-corrected quantile regression is utilized to provide a more accurate depiction of the gender wage gap distribution. The simulation of potential government child care benefits as an instrument helps address the selection bias issue. Findings reveal the persistent but inconsistent effects of selection bias across wage quantiles and time. The presence of negative selection for women entering the workforce is identified, and the absence of this bias would result in an even higher unexplained portion of the gender wage gap.\n\nMoving to the second chapter, a thorough investigation of the heterogeneity of the gender wage gap is conducted. High-dimensional models are employed to explore the diverse factors contributing to the wage gap. Advanced machine learning algorithms are utilized as robustness checks to address possible multicollinearity problems. The analysis reveals significant reductions in the gender wage gap attributed to age and several occupations, while penalties related to family structure persist.\n\nFinally, the third chapter explores the under-researched area of job-education mismatch and its impact on the gender wage gap. The study focuses on the differences in vertical and horizontal matching between women and men. Self-reported measurements of both vertical and horizontal mismatch, as well as an objective index of horizontal mismatch, are utilized. Results indicate that, unlike other countries, vertical mismatch does not contribute significantly to the gender wage gap. Furthermore, the role of horizontal mismatch is economically insignificant in relation to the overall gap.\n\nThis dissertation enhances our understanding of the Canadian gender wage gap by addressing important aspects such as selection bias, heterogeneity, and job-education mismatch. The findings contribute to the existing literature on gender inequality, offering valuable insights for policy interventions and future research in this field.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.007 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.008 | 0.014 |
| Science and technology studies | 0.007 | 0.001 |
| Scholarly communication | 0.004 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.051 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".