Top incomes and the gender divide
Bibliographic record
Abstract
In the recent research on top incomes, there has been little discussion of gender. How many of the top 1 and 10 per cent are women? A great deal is known about gender differentials in earnings, but how far does this carry over to the distribution of total incomes, bringing self-employment and capital income into the picture? We investigate the gender divide at the top of the income distribution using tax record data for a sample of eight countries with individual taxation. We show that women are under-represented at the top of the distribution. They account for between a fifth and a third of those in the top 10 per cent. Higher up the income distribution, the proportion is lower, with women constituting between 14 and 22 per cent of the top 1 per cent. The presence of women in the top income groups has generally increased over time, but the rise becomes smaller at the very top. As a result, the gradient with income has become more marked: the under-representation of women today increases more sharply. Examination of the shape of the income distribution by fitting a Pareto distribution shows that at the end of the period women disappear faster than men as one moves up the income scale in all countries. In this sense, there appears to be something of a "glass ceiling" for women. In the case of Canada, Denmark, Norway and New Zealand, there appears to have been a reversal over time, with the slope of the upper tail having been steeper for women in the past. In seeking to explain this, we highlight the role of income composition, where we show that there have been significant changes over time, underlining the fact that it is not sufficient to look only at earned income.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.002 | 0.003 |
| Scholarly communication | 0.002 | 0.003 |
| Open science | 0.000 | 0.002 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.015 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".