Learning and earning in Africa : where are the returns to education high?
Bibliographic record
Abstract
This paper investigates the role of learning - through formal schooling and time spent in the labor market - in explaining labor market outcomes of urban workers in Ghana and Tanzania. We investigate these issues using a new data set measuring incomes of both formal sector wage workers and the self-employed in the informal sector. In both countries we find significant, convex returns to education and large earnings differentials between sectors when we pool the data and do not control for selection. In Ghana there is a particularly steep age-earnings profile. We investigate how far a Harris-Todaro model of market segmentation or a Roy model of selection can explain the patterns observed in the data. We find highly significant differences across occupations and important effects from selection in both countries. The data is consistent with a pattern by which higher ability individuals queue for the high wage formal sector jobs such that the age earnings profile is convex for the self-employed in Ghana once we control for selection. The returns to education are far higher in the large firm sector than in others and in this sector they are linear not convex. In both countries there is clear evidence of convexity in the returns to education for the self-employed and here the average returns are low. The data used in this paper were collected by the Centre for the Study of African Economies, Oxford, in collaboration with the Ghana Statistical Office (GSO) and the Tanzania National Bureau of Statistics (NBS). The research, and the surveys on which it is based, has been funded by the Department for International Development (DfID) and the Economic and Social Research Council (ESRC) of the UK and by the IDRC in Canada. We are greatly indebted to numerous collaborators for enabling this data to be collected, particularly Emilian Karugendo and Trudy Owens in Tanzania, and Moses Awoonor-Williams, Geeta Kingdon and Andrew Zeitlin in Ghana. Andrew Kerr provided valuable assistance with the coding. An earlier version of this paper benefited from the input of seminar participants in IZA/Berlin and Cornell. We have discussed the points made in this paper extensively with Mans Soderbom who has offered many valuable suggestions. All errors are ours.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.004 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".