Machine learning and the labour market: A portrait of occupational and worker inequities in Canada
Bibliographic record
Abstract
ABSTRACT Introduction Machine learning (ML) is increasingly used by Canadian workplaces. Concerningly, the impact of ML may be inequitable and disrupt social determinants of health. The aim of this study is to estimate the number of workers in occupations highly exposed to ML and describe differences in ML exposure represents according to occupational and worker sociodemographic factors. Methods Canadian occupations were scored according to the extent to which they were made up of job tasks that could be performed by ML. Eight years of data from Canada’s Labour Force Survey were pooled and the number of Canadians in occupations with high or low exposed to machine learning were estimated. The relationship between gender, hourly wages, educational attainment and occupational job skills, experience and training requirements and ML exposure was examined using stratified logistic regression models. Results Approximately, 1.9 million Canadians are working in occupations with high ML exposure and 744,250 workers were employed in occupations with low ML exposure. Women were more likely to be employed in occupations with high ML exposure than men. Workers with greater educational attainment and in occupations with higher wages and greater job skills requirements were more likely to experience high ML exposure. Women, especially those with less educational attainment and in jobs with greater job skills, training and experience requirements, were disproportionately exposed to ML. Conclusion ML has the potential to widen inequities in the working population. Disadvantaged segments of the workforce may be most likely to be employed in occupations with high ML exposure. ML may have a gendered effect and disproportionately impact certain groups of women when compared to men. We provide a critical evidence base to develop strategic responses that ensure inclusion in a working world where ML is commonplace. KEY MESSAGES What is already known on this topic The Canadian labour market is undergoing an artificial intelligence (AI) revolution that has the potential to have widespread impact on a range of occupations and worker groups. It is unclear how which the adoption of machine learning (ML), an AI subfield, within the working world might contribute to inequities within the labour market. What this study adds Segments of the workforce which have been previously disadvantaged may be most likely to work in occupations most likely to be affected by ML. ML may have a gendered effect and disproportionately impact some groups of women when compared to men. How this study might affect research, practice or policy Findings can inform targeted policies and programs that optimize the economic benefits of ML while addressing disparities that can emerge because of the adoption of the technology on workers.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.004 | 0.010 |
| Science and technology studies | 0.005 | 0.002 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".