Should Workers Care about Firm Size?1 Ana Ferrer University of British Columbia2
Bibliographic record
Abstract
The question of wage differentials by firm size has been studied for several decades with no commonly accepted explanations for why large firms pay more. In this paper, we reexamine the relationship between firm-size and wage outcomes by estimating the returns to unmea-sured ability between large and small firms. Our empirical methodology, based on non linear instrumental variable estimations, allows us to directly estimate the returns to unmeasured ability by firm size and therefore to test the two main theories of wage determination proposed to explain the relationship between firm size and wages, namely ability sorting and job screen-ing. We use data from the Survey of Labour and Income Dynamics (SLID) which provides longitudinal information on workers and firms characteristics including establishment and firm size. We find significant differences in the returns to unmeasured ability across firm size. In particular, we find that the returns to unmeasured ability seem to follow a non linear pattern. The returns to unmeasured ability are significantly higher in medium size (above 500 but below 1000 workers) firms relative to small firms. However, the returns to unmeasured ability are not significantly greater in large firms relative to medium or small firms. Overall, it seems that ability sorting dominates for moves from small to medium size firms in that ability is more productive and therefore more rewarded in the latter than the former. On the other hand, when firms become “too large”, the monitoring costs hypothesis seems to dominate in that ability is not more rewarded than in smaller firms.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.013 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.003 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.035 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".