The Labor Market, the Decision to Become an Entrepreneur, and the Firm Size Distribution
Bibliographic record
Abstract
Why do some people become entrepreneurs, how do institutions affect this choice, and how does this affect the firm size distribution and aggregate productivity? This paper addresses this question using a matching model with occupational choice and heterogeneity in both ability as a worker and ex ante unknown productivity of firm start-ups. This rich setting allows to address effects of heterogeneity and diverse types of institutions, like labor market institutions, entry restrictions, taxes, which can possibly differ by firm size and thereby allow addressing informality. Importantly, the model allows for a comparatively flexible lower tail of the firm size distribution and can explain the existence and persistence of small, low-productivity firms with low profits: their owners have low outside options in the labor market. Key effects from a preliminary analysis are the following: labor market conditions affect incentives to start firms differently for workers and the unemployed, with repercussions on aggregate productivity; and they affect the expected value of firm creation due to the possibility of failure. Labor market frictions can have a new effect here: they shape prospective entrepreneurs' value of failure. Given that failure of new projects is common, they can strongly affect not only entry rates, but also the type of firms that enter.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.006 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.011 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".