Quantifying the Creator Economy: A Large-Scale Analysis of Patreon
Bibliographic record
Abstract
In recent years, the “creator economy” has emerged as a disruptive force in creative industries. Independent creators can now reach large and diverse audiences through online platforms, and membership platforms have emerged to connect these creators with fans who are willing to financially support them. However, the structure and dynamics of how membership platforms function on a large scale remain poorly understood. In this work, we develop an analysis framework for the study of membership platforms and apply it to the complete set of Patreon pledges exceeding $2 billion since its inception in 2013 until the end of 2020. We analyze Patreon activity through three perspectives: patrons (demand), creators (supply), and the platform as a whole. We find several important phenomena that help explain how membership platforms operate. Patrons who pledge to a narrow set of creators are more loyal, but churn off the platform more often. High-earning creators attract large audiences, but these audiences are less likely to pledge to other creators. Over its history, Patreon diversified into many topics and launched higher-earning creators over time. Our analysis framework and results shed light on the functioning of membership platforms and have implications for the creator economy.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.013 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.005 | 0.005 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.005 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.008 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".