Artificial intelligence tools to promote social good in gig markets
Bibliographic record
Abstract
The Artificial Intelligence (A.I.) industry has been essential to creating new jobs for the deployment of real-world solutions. As a result, the implementation of these new jobs involves the execution of multiple human intelligence micro-tasks, such as data labeling tasks for training Machine Learning models. The workers who perform those tasks, also known as crowd workers, usually are independent workers within crowdsourcing platforms. These platforms are subject to the free market, where the forces of supply and demand produce various power dynamics among stakeholders. As a result, disassociation between stakeholders often generates unbalanced power dynamics where workers are paid below minimum wage and are intimidated to keep their reputation or face termination. Within this thesis, I introduce computational techniques to audit the workplace conditions of crowd workers and design tools to address these power imbalances, as a positive and more efficient alternative for the labor conditions of crowd workers. Developing these objectives through the design and evaluation of tools in digital labor platforms, the first "Invisible Labor Tracker'' is a web browser plugin that audits and brings light to an important power dynamic: forcing others to do invisible labor (i.e., do unpaid tasks). Through my tool, I discovered that workers dedicate on average a third of their time to invisible labor, with a very large portion being used to check their payments, as well as being vigilantly on call for "good employers''. The second system, "Reputation Agent'', is an intelligent tool that helps workers to address power dynamics around being unjustly evaluated. The system detects when employers write unfair evaluations about workers, and in such cases, the tool prompts employers to reflect and focus on the performance metrics that are within workers' control. My third system, called "CultureFit'', is an intelligent tool that addresses power dynamics around workers having to change culturally for employers. Instead of forcing workers to change, my system detects a crowd worker's cultural background and then learns the type of cultural interface settings that are best suited to dispatch labor to the worker. Throughout my thesis, I will demonstrate the sustainability of systems that point to a future where A.I. can be used to audit and address power imbalances in the workplace. --Author's abstract
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".