Skill Management in Large‐Scale Service Marketplaces
Bibliographic record
Abstract
Large‐scale, web‐based service marketplaces have recently emerged as a new resource for customers who need quick resolutions for their short‐term problems. Due to the temporary nature of the relations between customers and service providers (agents) in these marketplaces, customers may not have an opportunity to assess the ability of an agent before their service completion. On the other hand, the moderating firm has a more sustained relationship with agents, and thus it can provide customers with more information about the abilities of agents through skill screening mechanisms. In this study, we consider a marketplace where the moderating firm can run two skills tests on agents to assess if their skills are above certain thresholds. Our main objective is to evaluate the effectiveness of skill screening as a revenue maximization tool. We, specifically, analyze how much benefit the firm obtains after each additional skill test. We find that skill screening leads to negligible revenue improvements in marketplaces where agent skills are highly compatible and the average service times are similar for all customers. As the compatibility of agent skills weakens or the customers start to vary in their processing time needs, we show that the firm starts to experience sizable improvements in revenue from skill screening. Apparently, the firm can reap the most of these substantial benefits when it runs only one test. For instance, in marketplaces where agents posses uncorrelated skills, the second skill test only brings an additional 2% improvement in revenue. Accounting for possible skill screening costs, we then show the optimality of offering only one test when the compatibility between agent skills is sufficiently low. The results of this study also have important implications in terms of the right level of intervention in the marketplaces we study.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".