A Crowd-Based Evaluation Model in a Business School Setting
Bibliographic record
Abstract
In nearly every social and economic setting evaluations of contribution are required for status attribution and tangible rewards and penalties. In some settings evaluations are made exclusively by status or role superiors. In many others evaluations are made by peers as well as superiors, with the evaluation of one often influencing that of the other. A significant body of literature has also determined how ascriptive characteristic including race, gender, and homophily influence interpersonal evaluation. Evaluation has significant ramifications to the extent it impacts the social, economic, and psychological welfare of individuals. The integrity of the evaluation system writ large also has implications for organizational justice and thus group or organizational processes. Understanding how to create evaluation systems that are equitable is thus a fundamental issue all evaluators, leaders and organizations must confront. In this paper we present results from an experiment in which we implement a crowd-based evaluation system in a business school setting. Our model includes detailed ascriptive and achieved characteristics of the evaluator, evaluated, and evaluator and evaluated (e.g., homophily) to determine whether crowd-based evaluation is a potentially equitable basis of evaluation.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.017 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".