An empirical investigation of the evaluators' scoring of vendors' responses to an RFP of a large healthcare system
Bibliographic record
Abstract
Request for Proposal (RFP) is a solicitation of proposals from vendors and they are often judged by human experts from varying backgrounds and experiences. This is typically done because large technical RFPs require a diverse group of evaluators who will bring their skills and experience to bear. However, different people with different backgrounds may evaluate proposals in different ways. In this paper, we examine the variability between expert proposal ratings through a case study, and try to determine how diversity of expertise affects the consistency of scoring results. We evaluate this by using scores of an RFP of a large industrial system by 20 different human experts from six different backgrounds. Our results suggest that there is not always a significant difference among raters in the scores given to vendors. Thus, human experts with varying but related backgrounds may judge in a similar manner and can be used to evaluate the RFP, provided there is a high level agreement between experts and the RFP is created through their consensus. This empirical evidence does not exist in the literature and it is a novel contribution of this paper.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.106 | 0.365 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".