Examining the assumption of measurement invariance in job performance ratings across time: The role of rater experience
Bibliographic record
Abstract
Abstract Performance appraisals are widely used in organizations and most typically involve raters evaluating groups of subordinates along a set of items designed to represent job performance over a predetermined period (e.g., annually). A defining but often overlooked characteristic of performance appraisals is that they are cyclical. Since raters conduct appraisals over many cycles, it may be that measures of job performance are not equivalent across time. This is important because changes or differences in aggregated performance ratings can only be meaningfully interpreted if raters' definitions of job performance, interpretation of what the items mean, and their view of what constitutes the different levels of performance remain unchanged over time, unaffected by their experience with appraisals. Although critical to the interpretation of job performance scores, measurement invariance concerns are generally absent from the literature. The current research investigated the extent to which rater experience affected the conceptualization and measurement of performance using performance data from a major South American company which comprised information from raters and ratees through several appraisal cycles. In the between‐rater design, measurement invariance was analyzed using ratings of one performance appraisal cycle from 514 raters divided into groups according to their level of experience. The within‐rater design analyzed ratings from the same 80 raters in their first three appraisal cycles. In the between‐rater analysis, data supported measurement invariance across raters with different levels of experience. Results from the within‐rater analysis suggested that the job performance factor structure was not the same across cycles. Implications for research and practice are discussed.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.197 | 0.508 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.005 |
| Scholarly communication | 0.003 | 0.004 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".