Can ChatGPT Outperform Humans in Faking a Personality Assessment While Avoiding Detection?
Bibliographic record
Abstract
ABSTRACT Large language models (LLMs), such as ChatGPT, have reshaped opportunities and challenges across various fields, including human resources (HR). Concerns have arisen about the potential for personality assessment manipulation using LLMs, posing a risk to the validity of these tools. This threat is a reality: recent research suggests that many candidates are using AI to complete pre‐hire assessments. This study addresses this problem by examining whether ChatGPT can outperform humans in faking personality assessments while avoiding detection. To explore this, two experiments were conducted focusing on assessing job‐relevant traits, with and without coaching, and with two methods of identifying faking, specifically using an impression management (IM) measure and an overclaiming questionnaire (OCQ). For each study, we used responses from 100 working adults recruited via the Prolific platform, which were compared to 100 replications from ChatGPT. The results revealed that while ChatGPT showed some ability to manipulate assessments, without coaching it did not consistently outperform humans. Coaching had a minimal impact on reducing IM scores for either humans or ChatGPT, but reduced OCQ bias scores for ChatGPT. These findings highlight the limitations of current faking detection measures and emphasize the need for further research to refine methods for ensuring the integrity of personality assessments in HR, particularly as artificial intelligence becomes more available to candidates.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.020 | 0.073 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".