Ethical Aspects of Measuring Intelligence: Towards Competence and Fairness
Bibliographic record
Abstract
The article is focused on the problem of intelligence measurement, with an emphasis on the ethical aspects of developing and using tests. The history of intelligence measurement provides a variety of examples, problematic from an ethical point of view, which have repeatedly led to negative consequences for both individuals and entire communities. The purpose of this article is to describe current ethical issues in the field of intelligence measurement, their background and historical examples. We discuss the ethical issues in terms of: (1) global approaches to operationalizing intelligence; 2) possible human rights violations resulting from the use of intelligence tests; 3) the fairness of intelligence tests for different groups of respondents; and 4) assessment of test quality in test selection. These issues are examined through the prism of the ethical principles of psychologists, such as respect, honesty, competence, and responsibility. Despite the extensive history of measuring intelligence and research in this area, ethical issues raised decades ago have not lost their relevance. Since ethical questions often do not have clear-cut answers, we believe that engaging in discussions about ethical issues in intelligence testing and exploring potential solutions is itself important and warranted. The content and conclusions of this article may be useful for both researchers and practitioners to make informed decisions in the context of intelligence measurement.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".