Exploring Fairness Debt Through Evidence from Studies on Algorithmic Discrimination
Bibliographic record
Abstract
Context. The increasing deployment of artificial intelligence in decision-making processes has highlighted biases embedded in algorithms, leading to discriminatory outcomes that impact society. This phenomenon has led to the proposal of fairness debt, a type of software debt that emerges when biases are not addressed early in development and accumulate over time, creating costly, complex issues that can perpetuate societal inequities. Goal. This study aimed to explore the concept of fairness debt within algorithm-driven decision-making systems, identifying how deferred fairness considerations lead to accumulated biases and affect society. Method. We conducted a mapping study, analyzing 86 papers on algorithmic discrimination to classify root causes, effects, and instances of discrimination associated with fairness debt. Findings. Our findings indicate that several root causes, such as training bias, historical bias, and design bias, contribute to the accumulation of fairness debt, with manifestations in issues like sexism, racism, ableism, and ageism within systems. These accumulated biases can exacerbate social inequalities, limit algorithmic reliability, and perpetuate stereotypes, demonstrating fairness debt’s extensive societal impact. Discussions. Our study demonstrates that fairness debt extends beyond technical issues, encompassing broader societal consequences that demand proactive mitigation strategies, emphasizing the risks associated with deferring fairness considerations and highlighting the importance of fairness as a foundational principle in modern software systems, in particular AI systems. Conclusion. We explore algorithm discrimination through the lens of fairness debt, showcasing e that social concerns of algorithms can be interpreted as main component of software engineering practice, considering the socio-technical factors embedded in it and helping researchers and developers exploring discrimination in their systems and practices.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".