Bench Stacking and Biases: The ICJ’s Partial Decision in Yugoslavia v. NATO Members in Comparative Perspective
Bibliographic record
Abstract
Since the Genocide Convention was adopted by the General Assembly in 1948, eight cases have been brought to the ICJ by invoking Article IX of the Genocide Convention as a basis of the Court’s jurisdiction. Only two cases have reached their conclusion based on the merits of the case, with others decided during preliminary proceedings, while still others remain ongoing. There have been numerous studies of ICJ impartiality, with particular focus on judges’ voting records, using large amounts of data to discern any trends and biases. This article is the first attempt to comparatively analyze ICJ genocide cases using an interpretive lens through a close reading and detailed textual analysis of the majority opinion in five cases, along with elements of oral proceedings, declarations, and separate and dissenting opinions. By comparing the ICJ’s decisions during the provisional measures phase in Bosnia v. Serbia, Yugoslavia v. NATO members, The Gambia v. Myanmar, Ukraine v. Russia, and South Africa v. Israel, evidence suggests the ICJ delivered a biased decision against Yugoslavia due to the Court’s handling of genocidal intent, its omission of significant principles that were cited in the other four cases, and the stacking of the bench with judges from the respondent states.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".