Case Study: The International Criminal Tribunal for the Former Yugoslavia’s Court Transcripts in Bosnian/Croatian/Serbian—Part 1: Needs, Feasibility, and Output Assessment
Notice bibliographique
Résumé
International Criminal Tribunal for the Former Yugoslavia (ICTY) remains the most important organization for the past, the present, and the future of the former Yugoslavia. Faced with a country that always lived under totalitarian regimes with very little insight into actions of the groups and individuals who reaped unthinkable havoc on each other at the end of the twentieth century, the ICTY set undisputable historical record about events that took place during the 1991–1999 wars and put the country on an excellent track towards transformation for the better. But even 28 years since the establishment of the ICTY, the former Yugoslavia remains the hotbed of nationalism, ethnic divisions, genocide denial, and genocide justification. Court transcripts belong to the category of the permanent court record. The ICTY court transcripts have only been made in English and French, but not in Bosnian/Croatian/Serbian (B/C/S), the languages of the former Yugoslavia. This paper is going to examine the needs for the ICTY court transcripts in the B/C/S, could they have been made in the B/C/S from the very beginning of the institution and whether the existing ICTY court transcripts in the B/C/S are up to par for any of its audiences.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,014 | 0,028 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,004 | 0,004 |
| Études des sciences et des technologies | 0,015 | 0,004 |
| Communication savante | 0,007 | 0,002 |
| Science ouverte | 0,003 | 0,005 |
| Intégrité de la recherche | 0,004 | 0,004 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,004 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».