Community Interpreting Database Pilot Corpus (ComInDat)
Notice bibliographique
Résumé
Audio and video recordings of various types of community interpreted discourse (doctor-patient communication, simulated doctor-patient communication, courtroom communication) in German (simulated and authentic doctor-patient communication) and US (courtroom communication) institutions with varying community languages. Video recordings only exist for the simulated communication. For the authentic interpreted doctor-patient communication, no audio files will be made available. The ComInDat pilot corpus contains sample data from three different projects: the DiK corpus of Portuguese/German and Turkish/German interpreted doctor-patient communication in hospitals (Bührig & Meyer 2004), he IiSCC-corpus, a corpus of interpreted court proceedings in different language constellations (Spanish/English, Russian/English, Haitian Creole/English and Polish/English) (Angermeyer 2006), a corpus of simulated interpreted doctor-patient interactions in different language constellations (Russian/German, Polish/German and Romanian/German) from a training seminar for bilingual nursing staff ("SimDiK", Bührig, Kliche, Meyer & Pawlack 2012). More information about the background of the corpus and the details of its design can be found in (Angermeyer, Meyer & Schmidt 2012). For more information about the project, please contact Philipp Angermeyer. Angermeyer, P., Meyer, B. and Schmidt, T. (2012). Sharing Community Interpreting Corpora: A pilot study. In: Schmidt, T. and Wörner, K. (eds.) Multilingual Corpora and Multilingual Corpus Analysis. Amsterdam: Benjamins, 275-294. <strong>CLARIN Metadata summary for Community Interpreting Database Pilot Corpus (ComInDat) (CMDI-based)</strong> <strong>Title: </strong>Community Interpreting Database Pilot Corpus (ComInDat)<br> <strong>Description: </strong>Audio and video recordings of various types of community interpreted discourse (doctor-patient communication, simulated doctor-patient communication, courtroom communication) in German (simulated and authentic doctor-patient communication) and US (courtroom communication) institutions with varying community languages. Video recordings only exist for the simulated communication. For the authentic interpreted doctor-patient communication, no audio files will be made available.<br> <strong>Publication date: </strong>2013-06-10<br> <strong>Data owner: </strong> Philipp Angermeyer, Department of Languages, Literatures and Linguistics / York University / 4700 Keele Street / Canada M3J 1P3, pangerme@yorku.ca, Kristin Bührig, Institut für Germanistik I / Von-Melle-Park 6 / D-20146 Hamburg, kristin.buehrig@uni-hamburg.de, Bernd Meyer, Arbeitsbereich Interkulturelle Kommunikation / Fachbereich 06: Translations-, Sprach- und Kulturwissenschaft / Johannes Gutenberg-Universität Mainz / An der Hochschule 2 / D-76726 Germersheim, meyerb@uni-mainz.de<br> <strong>Contributors: </strong> Philipp Angermeyer, Department of Languages, Literatures and Linguistics / York University / 4700 Keele Street / Canada M3J 1P3, pangerme@yorku.ca (compiler), Kristin Bührig, Institut für Germanistik I / Von-Melle-Park 6 / D-20146 Hamburg, kristin.buehrig@uni-hamburg.de (compiler), Bernd Meyer, Arbeitsbereich Interkulturelle Kommunikation / Fachbereich 06: Translations-, Sprach- und Kulturwissenschaft / Johannes Gutenberg-Universität Mainz / An der Hochschule 2 / D-76726 Germersheim, meyerb@uni-mainz.de (compiler)<br> <strong>Project: </strong> The Integration of Text, Sound, and Image into the Corpus-Based Analysis of Interpreter-Mediated Interaction<br> <strong>Keywords: </strong> community interpreting, doctor-patient communication, courtroom communication, EXMARaLDA<br> <strong>Languages: </strong> German (deu), English (eng), Spanish (spa), Turkish (tur), Polish (pol), Portuguese (por), Romanian (ron), Russian (rus), Haitian (hat)<br> <strong>Size: </strong> 54 speakers (35 female, 16 male, 3 unknown), 14 communications, 12 recordings, 83 minutes, 17 transcriptions, 35051 words<br> <strong>Annotation types: </strong> transcription (manual): HIAT/CHAT, deu: German translation, eng: English translation, k: free comment, lang: utterance language, sup: suprasegmental information, trans: utterance translation status, akz: accentuation/stress, pol: Polish translation<br> <strong>Temporal Coverage: </strong> 1999-07-01/2010-03-09<br> <strong>Spatial Coverage: </strong> Hamburg, DE; New York, US; Neumünster, DE<br> <strong>Genre: </strong> discourse<br> <strong>Modality: </strong> spoken<br> <strong>References: </strong> Angermeyer, P., Meyer, B. and Schmidt, T. (2012). Sharing Community Interpreting Corpora: A pilot study. In: Schmidt, T. and Wörner, K. (eds.) Multilingual Corpora and Multilingual Corpus Analysis. Amsterdam: Benjamins, 275-294.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,002 | 0,001 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,001 | 0,000 |
| Bibliométrie | 0,001 | 0,001 |
| Études des sciences et des technologies | 0,001 | 0,001 |
| Communication savante | 0,000 | 0,002 |
| Science ouverte | 0,004 | 0,005 |
| Intégrité de la recherche | 0,001 | 0,009 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,006 | 0,017 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; les deux têtes enseignantes s’accordent sur ce qui est montré ici.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».