Notice bibliographique
Résumé
<strong>This dataset is distributed under a Creative Commons Attribution Non Commercial 4.0 International license. Use for research purposes only!</strong> The dataset contains 13,171 variable double-object and prepositional datives extracted from the International Corpus of English series and the Corpus of Global web-based English sampling from nine national varieties of English: British English Canadian English New Zealand English Irish English Hong Kong English Philippine English Singapore English Indian English Jamaican English <strong>The dataframe contains the following columns:</strong> <strong>1 TokenID</strong>: Unique identifier for the individual token <strong>2 Variety</strong>: The variety from which the token is taken <strong>3 Nativity</strong>: Native or non-native variety of English (L1 vs. L2) <strong>4 Corpus</strong>: The corpus from which the token stems <strong>5 Subcorpus</strong>: Combination of Corpus and Variety <strong>6 FileID</strong>: ID of the corpus file in which the token was found. Format: VARIETY:FILENAME <strong>7 TextID</strong>: ID of the corpus text in which the token was found. Individual files in ICE can have multiple texts. Format: VARIETY:FILENAME:TEXTNUMBER <strong>8 LineID</strong>: ID of the line in the text in which the token sentence was found. Format: VARIETY:FILENAME:TEXTNUMBER:LINENUMBER <strong>9 SpeakerID</strong>: ID of the speaker of the sentence. Speakers in spoken texts are indicated with capital letters. Authors of written texts have ID ‘A’. Format: VARIETY:FILENAME:TEXTNUMBER:SPEAKERID <strong>10 UnitMarker</strong>: UnitMarker of the utterance in the text. Format UTTERANCE NUMBER:TEXTNUMBER:SPEAKERID <strong>11 GenreFine</strong>: 14-level distinction: The 12-level ICE sub-register in which the token was found and the two levels in GloWbE (blog vs. general). Levels: See ICE documentation <strong>12 GenreCoarse</strong>: 5-level distinction: The 4-level ICE register in which the token was found and GloWbE (online = 1 level). Levels: See ICE documentation <strong>13 Mode</strong>: The mode (‘spoken’, ‘written’) of the token. <strong>14 Register</strong>: The 4-level Register along two axes – spoken vs. written / informal vs. formal <strong>15 PriorContextPlain</strong>: The plain text version of the 100 words preceding the dative token. <strong>16 PriorContextTag</strong>: The POS-tagged version of the 100 words preceding the dative token. <strong>17 SentencePlain</strong>: The plain text version of the sentence containing the dative token. <strong>18 SentenceTag</strong>: The POS-tagged version of the sentence containing the dative token. <strong>19 WholeConstructionPlain</strong>: The plain text version of the VP containing the dative token (i.e. verb + object + object). <strong>20 WholeConstructionTag</strong>: The POS-tagged version of the VP containing the dative token. <strong>21 Verb</strong>: The lemma of the verbal head (<em>give </em>in <em>gave it some thought</em>) <strong>22 VerbForm</strong>: The verb form of the verbal head (<em>gave </em>in <em>gave it some thought</em>) <strong>23 RecipientShort</strong>: The short plain text version of the recipient without hesitations or repetitions <strong>24 ThemeShort</strong>: The short plain text version of the theme without hesitations or repetitions <strong>25 RecipientLong</strong>: The long plain text version of the recipient with hesitations or repetitions <strong>26 ThemeLong</strong>: The long plain text version of the theme with hesitations or repetitions <strong>27 RecHeadPlain</strong>: The plain text version of the recipient head <strong>28 RecHeadTag</strong>: The POS-tagged version of the recipient head <strong>29 RecHeadLemma</strong>: The lemma of the recipient head <strong>30 ThemeHeadPlain</strong>: The plain text version of the theme head <strong>31 ThemeHeadTag</strong>: The POS-tagged version of the theme head <strong>32 ThemeHeadLemma</strong>: The lemma of the theme head <strong>33 VerbThemeLemma</strong>: Combination of the verb lemma and the theme head. Format: VERB_THEME <strong>34 VerbSense</strong>: Semantics of the verb based on the whole construction combined with the verb lemma. Format: VERB.VERBSEMANTICS <strong>35 VerbSemantics</strong>: 5-level distinction of verb semantics (‘a’, ‘t’, ‘p’, ‘f’, ‘c’). <strong>36 Resp</strong>: The variant order. Levels: ‘do’ (=ditransitive), ‘pd’ (=prepositional) <strong>37 RecAnimacy</strong>: 6-level distinction of recipient animacy following previous research: human (a1) > animal (a2) > collective (c) > locative (l) > temporal (t) > inanimate (i) <strong>38 ThemeAnimacy</strong>: 6-level distinction of theme animacy following previous research: human (a1) > animal (a2) > collective (c) > locative (l) > temporal (t) > inanimate (i) <strong>39 RecWordLth</strong>: Length of recipient NP in words <strong>40 RecLetterLth</strong>: Length of recipient NP in orthographic characters <strong>41 ThemeWordLth</strong>: Length of theme NP in words <strong>42 ThemeLetterLth</strong>: Length of theme NP in orthographic characters <strong>43 RecComplexity </strong> 15-level distinction of recipient complexity indicating type and number of posthead dependents, restricted to the ICE components. (GloWbE components make simplified distinction between ‘simple’ and ‘complex’). Levels: ‘s’ = simple (no postmodifications), ‘co’ = coordinated, ‘ge’ = general extender, ‘gn’ = genitive, ‘postad’ = postmodifying adverbial/adjective, ‘pp’ = modifying prepositional phrase, ‘appnom’ = nominal apposition, ‘rc’ = relative clause, ‘cp’ = complement clause, ‘advc’ = adverbial clause, ‘nonfin’ = nonfinite clause, ‘tpp’ = two nominal posthead dependents, ‘tvp’ = two posthead dependents involving at least one VP, ‘mpp’ = more than two nominal posthead dependents, ‘mvp’ = more than two posthead dependents involving at least one VP <strong>44 ThemeComplexity </strong> 15-level distinction of theme complexity indicating type and number of posthead dependents, restricted to the ICE components. (GloWbE components make simplified distinction between ‘simple’ and ‘complex’). Levels: ‘s’ = simple (no postmodifications), ‘co’ = coordinated, ‘ge’ = general extender, ‘gn’ = genitive, ‘postad’ = postmodifying adverbial/adjective, ‘pp’ = modifying prepositional phrase, ‘appnom’ = nominal apposition, ‘rc’ = relative clause, ‘cp’ = complement clause, ‘advc’ = adverbial clause, ‘nonfin’ = nonfinite clause, ‘tpp’ = two nominal posthead dependents, ‘tvp’ = two posthead dependents involving at least one VP, ‘mpp’ = more than two nominal posthead dependents, ‘mvp’ = more than two posthead dependents involving at least one VP <strong>45 RecNPExprType</strong>: Syntactic category of the recipient NP Levels: ‘dem’ = bare demonstrative; ‘nc’ = common noun; ‘np’ = proper noun; ‘pprn’ = personal pronoun; ‘iprn’ = impersonal pronoun; ‘rprn’ = reflexive pronoun; ‘vp’ = gerund (-ing) NP; ‘wh’ = NP headed by wh- word <strong>46 ThemeNPExprType</strong>: Syntactic category of the theme NP Levels: ‘dem’ = bare demonstrative; ‘nc’ = common noun; ‘np’ = proper noun; ‘pprn’ = personal pronoun; ‘iprn’ = impersonal pronoun; ‘rprn’ = reflexive pronoun; ‘vp’ = gerund (-ing) NP; ‘wh’ = NP headed by wh- word <strong>47 RecGivenness</strong>: Givenness of the recipient NP. Levels: ‘given’, ‘new’ <strong>48 ThemeGivenness</strong>: Givenness of the theme NP. Levels: ‘given’, ‘new’ <strong>49 RecDefiniteness</strong>: Definiteness of the recipient NP. Levels: ‘def’, ‘indef’ <strong>50 ThemeDefiniteness</strong>: Definiteness of the theme NP. Levels: ‘def’, ‘indef’ <strong>51 RecBinComplexity</strong>: Binary predictor of recipient complexity indicating following postmodifications after the head noun. Levels: ‘simple’, ‘complex’ <strong>52 ThemeBinComplexity</strong>: Binary predictor of theme complexity indicating following postmodifications after the head noun. Levels: ‘simple’, ‘complex’ <strong>53 RecPerson</strong>: Person of recipient. Levels: ‘local’, ‘non-local’ <strong>54 ThemeConcreteness</strong>: Concreteness of theme based on verb semantics. Levels: ‘concrete’, ‘non-concrete’ <strong>55 TypeTokenRatio</strong>: Type-token ratio of the 100 word context surrounding the token <strong>56 RecHeadFreq</strong>: Frequency of recipient head lemma in GloWbE <strong>57 ThemeHeadFreq</strong>: Frequency of theme head lemma in GloWbE <strong>58 RecThematicity</strong>: Normalized frequency of recipient head lemma in its text (per 2000 words) <strong>59 ThemeThematicity</strong>: Normalized frequency of theme head lemma in its text (per 2000 words) <strong>60 PrimeType</strong>: The response type of the preceding dative token, if any. Levels: ‘do, ‘pd, ‘NA’ <strong>61 Persistence</strong>: Indicates whether preceding dative token, if any, is the same or not. Levels: ‘none’, ‘yes’, ‘no’ <strong>62 SameUtterance</strong>: Indicates whether the preceding dative token occurred in the same utterance or not. Necessary for manual coding of persistence. <strong>63 DistanceToPrevious</strong>: Number of utterances between current and preceding dative token. ‘None’ if no preceding dative token. <strong>64 RecPron</strong>: Binary factor of recipient pronominality. Levels: ‘pron’, ‘non-pron’ <strong>65 ThemePron</strong>: Binary factor of theme pronominality: Levels: ‘pron’, ‘non-pron’ <strong>66 RecBinAnimacy</strong>: Binary factor of recipient animacy. Levels of RecAnimacy conflated to: ‘animate’, ‘inanimate’ <strong>67 ThemeBinAnimacy</strong>: Binary factor of theme animacy. Levels of ThemeAnimacy conflated to: ‘animate’, ‘inanimate’ <strong>68 logRecLetterLth</strong>: Natural logarithm of recipient length in orthographic characters <strong>69 logThemeLetterLth</strong>: Natural logarithm of theme length in orthographic characters <strong>70 WeightRatio</strong>: Ratio of object lengths: Recipient length in characters divided by theme length in characters <strong>71 logWeightRatio</
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,009 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,004 | 0,001 |
| Communication savante | 0,002 | 0,000 |
| Science ouverte | 0,002 | 0,002 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,043 | 0,006 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; les deux têtes enseignantes s’accordent sur ce qui est montré ici.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».