Les communautés francophones de l’Ouest canadien : de la constitution des corpus de français parlé aux perspectives de revitalisation
Bibliographic record
Abstract
La recherche sur les variétés de français parlées dans l’Ouest canadien a connu ces dernières années un essor lent mais certain. Il reste que bien des aspects de ces parlers français sont encore méconnus et qu’il est nécessaire de poursuivre les travaux de collecte, de description et d’analyse afin de donner un portrait linguistique actualisé des communautés francophones de ces régions du Canada. C’est le principal objectif que s’est fixé l’équipe de six chercheurs qui aborde, selon différents points de vue, la question des particularités des variétés de français de l’Ouest canadien et des communautés où elles sont en usage. En procédant d’est en ouest (Manitoba, Saskatchewan, Alberta), le tour d’horizon des recherches et des réalisations en cours proposé dans cet article se veut essentiellement descriptif et se situe donc, souvent mais pas uniquement, en amont de l’analyse. Il ne néglige pourtant pas les aspects réflexifs qui accompagnent nécessairement la prise en main d’un corpus oral déjà constitué ou l’élaboration d’un nouveau corpus. Les corpus présentés ici offrent de nombreuses perspectives d’analyse qui s’arriment à d’autres projets en cours de grande envergure (PFC,Le français à la mesure d’un continent) dont la finalité commune est l’enrichissement des connaissances sur les variétés de français de la francophonie.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.007 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.007 | 0.012 |
| Science and technology studies | 0.010 | 0.010 |
| Scholarly communication | 0.010 | 0.003 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.010 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".