Bibliographic record
Abstract
The Dialect Topography of Canada has reached a kind of plateau.After ten years of data-gathering, from 1992 to 2002, we have assembled large databases on language variants in regions across Canada.The databases are accessible at dialect.topography.chass.utoronto.ca.The website, constructed by Dr. Tony Pi, is free of charge and user-friendly, with tutorials and analytic aids.We are not presently engaged in Dialect Topography surveys in other regions.In years to come, there will undoubtedly be more regional surveys and new surveys of the original regions, but the time gap between the existing ones and the ones that will follow entails that they will relate to one another not as additional contemporaneous surveys but as real-time comparisons.In this article, I illustrate the breadth of coverage by investigating three geolinguistic patterns that have emerged from our research.I begin with a brief introduction to the methods and goals of Dialect Topography.In so doing, I cannot avoid noting a salubrious coincidence.The first public presentation on Dialect Topography took place at Universite de Moncton, at a meeting of the Atlantic Provinces Linguistic Association in 1992.The presentation on which this article is based, which represents a kind of stock-taking on what we have accomplished with Dialect Topography at this juncture, also took place at Universite de Moncton.That first presentation, fourteen years ago, resulted in an article that provided an introduction to Dialect Topography (Chambers 1994).That article is fuller and more discursive than space allows here, and I am pleased to refer readers to it to fill in any gaps I leave here.The "distance" between that first presentation and this one symbolically represents a huge investment of time and effort by a team of dedicated scholars.]Our bond comes not only from the many hours we spent working together but also in the shared belief that we have left behind a resource that has almost limitless potential.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".