MétaCan
Menu
Cohort builder

4,299,418 works, Canadian by any of four routes.

Every filter state is a URL; the URL is the query; the query is citable via /q/⟨hash⟩. The page, the API and the export parse the same parameters.

The current cohort, streamed from the database: every work column, the machine labels, the provisional scores, and the per-row validation status. Exports are capped at 100,000 rows. Mints a permanent /q/ link for this exact query. The same filters always produce the same link, whoever asks.

Search term
Author
Year range
Sort
Language
Type
Field
Venue
Topic
Natural Language Processing Techniques
Retraction
Abstract
Evidence source
Study design
Label agreement
Label status

Direct Codex and Gemma labels are unvalidated and sparse. Distilled predictions cover the full frame and are also unvalidated. Choose the evidence source explicitly; absence of a direct label is never a negative label.

affaffiliation
fundfunder
venuejournal
aboutaboutness

The four routes compose: require the funder route and exclude affiliation to get the funder-only stratum no affiliation-based frame ever sees.

3,084 results · 1 filter active ·
Results by year
20002025
Publication date
Categories
Machine labels · sparse coverage
Evidence
Language
Type
Citations
An unlabeled work is unknown, not a negative. Label coverage is reported on every query.
3,084 works in the cohort · of 4,299,418page 4 of 62

Labels cover 10 of 3,084 works in this cohort. The rest are unlabeled, which is not a negative label: the label table is sparse today and grows as labeling rounds land.

Distilled predictions cover 3,084 of 3,084 works in this cohort. Predictions are machine_predicted_unvalidated. The Gemma side is a direct model label for every work (title-only); the Codex side is a distilled, calibrated classifier. Candidate is the union; consensus is the intersection.

affunlabeled
Techniques in Complex Semantic Fieldwork
M. Ryan Bochnak, Lisa Matthewson
2019· article· en· Annual Review of Linguistics· Computer Science
machine prediction:candidate · noneconsensus · none
79
citations
afffundunlabeled
TransType
Philippe Langlais, George Foster, Guy Lapalme
2000· article· en· Computer Science
machine prediction:candidate · noneconsensus · none
79
citations
affunlabeled
Measuring Machine Translation Errors in New Domains
Ann Irvine, John Morgan, Marine Carpuat, Hal Daumé, Dragos Stefan Munteanu
2013· article· en· Transactions of the Association for Computational Linguistics· Computer Science
machine prediction:candidate · noneconsensus · none
77
citations
affunlabeled
Translating with non-contiguous phrases
Michel Simard, Nicola Cancedda, Bruno Cavestro, Marc Dymetman, Éric Gaussier, Cyril Goutte +3 more
2005· article· en· Computer Science
machine prediction:candidate · noneconsensus · none
74
citations
affno abstractunlabeled
Analyzing linguistic complexity and scientific impact
Chao Lu, Yi Bu, Xianlei Dong, Jie Wang, Ying Ding, Vincent Larivière +3 more
2019· article· en· Journal of Informetrics· Computer Science
machine prediction:candidate · bibliometricsconsensus · none
74
citations
affunlabeled
New Tools for Web-Scale N-grams
Dekang Lin, Kenneth Church, Heng Ji, Satoshi Sekine, David Yarowsky, Shane Bergsma +6 more
2010· article· en· Computer Science
machine prediction:candidate · noneconsensus · none
74
citations
affunlabeled
Lexical Cohesion Analysis of Political Speech
Beata Beigman Klebanov, Daniel Diermeier, Eyal Beigman
2008· article· en· Political Analysis· Computer Science
machine prediction:candidate · noneconsensus · none
70
citations
affunlabeled
Semantic Role Labeling
2011· article· en· Institutional Research Information System (Università degli Studi di Trento)· Computer Science
machine prediction:candidate · noneconsensus · none
70
citations
afffundunlabeled
Argumentation Quality Assessment: Theory vs. Practice
Henning Wachsmuth, Nona Naderi, Ivan Habernal, Yufang Hou, Graeme Hirst, Iryna Gurevych +1 more
2017· article· en· Computer Science
machine prediction:candidate · noneconsensus · none
65
citations
affno abstractunlabeled
Classification of semantic relations between nominals
Roxana Gîrju, Preslav Nakov, Vivi Năstase, Stan Śzpakowicz, Peter D. Turney, Deniz Yüret
2009· article· en· Language Resources and Evaluation· Computer Science
machine prediction:candidate · noneconsensus · none
63
citations
affvenueunlabeled
Has Computerization Changed Translation?
Brian Mossop
2006· article· en· Meta Journal des traducteurs· Computer Science
machine prediction:candidate · noneconsensus · none
63
citations
affunlabeled
Translating Unknown Words by Analogical Learning
Philippe Langlais, Alexandre Patry
2007· article· en· Empirical Methods in Natural Language Processing· Computer Science
machine prediction:candidate · noneconsensus · none
63
citations
affunlabeled
Lethal Ambiguity
Martha McGinnis
2004· article· en· Linguistic Inquiry· Computer Science
machine prediction:candidate · noneconsensus · none
62
citations
affunlabeled
Substring-Based Transliteration
Tarek Sherif, Grzegorz Kondrak
2007· article· en· Meeting of the Association for Computational Linguistics· Computer Science
machine prediction:candidate · noneconsensus · none
61
citations
afffundunlabeled
LADEC: The Large Database of English Compounds
Christina L. Gagné, Thomas L. Spalding, Daniel Schmidtke
2019· article· en· Behavior Research Methods· Computer Science
machine prediction:candidate · noneconsensus · none
60
citations
affno abstractunlabeled
Evaluation of parallel text alignment systems
Jean Véronis, Philippe Langlais
2000· book-chapter· en· Text, speech and language technology· Computer Science
machine prediction:candidate · noneconsensus · none
59
citations
affunlabeled
Understanding the Origins of Bias in Word Embeddings
Marc-Etienne Brunet, Colleen Alkalay-Houlihan, Ashton Anderson, Richard S. Zemel
2019· article· en· International Conference on Machine Learning· Computer Science
machine prediction:candidate · noneconsensus · none
58
citations

How this was built: Screen · Findings · About