MétaCan
Menu
Cohort builder

4,299,418 works, Canadian by any of four routes.

Every filter state is a URL; the URL is the query; the query is citable via /q/⟨hash⟩. The page, the API and the export parse the same parameters.

The current cohort, streamed from the database: every work column, the machine labels, the provisional scores, and the per-row validation status. Exports are capped at 100,000 rows. Mints a permanent /q/ link for this exact query. The same filters always produce the same link, whoever asks.

Search term
Author
Year range
Sort
Language
Type
Field
Venue
Topic
Natural Language Processing Techniques
Retraction
Abstract
Evidence source
Study design
Label agreement
Label status

Direct Codex and Gemma labels are unvalidated and sparse. Distilled predictions cover the full frame and are also unvalidated. Choose the evidence source explicitly; absence of a direct label is never a negative label.

affaffiliation
fundfunder
venuejournal
aboutaboutness

The four routes compose: require the funder route and exclude affiliation to get the funder-only stratum no affiliation-based frame ever sees.

3,084 results · 1 filter active ·
Results by year
20002025
Publication date
Categories
Machine labels · sparse coverage
Evidence
Language
Type
Citations
An unlabeled work is unknown, not a negative. Label coverage is reported on every query.
3,084 works in the cohort · of 4,299,418page 3 of 62

Labels cover 10 of 3,084 works in this cohort. The rest are unlabeled, which is not a negative label: the label table is sparse today and grows as labeling rounds land.

Distilled predictions cover 3,084 of 3,084 works in this cohort. Predictions are machine_predicted_unvalidated. The Gemma side is a direct model label for every work (title-only); the Codex side is a distilled, calibrated classifier. Candidate is the union; consensus is the intersection.

affno abstractunlabeled
Cross-Language Information Retrieval
Jian‐Yun Nie
2010· book· en· Synthesis lectures on human language technologies· Computer Science
machine prediction:candidate · noneconsensus · none
109
citations
affno abstractunlabeled
A Statistical Corpus-Based Term Extractor
Patrick Pantel, Dekang Lin
2001· book-chapter· en· Lecture notes in computer science· Computer Science
machine prediction:candidate · noneconsensus · none
107
citations
affunlabeled
Learning Noun Phrase Query Segmentation
Shane Bergsma, Qin Iris Wang
2007· article· en· Empirical Methods in Natural Language Processing· Computer Science
machine prediction:candidate · noneconsensus · none
106
citations
afffundunlabeled
Segmenting documents by stylistic character
N. L. Graham, Graeme Hirst, Bhaskara Marthi
2005· article· en· Natural Language Engineering· Computer Science
machine prediction:candidate · noneconsensus · none
104
citations
afffundunlabeled
Computing Lexical Contrast
Saif M. Mohammad, Bonnie J. Dorr, Graeme Hirst, Peter D. Turney
2012· article· en· Computational Linguistics· Computer Science
machine prediction:candidate · noneconsensus · none
100
citations
affunlabeled
AraT5: Text-to-Text Transformers for Arabic Language Generation
El Moatez Billah Nagoudi, AbdelRahim Elmadany, Muhammad Abdul-Mageed
2022· article· en· Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)· Computer Science
machine prediction:candidate · noneconsensus · none
97
citations
affvenueunlabeled
Paraphrasing for Style
Wei Xu, Alan Ritter, Bill Dolan, Ralph Grishman, Colin Cherry
2012· article· en· NPARC· Computer Science
machine prediction:candidate · noneconsensus · none
97
citations
aboutno affunlabeled
Intelligent Language Tutors
V. Melissa Holland, Michelle Sams, Jonathan D. Kaplan
2013· book· it· Computer Science
machine prediction:candidate · noneconsensus · none
94
citations
afffundunlabeled
Computing word-pair antonymy
Saif M. Mohammad, Bonnie J. Dorr, Graeme Hirst
2008· article· en· Computer Science
machine prediction:candidate · noneconsensus · none
93
citations
affunlabeled
Probabilistic Frame Induction
Jackie Chi Kit Cheung, Hoifung Poon, Lucy Vanderwende
2013· article· en· arXiv (Cornell University)· Computer Science
machine prediction:candidate · noneconsensus · none
87
citations
affunlabeled
Diagnosing Cyclicity in Sluicing
Calixto Agüero-Bautista
2007· article· en· Linguistic Inquiry· Computer Science
machine prediction:candidate · noneconsensus · none
85
citations
affunlabeled
French patterns for expressing concept relations
Elizabeth Marshman, T.J. Morgan, Ingrid Meyer
2002· article· en· Terminology International Journal of Theoretical and Applied Issues in Specialized Communication· Computer Science
machine prediction:candidate · noneconsensus · none
83
citations
affunlabeled
Names and similarities on the web
Marius Paşca, Dekang Lin, Jeffrey P. Bigham, Andrei Lifchits, Alpa Jain
2006· article· en· Computer Science
machine prediction:candidate · noneconsensus · none
83
citations
affunlabeled
Pulling their weight
Paul Cook, Afsaneh Fazly, Suzanne Stevenson
2007· article· en· Computer Science
machine prediction:candidate · noneconsensus · none
81
citations
affno abstractunlabeled
Topicalization in Asian Languages
Liejiong Xu
2006· other· en· Computer Science
machine prediction:candidate · noneconsensus · none
80
citations
affunlabeled
Techniques in Complex Semantic Fieldwork
M. Ryan Bochnak, Lisa Matthewson
2019· article· en· Annual Review of Linguistics· Computer Science
machine prediction:candidate · noneconsensus · none
79
citations

How this was built: Screen · Findings · About