MétaCan
Menu
Cohort builder

4,299,418 works, Canadian by any of four routes.

Every filter state is a URL; the URL is the query; the query is citable via /q/⟨hash⟩. The page, the API and the export parse the same parameters.

The current cohort, streamed from the database: every work column, the machine labels, the provisional scores, and the per-row validation status. Exports are capped at 100,000 rows. Mints a permanent /q/ link for this exact query. The same filters always produce the same link, whoever asks.

Search term
Author
Year range
Sort
Language
Type
Field
Venue
Topic
Natural Language Processing Techniques
Retraction
Abstract
Evidence source
Study design
Label agreement
Label status

Direct Codex and Gemma labels are unvalidated and sparse. Distilled predictions cover the full frame and are also unvalidated. Choose the evidence source explicitly; absence of a direct label is never a negative label.

affaffiliation
fundfunder
venuejournal
aboutaboutness

The four routes compose: require the funder route and exclude affiliation to get the funder-only stratum no affiliation-based frame ever sees.

3,084 results · 1 filter active ·
Results by year
20002025
Publication date
Categories
Machine labels · sparse coverage
Evidence
Language
Type
Citations
An unlabeled work is unknown, not a negative. Label coverage is reported on every query.
3,084 works in the cohort · of 4,299,418page 8 of 62

Labels cover 10 of 3,084 works in this cohort. The rest are unlabeled, which is not a negative label: the label table is sparse today and grows as labeling rounds land.

Distilled predictions cover 3,084 of 3,084 works in this cohort. Predictions are machine_predicted_unvalidated teacher distillation outputs. Candidate is the union; consensus is the intersection.

affunlabeled
MSR SPLAT, a language analysis toolkit
Chris Quirk, Pallavi Choudhury, Jianfeng Gao, Hisami Suzuki, Kristina Toutanova, Michael Gamon +3 more
2012· article· en· Computer Science
distilled prediction:candidate · noneconsensus · none
32
citations
venueno affunlabeled
The Unit of Translation: Statistics Speak
Harry J. Huang, Canzhong Wu
2009· article· en· Meta Journal des traducteurs· Computer Science
distilled prediction:candidate · noneconsensus · none
32
citations
affunlabeled
Data Sets: Word Embeddings Learned from Tweets and General Data
Quanzhi Li, Sameena Shah, Xiaomo Liu, Armineh Nourbakhsh
2017· article· en· Proceedings of the International AAAI Conference on Web and Social Media· Computer Science
distilled prediction:candidate · open_scienceconsensus · none
32
citations
affno abstractunlabeled
Conceptualizing Student Models for ICALL
Luiz Amaral, Detmar Meurers
2007· book-chapter· en· Lecture notes in computer science· Computer Science
distilled prediction:candidate · metaepi_narrowconsensus · none
32
citations
affunlabeled
plWordNet 3.0 - a Comprehensive Lexical-Semantic Resource.
Marek Maziarz, Maciej Piasecki, Ewa Rudnicka, Stan Śzpakowicz, Paweł Kędzia
2016· article· en· International Conference on Computational Linguistics· Computer Science
distilled prediction:candidate · noneconsensus · none
32
citations
affunlabeled
Adapting a synonym database to specific domains
Davide Turcato, Fred Popowich, Janine Toole, Dan Fass, Devlan Nicholson, Gordon Tisher
2000· article· en· Computer Science
distilled prediction:candidate · noneconsensus · none
31
citations
affno abstractunlabeled
Introduction to special issue on post-editing
Sharon O’Brien, Michel Simard
2014· article· en· Machine Translation· Computer Science
distilled prediction:candidate · noneconsensus · none
31
citations
aboutno affunlabeled
Eskimo-Aleut
Anna Berge
2016· reference-entry· en· Oxford Research Encyclopedia of Linguistics· Computer Science
distilled prediction:candidate · metaresearch+metaepi_narrow+open_science+research_integrityconsensus · none
31
citations
affno abstractunlabeled
Question Answering By Passage Selection
Charles L. A. Clarke, Gordon V. Cormack, Thomas R. Lynam, Egidio L. Terra
2006· book-chapter· en· Text, speech and language technology· Computer Science
distilled prediction:candidate · metaepi_narrowconsensus · none
31
citations
venueno affunlabeled
The DLSIUAES Team's Participation in the TAC 2008 Tracks
Alexandra Balahur, Elena Lloret, Óscar Ferrández, Andrés Montoyo, Manuel Palomar, Rafael Muñoz
2008· article· en· Theory and applications of categories· Computer Science
distilled prediction:candidate · noneconsensus · none
31
citations
aboutno affunlabeled
Text Deixis in Narrative Sequences
Josep E. Ribera
2007· article· en· International journal of english studies, Vol· Computer Science
distilled prediction:candidate · noneconsensus · none
31
citations
affunlabeled
Expression of uncertainty in linguistic data
Alain Auger, J. Roy
2008· article· en· International Conference on Information Fusion· Computer Science
distilled prediction:candidate · noneconsensus · none
30
citations
affunlabeled
Confidence estimation for NLP applications
Simona Gandrabur, George Foster, Guy Lapalme
2006· article· en· ACM Transactions on Speech and Language Processing· Computer Science
distilled prediction:candidate · noneconsensus · none
30
citations
venueno affunlabeled
Lexical Cohesion and Translation Equivalence
Kazem Lotfipour-Saedi
2002· article· en· Meta Journal des traducteurs· Computer Science
distilled prediction:candidate · noneconsensus · none
29
citations
affno abstractunlabeled
Hyperrelations in version space
Hui Wang, Ivo Düntsch, Günther Gediga, Andrzej Skowron
2003· article· en· International Journal of Approximate Reasoning· Computer Science
distilled prediction:candidate · noneconsensus · none
29
citations
affno abstractunlabeled
Machine Translation of Closed Captions
Fred Popowich, Paul McFetridge, Davide Turcato, Janine Toole
2000· article· en· Machine Translation· Computer Science
distilled prediction:candidate · noneconsensus · none
28
citations

How this was built: Screen · Findings · About