MétaCan
Menu
Cohort builder

4,299,418 works, Canadian by any of four routes.

Every filter state is a URL; the URL is the query; the query is citable via /q/⟨hash⟩. The page, the API and the export parse the same parameters.

The current cohort, streamed from the database: every work column, the machine labels, the provisional scores, and the per-row validation status. Exports are capped at 100,000 rows. Mints a permanent /q/ link for this exact query. The same filters always produce the same link, whoever asks.

Search term
Author
Year range
Sort
Language
Type
Field
Venue
Topic
Natural Language Processing Techniques
Retraction
Abstract
Evidence source
Study design
Label agreement
Label status

Direct Codex and Gemma labels are unvalidated and sparse. Distilled predictions cover the full frame and are also unvalidated. Choose the evidence source explicitly; absence of a direct label is never a negative label.

affaffiliation
fundfunder
venuejournal
aboutaboutness

The four routes compose: require the funder route and exclude affiliation to get the funder-only stratum no affiliation-based frame ever sees.

3,084 results · 1 filter active ·
Results by year
20002025
Publication date
Categories
Machine labels · sparse coverage
Evidence
Language
Type
Citations
An unlabeled work is unknown, not a negative. Label coverage is reported on every query.
3,084 works in the cohort · of 4,299,418page 2 of 62

Labels cover 10 of 3,084 works in this cohort. The rest are unlabeled, which is not a negative label: the label table is sparse today and grows as labeling rounds land.

Distilled predictions cover 3,084 of 3,084 works in this cohort. Predictions are machine_predicted_unvalidated. The Gemma side is a direct model label for every work (title-only); the Codex side is a distilled, calibrated classifier. Candidate is the union; consensus is the intersection.

affno abstractunlabeled
Multiple-Attribute Text Rewriting.
Guillaume Lample, Sandeep Subramanian, Eric M. Smith, Ludovic Denoyer, Marc’Aurelio Ranzato, Y-Lan Boureau
2018· article· en· International Conference on Learning Representations· Computer Science
machine prediction:candidate · noneconsensus · none
192
citations
afffundunlabeled
Concept discovery from text
Dekang Lin, Patrick Pantel
2002· article· en· Computer Science
machine prediction:candidate · noneconsensus · none
185
citations
affunlabeled
Classifying arguments by scheme
Vanessa Wei Feng, Graeme Hirst
2011· article· en· Computer Science
machine prediction:candidate · noneconsensus · none
178
citations
affno abstractunlabeled
Ontology and the Lexicon
Graeme Hirst
2009· book-chapter· en· Computer Science
machine prediction:candidate · noneconsensus · none
174
citations
fundvenueno affunlabeled
Statistical Phrase-Based Post-Editing
Michel Simard, Cyril Goutte, Pierre Isabelle
2007· article· en· NPARC· Computer Science
machine prediction:candidate · noneconsensus · none
168
citations
affunlabeled
Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets
Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh +45 more
2022· article· en· Transactions of the Association for Computational Linguistics· Computer Science
machine prediction:candidate · metaresearchconsensus · none
167
citations
afffundunlabeled
Learning to Understand Phrases by Embedding the Dictionary
Felix Hill, Kyunghyun Cho, Anna Korhonen, Yoshua Bengio
2016· article· en· Transactions of the Association for Computational Linguistics· Computer Science
machine prediction:candidate · noneconsensus · none
161
citations
aboutno affunlabeled
Toponym resolution in text
Jochen L. Leidner
2007· article· en· ACM SIGIR Forum· Computer Science
machine prediction:candidate · noneconsensus · none
158
citations
affunlabeled
The Mental Representation of Semitic Words
Jean-François Prunet, Renée Béland, Ali Idrissi
2000· article· en· Linguistic Inquiry· Computer Science
machine prediction:candidate · noneconsensus · none
157
citations
affunlabeled
Bootstrapping statistical parsers from small datasets
Mark Steedman, Miles Osborne, Anoop Sarkar, Stephen Clark, Rebecca Hwa, Julia Hockenmaier +3 more
2003· article· en· Computer Science
machine prediction:candidate · noneconsensus · none
156
citations
affunlabeled
Neural-Symbolic Learning and Reasoning: Contributions and Challenges
Artur d’Avila Garcez, Tarek R. Besold, Luc De Raedt, Péter Földiák, Pascal Hitzler, Thomas Icard +4 more
2015· article· en· City Research Online (City University London)· Computer Science
machine prediction:candidate · noneconsensus · none
127
citations
affunlabeled
Discovering word senses from text
Patrick Pantel, Dekang Lin
2002· article· en· Computer Science
machine prediction:candidate · noneconsensus · none
120
citations

How this was built: Screen · Findings · About