MétaCan
Menu
Cohort builder

4,299,418 works, Canadian by any of four routes.

Every filter state is a URL; the URL is the query; the query is citable via /q/⟨hash⟩. The page, the API and the export parse the same parameters.

The current cohort, streamed from the database: every work column, the machine labels, the provisional scores, and the per-row validation status. Exports are capped at 100,000 rows. Mints a permanent /q/ link for this exact query. The same filters always produce the same link, whoever asks.

Search term
Author
Year range
Sort
Language
Type
Field
Venue
Topic
Natural Language Processing Techniques
Retraction
Abstract
Evidence source
Study design
Label agreement
Label status

Direct Codex and Gemma labels are unvalidated and sparse. Distilled predictions cover the full frame and are also unvalidated. Choose the evidence source explicitly; absence of a direct label is never a negative label.

affaffiliation
fundfunder
venuejournal
aboutaboutness

The four routes compose: require the funder route and exclude affiliation to get the funder-only stratum no affiliation-based frame ever sees.

3,084 results · 1 filter active ·
Results by year
20002025
Publication date
Categories
Machine labels · sparse coverage
Evidence
Language
Type
Citations
An unlabeled work is unknown, not a negative. Label coverage is reported on every query.
3,084 works in the cohort · of 4,299,418page 1 of 62

Labels cover 10 of 3,084 works in this cohort. The rest are unlabeled, which is not a negative label: the label table is sparse today and grows as labeling rounds land.

Distilled predictions cover 3,084 of 3,084 works in this cohort. Predictions are machine_predicted_unvalidated. The Gemma side is a direct model label for every work (title-only); the Codex side is a distilled, calibrated classifier. Candidate is the union; consensus is the intersection.

affunlabeled
Attention Is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez +2 more
2025· preprint· en· Computer Science
machine prediction:candidate · noneconsensus · none
6,569
citations
affunlabeled
Analyzing Linguistic Data
R. Harald Baayen
2008· book· en· Cambridge University Press eBooks· Computer Science
machine prediction:candidate · noneconsensus · none
3,194
citations
affunlabeled
A Neural Probabilistic Language Model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent
2000· article· en· Computer Science
machine prediction:candidate · noneconsensus · none
1,162
citations
affno abstractunlabeled
Association for Computational Linguistics
G. Hirst
2006· book-chapter· en· Elsevier eBooks· Computer Science
machine prediction:candidate · insufficient_payloadconsensus · insufficient_payload
1,154
citations
affunlabeled
The Winograd Schema Challenge
Hector J. Levesque
2011· article· en· Computer Science
machine prediction:candidate · noneconsensus · none
865
citations
afffundunlabeled
Discovering word senses from text
Patrick Pantel, Dekang Lin
2002· article· en· Computer Science
machine prediction:candidate · noneconsensus · none
597
citations
aboutno affunlabeled
From TreeBank to PropBank
Paul Kingsbury, Martha Palmer
2002· article· en· Computer Science
machine prediction:candidate · insufficient_payloadconsensus · none
539
citations
affunlabeled
SemEval-2010 task 8
Iris Hendrickx, Su Nam Kim, Zornitsa Kozareva, Preslav Nakov, Diarmuid Ó Séaghdha, Sebastian Padó +3 more
2009· article· en· Computer Science
machine prediction:candidate · noneconsensus · none
497
citations
affno abstractunlabeled
Indeterminate Pronouns: The View from Japanese
Angelika Kratzer, Junko Shimoyama
2017· book-chapter· en· Studies in natural language and linguistic theory· Computer Science
machine prediction:candidate · noneconsensus · none
443
citations
affunlabeled
Confidence estimation for machine translation
John Blatz, Erin Fitzgerald, George Foster, Simona Gandrabur, Cyril Goutte, Alex Kulesza +2 more
2004· article· en· Computer Science
machine prediction:candidate · noneconsensus · none
356
citations
affunlabeled
RWKV: Reinventing RNNs for the Transformer Era
Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Stella Biderman +26 more
2023· article· en· Computer Science
machine prediction:candidate · noneconsensus · none
303
citations
affunlabeled
Near-Synonymy and Lexical Choice
Philip Edmonds, Graeme Hirst
2002· article· en· Computational Linguistics· Computer Science
machine prediction:candidate · noneconsensus · none
258
citations
affunlabeled
Mixture-model adaptation for SMT
George Foster, Roland Kühn
2007· article· en· Computer Science
machine prediction:candidate · noneconsensus · none
253
citations
afffundunlabeled
Standardized annotation of translated open reading frames
Jonathan M. Mudge, Jorge Ruiz‐Orera, John R. Prensner, Marie A. Brunet, Ferriol Calvet, Irwin Jungreis +29 more
2022· letter· en· Nature Biotechnology· Computer Science
machine prediction:candidate · noneconsensus · none
250
citations

How this was built: Screen · Findings · About