MétaCan
Menu
Cohort builder

4,299,418 works, Canadian by any of four routes.

Every filter state is a URL; the URL is the query; the query is citable via /q/⟨hash⟩. The page, the API and the export parse the same parameters.

The current cohort, streamed from the database: every work column, the machine labels, the provisional scores, and the per-row validation status. Exports are capped at 100,000 rows. Mints a permanent /q/ link for this exact query. The same filters always produce the same link, whoever asks.

Search term
Author
Year range
Sort
Language
Type
Field
Venue
Topic
Topic Modeling
Retraction
Abstract
Evidence source
Study design
Label agreement
Label status

Direct Codex and Gemma labels are unvalidated and sparse. Distilled predictions cover the full frame and are also unvalidated. Choose the evidence source explicitly; absence of a direct label is never a negative label.

affaffiliation
fundfunder
venuejournal
aboutaboutness

The four routes compose: require the funder route and exclude affiliation to get the funder-only stratum no affiliation-based frame ever sees.

2,769 results · 1 filter active ·
Results by year
20002025
Publication date
Categories
Machine labels · sparse coverage
Evidence
Language
Type
Citations
An unlabeled work is unknown, not a negative. Label coverage is reported on every query.
2,769 works in the cohort · of 4,299,418page 2 of 56

Labels cover 6 of 2,769 works in this cohort. The rest are unlabeled, which is not a negative label: the label table is sparse today and grows as labeling rounds land.

Distilled predictions cover 2,769 of 2,769 works in this cohort. Predictions are machine_predicted_unvalidated. The Gemma side is a direct model label for every work (title-only); the Codex side is a distilled, calibrated classifier. Candidate is the union; consensus is the intersection.

affunlabeled
MasakhaNER: Named Entity Recognition for African Languages
David Ifeoluwa Adelani, Jade Abbott, Graham Neubig, Daniel D’souza, Julia Kreutzer, Constantine Lignos +55 more
2021· article· en· Transactions of the Association for Computational Linguistics· Computer Science
machine prediction:candidate · noneconsensus · none
235
citations
affunlabeled
Information Retrieval by Semantic Similarity
Angelos Hliaoutakis, Giannis Varelas, Epimenidis Voutsakis, Euripides G. M. Petrakis, Evangelos Milios
2006· article· en· International Journal on Semantic Web and Information Systems· Computer Science
machine prediction:candidate · noneconsensus · none
228
citations
aboutno affunlabeled
A Deep Reinforcement Learning Chatbot
Iulian Vlad Serban, Chinnadhurai Sankar, Mathieu Germain, Saizheng Zhang, Zhouhan Lin, Sandeep Subramanian +12 more
2017· preprint· en· arXiv (Cornell University)· Computer Science
machine prediction:candidate · noneconsensus · none
200
citations
affunlabeled
VeriGen: A Large Language Model for Verilog Code Generation
Shailja Thakur, Baleegh Ahmad, Hammond Pearce, Benjamin Tan, Brendan Dolan-Gavitt, Ramesh Karri +1 more
2024· article· en· ACM Transactions on Design Automation of Electronic Systems· Computer Science
machine prediction:candidate · noneconsensus · none
191
citations
affunlabeled
A Neural Autoregressive Topic Model
Hugo Larochelle, Stanislas Lauly
2012· article· en· Computer Science
machine prediction:candidate · noneconsensus · none
189
citations
afffundunlabeled
Similarity of Semantic Relations
2006· article· en· Computational Linguistics· Computer Science
machine prediction:candidate · noneconsensus · none
179
citations
fundno affunlabeled
Generated Knowledge Prompting for Commonsense Reasoning
Jiacheng Liu, Alisa Liu, Ximing Lu, Sean Welleck, Peter West, Ronan Le Bras +2 more
2022· article· en· Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)· Computer Science
machine prediction:candidate · noneconsensus · none
176
citations
affunlabeled
Deep Learning for Text Style Transfer: A Survey
Di Jin, Zhijing Jin, Zhiting Hu, Olga Vechtomova, Rada Mihalcea
2021· article· en· Computational Linguistics· Computer Science
machine prediction:candidate · noneconsensus · none
162
citations
fundno affunlabeled
Symbolic Knowledge Distillation: from General Language Models to Commonsense Models
Peter West, Chandra Bhagavatula, Jack Hessel, Jena D. Hwang, Liwei Jiang, Ronan Le Bras +3 more
2022· article· en· Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies· Computer Science
machine prediction:candidate · noneconsensus · none
150
citations
affunlabeled
Truth and deception at the rhetorical structure level
Victoria L. Rubin, Tatiana Lukoianova
2014· article· en· Journal of the Association for Information Science and Technology· Computer Science
machine prediction:candidate · noneconsensus · none
144
citations

How this was built: Screen · Findings · About