MétaCan
Menu
Cohort builder

4,299,418 works, Canadian by any of four routes.

Every filter state is a URL; the URL is the query; the query is citable via /q/⟨hash⟩. The page, the API and the export parse the same parameters.

The current cohort, streamed from the database: every work column, the machine labels, the provisional scores, and the per-row validation status. Exports are capped at 100,000 rows. Mints a permanent /q/ link for this exact query. The same filters always produce the same link, whoever asks.

Search term
Author
Year range
Sort
Language
Type
Field
Venue
Topic
Data Mining Algorithms and Applications
Retraction
Abstract
Evidence source
Study design
Label agreement
Label status

Direct Codex and Gemma labels are unvalidated and sparse. Distilled predictions cover the full frame and are also unvalidated. Choose the evidence source explicitly; absence of a direct label is never a negative label.

affaffiliation
fundfunder
venuejournal
aboutaboutness

The four routes compose: require the funder route and exclude affiliation to get the funder-only stratum no affiliation-based frame ever sees.

1,100 results · 1 filter active ·
Results by year
20002025
Publication date
Categories
Machine labels · sparse coverage
Evidence
Language
Type
Citations
An unlabeled work is unknown, not a negative. Label coverage is reported on every query.
1,100 works in the cohort · of 4,299,418page 1 of 22

Labels cover 1 of 1,100 works in this cohort. The rest are unlabeled, which is not a negative label: the label table is sparse today and grows as labeling rounds land.

Distilled predictions cover 1,100 of 1,100 works in this cohort. Predictions are machine_predicted_unvalidated. The Gemma side is a direct model label for every work (title-only); the Codex side is a distilled, calibrated classifier. Candidate is the union; consensus is the intersection.

affunlabeled
Interestingness measures for data mining
Liqiang Geng, Howard J. Hamilton
2006· review· en· ACM Computing Surveys· Computer Science
machine prediction:candidate · noneconsensus · none
1,125
citations
affno abstractunlabeled
The SPMF Open-Source Data Mining Library Version 2
Philippe Fournier‐Viger, Jerry Chun‐Wei Lin, Antonio Gomariz, Ted Gueniche, Azadeh Soltani, Zhi‐Hong Deng +1 more
2016· book-chapter· en· Lecture notes in computer science· Computer Science
machine prediction:candidate · noneconsensus · none
551
citations
affno abstractunlabeled
Mining Access Patterns Efficiently from Web Logs
Jian Pei, Jiawei Han, Behzad Mortazavi-Asl, Zhu Hua
2000· book-chapter· en· Lecture notes in computer science· Computer Science
machine prediction:candidate · noneconsensus · none
498
citations
affunlabeled
SPMF: a Java open-source pattern mining library
Philippe Fournier‐Viger, Antonio Gomariz, Ted Gueniche, Azadeh Soltani, Chengwei Wu, Vincent S. Tseng
2014· article· en· Journal of Machine Learning Research· Computer Science
machine prediction:candidate · noneconsensus · none
417
citations
affunlabeled
Privacy preserving frequent itemset mining
Stanley Robson de Medeiros Oliveira, Osmar R. Zai͏̈ane
2002· article· en· Computer Science
machine prediction:candidate · noneconsensus · none
286
citations
affunlabeled
Mining frequent patterns by pattern-growth
Jiawei Han, Jian Pei
2000· article· en· ACM SIGKDD Explorations Newsletter· Computer Science
machine prediction:candidate · noneconsensus · none
275
citations
affno abstractunlabeled
Clustering
Naomi Altman, Martin Krzywinski
2017· article· en· Nature Methods· Computer Science
machine prediction:candidate · noneconsensus · none
186
citations
afffundunlabeled
Document clustering with committees
Patrick Pantel, Dekang Lin
2002· article· en· Computer Science
machine prediction:candidate · noneconsensus · none
174
citations
affunlabeled
Beyond intratransaction association analysis
Hongjun Lü, Ling Feng, Jiawei Han
2000· article· en· ACM Transactions on Information Systems· Computer Science
machine prediction:candidate · noneconsensus · none
153
citations
affno abstractunlabeled
Top Down FP-Growth for Association Rule Mining
Ke Wang, Tang Liu, Jiawei Han, Junqiang Liu
2002· book-chapter· en· Lecture notes in computer science· Computer Science
machine prediction:candidate · noneconsensus · none
136
citations
affunlabeled
Constrained frequent pattern mining
Jian Pei, Jiawei Han
2002· article· en· ACM SIGKDD Explorations Newsletter· Computer Science
machine prediction:candidate · noneconsensus · none
133
citations
affunlabeled
Efficient algorithms for densest subgraph discovery
Yixiang Fang, Kaiqiang Yu, Reynold Cheng, Laks V. S. Lakshmanan, Xuemin Lin
2019· article· en· Proceedings of the VLDB Endowment· Computer Science
machine prediction:candidate · noneconsensus · none
115
citations

How this was built: Screen · Findings · About