MétaCan
Menu
Cohort builder

4,299,418 works, Canadian by any of four routes.

Every filter state is a URL; the URL is the query; the query is citable via /q/⟨hash⟩. The page, the API and the export parse the same parameters.

The current cohort, streamed from the database: every work column, the machine labels, the provisional scores, and the per-row validation status. Exports are capped at 100,000 rows. Mints a permanent /q/ link for this exact query. The same filters always produce the same link, whoever asks.

Search term
Author
Year range
Sort
Language
Type
Field
Venue
Topic
Data Quality and Management
Retraction
Abstract
Evidence source
Study design
Label agreement
Label status

Direct Codex and Gemma labels are unvalidated and sparse. Distilled predictions cover the full frame and are also unvalidated. Choose the evidence source explicitly; absence of a direct label is never a negative label.

affaffiliation
fundfunder
venuejournal
aboutaboutness

The four routes compose: require the funder route and exclude affiliation to get the funder-only stratum no affiliation-based frame ever sees.

867 results · 1 filter active ·
Results by year
20002025
Publication date
Categories
Machine labels · sparse coverage
Evidence
Language
Type
Citations
An unlabeled work is unknown, not a negative. Label coverage is reported on every query.
867 works in the cohort · of 4,299,418page 2 of 18

Labels cover 10 of 867 works in this cohort. The rest are unlabeled, which is not a negative label: the label table is sparse today and grows as labeling rounds land.

Distilled predictions cover 867 of 867 works in this cohort. Predictions are machine_predicted_unvalidated. The Gemma side is a direct model label for every work (title-only); the Codex side is a distilled, calibrated classifier. Candidate is the union; consensus is the intersection.

venueaboutno affunlabeled
Data Literacy - What is it and how can we make it happen?
Mark S. Frank, Johanna Walker, Judie Attard, Alan Freihof Tygel
2016· article· en· The Journal of Community Informatics· Decision Sciences
machine prediction:candidate · noneconsensus · none
64
citations
affunlabeled
The iBench integration metadata generator
Patricia C. Arocena, Boris Glavic, Radu Ciucanu, Renée J. Miller
2015· article· en· Proceedings of the VLDB Endowment· Decision Sciences
machine prediction:candidate · noneconsensus · none
59
citations
affunlabeled
Descriptive and prescriptive data cleaning
Anup Chalamalla, Ihab F. Ilyas, Mourad Ouzzani, Paolo Papotti
2014· article· en· Decision Sciences
machine prediction:candidate · metaresearchconsensus · none
53
citations
affunlabeled
KATARA
Xu Chu, John Morcos, Ihab F. Ilyas, Mourad Ouzzani, Paolo Papotti, Nan Tang +1 more
2015· article· en· Proceedings of the VLDB Endowment· Decision Sciences
machine prediction:candidate · noneconsensus · none
45
citations
affunlabeled
Ontology-based entity matching in attributed graphs
Hanchao Ma, Morteza Alipourlangouri, Yinghui Wu, Fei Chiang, Jiaxing Pi
2019· article· en· Proceedings of the VLDB Endowment· Decision Sciences
machine prediction:candidate · noneconsensus · none
42
citations
affunlabeled
Discovering linkage points over web data
Oktie Hassanzadeh, Ken Q. Pu, Soheil Hassas Yeganeh, Renée J. Miller, Lucian Popa, Mauricio A. Hernández +1 more
2013· article· en· Proceedings of the VLDB Endowment· Decision Sciences
machine prediction:candidate · noneconsensus · none
41
citations
affunlabeled
Qualitative data cleaning
Xu Chu, Ihab F. Ilyas
2016· article· en· Proceedings of the VLDB Endowment· Decision Sciences
machine prediction:candidate · metaresearchconsensus · none
40
citations
aboutno affunlabeled
Record linkage for pharmacoepidemiological studies in cancer patients
Myrthe P. P. van Herk‐Sukel, V.E.P.P. Lemmens, Lonneke V. van de Poll‐Franse, Ron M. C. Herings, J.W.W. Coebergh
2011· review· en· Pharmacoepidemiology and Drug Safety· Decision Sciences
machine prediction:candidate · metaresearchconsensus · none
39
citations
affunlabeled
Semandaq
Wenfei Fan, Floris Geerts, Xibei Jia
2008· article· en· Proceedings of the VLDB Endowment· Decision Sciences
machine prediction:candidate · noneconsensus · none
39
citations
affunlabeled
RONIN
Paul Ouellette, Aidan Sciortino, Fatemeh Nargesian, Bahar Ghadiri Bashardoost, Erkang Zhu, Ken Q. Pu +1 more
2021· article· en· Proceedings of the VLDB Endowment· Decision Sciences
machine prediction:candidate · noneconsensus · none
33
citations
affunlabeled
Data Quality
Shazia Sadiq, Tamraparni Dasu, Xin Luna Dong, Juliana Freire, Ihab F. Ilyas, Sebastian Link +4 more
2018· article· en· ACM SIGMOD Record· Decision Sciences
machine prediction:candidate · metaresearchconsensus · none
32
citations
affgemma · no categorygpt · no categorymodels split
Sampling for Big Data Profiling: A Survey
Zhicheng Liu, Aoqian Zhang
2020· article· en· IEEE Access· Decision Sciences
machine prediction:candidate · noneconsensus · none
32
citations
affno abstractunlabeled
Data Profiling
Ziawasch Abedjan, Lukasz Golab, Felix Naumann, Thorsten Papenbrock
2019· book· en· Synthesis lectures on data management· Decision Sciences
machine prediction:candidate · noneconsensus · none
31
citations
affunlabeled
SEMA-JOIN
Yeye He, Kris Ganjam, Xu Chu
2015· article· en· Proceedings of the VLDB Endowment· Decision Sciences
machine prediction:candidate · noneconsensus · none
31
citations

How this was built: Screen · Findings · About