MétaCan
Menu
Cohort builder

4,299,418 works, Canadian by any of four routes.

Every filter state is a URL; the URL is the query; the query is citable via /q/⟨hash⟩. The page, the API and the export parse the same parameters.

The current cohort, streamed from the database: every work column, the machine labels, the provisional scores, and the per-row validation status. Exports are capped at 100,000 rows. Mints a permanent /q/ link for this exact query. The same filters always produce the same link, whoever asks.

Search term
Author
Year range
Sort
Language
Type
Field
Venue
Topic
Data Quality and Management
Retraction
Abstract
Evidence source
Study design
Label agreement
Label status

Direct Codex and Gemma labels are unvalidated and sparse. Distilled predictions cover the full frame and are also unvalidated. Choose the evidence source explicitly; absence of a direct label is never a negative label.

affaffiliation
fundfunder
venuejournal
aboutaboutness

The four routes compose: require the funder route and exclude affiliation to get the funder-only stratum no affiliation-based frame ever sees.

867 results · 1 filter active ·
Results by year
20002025
Publication date
Categories
Machine labels · sparse coverage
Evidence
Language
Type
Citations
An unlabeled work is unknown, not a negative. Label coverage is reported on every query.
867 works in the cohort · of 4,299,418page 1 of 18

Labels cover 10 of 867 works in this cohort. The rest are unlabeled, which is not a negative label: the label table is sparse today and grows as labeling rounds land.

Distilled predictions cover 867 of 867 works in this cohort. Predictions are machine_predicted_unvalidated. The Gemma side is a direct model label for every work (title-only); the Codex side is a distilled, calibrated classifier. Candidate is the union; consensus is the intersection.

venueno affno abstractunlabeled
Systematic Approaches to a Successful Literature Review
Jackie MacDonald
2013· article· fr· Journal of the Canadian Health Libraries Association / Journal de l Association de bilbiothèques de la santé du Canada· Decision Sciences
machine prediction:candidate · metaresearchconsensus · metaresearch
871
citations
afffundunlabeled
HoloClean
Theodoros Rekatsinas, Xu Chu, Ihab F. Ilyas, Christopher Ré
2017· article· en· Proceedings of the VLDB Endowment· Decision Sciences
machine prediction:candidate · noneconsensus · none
454
citations
affgemma · metaresearchgpt · metaresearchmodels split
Challenges in administrative data linkage for research
Katie Harron, Chris Dibben, James Boyd, Anders Hjern, Mahmoud Azimaee, Maurício L. Barreto +1 more
2017· article· en· Big Data & Society· Decision Sciences
machine prediction:candidate · metaresearchconsensus · metaresearch
356
citations
affunlabeled
Improving data quality: consistency and accuracy
Gao Cong, Wenfei Fan, Floris Geerts, Xibei Jia, Shuai Ma
2007· article· en· Edinburgh Research Explorer (University of Edinburgh)· Decision Sciences
machine prediction:candidate · noneconsensus · none
312
citations
affno abstractunlabeled
Profiling relational data: a survey
Ziawasch Abedjan, Lukasz Golab, Felix Naumann
2015· article· en· The VLDB Journal· Decision Sciences
machine prediction:candidate · noneconsensus · none
273
citations
affunlabeled
Discovering data quality rules
Fei Chiang, Renée J. Miller
2008· article· en· Proceedings of the VLDB Endowment· Decision Sciences
machine prediction:candidate · noneconsensus · none
268
citations
affunlabeled
Discovering denial constraints
Xu Chu, Ihab F. Ilyas, Paolo Papotti
2013· article· en· Proceedings of the VLDB Endowment· Decision Sciences
machine prediction:candidate · noneconsensus · none
252
citations
affunlabeled
Detecting data errors
Ziawasch Abedjan, Xu Chu, Dong Deng, Raul Castro Fernandez, Ihab F. Ilyas, Mourad Ouzzani +3 more
2016· article· en· Proceedings of the VLDB Endowment· Decision Sciences
machine prediction:candidate · noneconsensus · none
237
citations
affunlabeled
Data lake management
Fatemeh Nargesian, Erkang Zhu, Renée J. Miller, Ken Q. Pu, Patricia C. Arocena
2019· article· en· Proceedings of the VLDB Endowment· Decision Sciences
machine prediction:candidate · noneconsensus · none
236
citations
afffundaboutunlabeled
Emerging problems of data quality in citizen science
Roman Lukyanenko, Jeffrey Parsons, Yolanda F. Wiersma
2016· editorial· en· Conservation Biology· Decision Sciences
machine prediction:candidate · metaresearchconsensus · metaresearch
209
citations
affunlabeled
Table union search on open data
Fatemeh Nargesian, Erkang Zhu, Ken Q. Pu, Renée J. Miller
2018· article· en· Proceedings of the VLDB Endowment· Decision Sciences
machine prediction:candidate · noneconsensus · none
193
citations
fundno affno abstractunlabeled
Foundations of Data Quality Management
Wenfei Fan, Floris Geerts
2012· article· en· Synthesis lectures on data management· Decision Sciences
machine prediction:candidate · noneconsensus · none
131
citations
affunlabeled
Extending dependencies with conditions
Loreto Bravo, Wenfei Fan, Shuai Ma
2007· article· en· Edinburgh Research Explorer (University of Edinburgh)· Decision Sciences
machine prediction:candidate · noneconsensus · none
108
citations
affunlabeled
Combining quantitative and logical data cleaning
Nataliya Prokoshyna, Jaroslaw Szlichta, Fei Chiang, Renée J. Miller, Divesh Srivastava
2015· article· en· Proceedings of the VLDB Endowment· Decision Sciences
machine prediction:candidate · noneconsensus · none
104
citations
fundno affgemma · metaresearchgpt · metaresearchmodels split
Evaluating bias due to data linkage error in electronic healthcare records
Katie Harron, Angie Wade, Ruth Gilbert, Berit Müller‐Pebody, Harvey Goldstein
2014· article· en· BMC Medical Research Methodology· Decision Sciences
machine prediction:candidate · metaresearchconsensus · metaresearch
102
citations
affunlabeled
Linked Movie Data Base
Oktie Hassanzadeh, Mariano P. Consens
2009· article· en· Decision Sciences
machine prediction:candidate · noneconsensus · none
93
citations
affunlabeled
Distributed data deduplication
Xu Chu, Ihab F. Ilyas, Paraschos Koutris
2016· article· en· Proceedings of the VLDB Endowment· Decision Sciences
machine prediction:candidate · noneconsensus · none
80
citations
affunlabeled
Messing up with BART
Patricia C. Arocena, Boris Glavic, Giansalvatore Mecca, Renée J. Miller, Paolo Papotti, Donatello Santoro
2015· article· en· Proceedings of the VLDB Endowment· Decision Sciences
machine prediction:candidate · noneconsensus · none
80
citations
affunlabeled
Meeting of the MINDS
Jamie Callan, James Allan, Charles L. A. Clarke, Susan Dumais, David A. Evans, Mark Sanderson +1 more
2007· article· en· ACM SIGIR Forum· Decision Sciences
machine prediction:candidate · noneconsensus · none
76
citations
affunlabeled
BlogScope
Nilesh Bansal, Nick Koudas
2007· article· en· Decision Sciences
machine prediction:candidate · noneconsensus · none
72
citations
affunlabeled
Hashed samples
Marios Hadjieleftheriou, Xiaohui Yu, Nick Koudas, Divesh Srivastava
2008· article· en· Proceedings of the VLDB Endowment· Decision Sciences
machine prediction:candidate · noneconsensus · none
68
citations
affno abstractgemma · no categorygpt · no categorymodels split
Data Ingestion for the Connected World.
John Meehan, Cansu Aslantas, Stan Zdonik, Nesime Tatbul, Jiang Du
2017· article· en· Conference on Innovative Data Systems Research· Decision Sciences
machine prediction:candidate · noneconsensus · none
68
citations
affunlabeled
Auto-join
Erkang Zhu, Yeye He, Surajit Chaudhuri
2017· article· en· Proceedings of the VLDB Endowment· Decision Sciences
machine prediction:candidate · noneconsensus · none
66
citations

How this was built: Screen · Findings · About