MétaCan
Menu
Cohort builder

4,299,418 works, Canadian by any of four routes.

Every filter state is a URL; the URL is the query; the query is citable via /q/⟨hash⟩. The page, the API and the export parse the same parameters.

The current cohort, streamed from the database: every work column, the machine labels, the provisional scores, and the per-row validation status. Exports are capped at 100,000 rows. Mints a permanent /q/ link for this exact query. The same filters always produce the same link, whoever asks.

Search term
Author
Year range
Sort
Language
Type
Field
Venue
Topic
Data Quality and Management
Retraction
Abstract
Evidence source
Study design
Label agreement
Label status

Direct Codex and Gemma labels are unvalidated and sparse. Distilled predictions cover the full frame and are also unvalidated. Choose the evidence source explicitly; absence of a direct label is never a negative label.

affaffiliation
fundfunder
venuejournal
aboutaboutness

The four routes compose: require the funder route and exclude affiliation to get the funder-only stratum no affiliation-based frame ever sees.

867 results · 1 filter active ·
Results by year
20002025
Publication date
Categories
Machine labels · sparse coverage
Evidence
Language
Type
Citations
An unlabeled work is unknown, not a negative. Label coverage is reported on every query.
867 works in the cohort · of 4,299,418page 10 of 18

Labels cover 10 of 867 works in this cohort. The rest are unlabeled, which is not a negative label: the label table is sparse today and grows as labeling rounds land.

Distilled predictions cover 867 of 867 works in this cohort. Predictions are machine_predicted_unvalidated teacher distillation outputs. Candidate is the union; consensus is the intersection.

affunlabeled
Time to Professionalise Data Stewardship
Maria Cruz, Melanie Imming, Hugh Shanahan, Marta Teperek, Celia van Gelder, Angus Whyte
2019· article· en· Figshare· Decision Sciences
distilled prediction:candidate · metaresearch+insufficient_payloadconsensus · insufficient_payload
1
citations
affunlabeled
Dataset for the Synaptotagmin-1 antibody screening study
2024· dataset· en· Zenodo (CERN European Organization for Nuclear Research)· Decision Sciences
distilled prediction:candidate · sts+scholarly_communication+open_science+insufficient_payloadconsensus · open_science+insufficient_payload
1
citations
affunlabeled
Record Fusion via Inference and Data Augmentation
Alireza Heidari, Ihab F. Ilyas, Theodoros Rekatsinas
2024· article· en· ACM / IMS Journal of Data Science· Decision Sciences
distilled prediction:candidate · metaresearch+scholarly_communication+open_scienceconsensus · scholarly_communication+open_science
1
citations
affaboutunlabeled
Leveraging Full Count Census Data through Record Linkage
Catherine A Fitch, Victoria Udalova, Luiza Antonie
2024· article· en· International Journal for Population Data Science· Decision Sciences
distilled prediction:candidate · scholarly_communication+open_scienceconsensus · scholarly_communication
1
citations
affunlabeled
In memoriam Alberto Oscar Mendelzon
Renée J. Miller
2005· article· en· ACM SIGMOD Record· Decision Sciences
distilled prediction:candidate · insufficient_payloadconsensus · insufficient_payload
1
citations
venueno affunlabeled
Leading in a Data Centric Society
Vincent Virk
2021· article· en· The Journal of Intelligence Conflict and Warfare· Decision Sciences
distilled prediction:candidate · noneconsensus · none
1
citations
affunlabeled
Editorial: Introducing this Special Issue on Data Librarianship
Kristi Thompson
2017· editorial· en· International Journal of Librarianship· Decision Sciences
distilled prediction:candidate · metaresearch+metaepi_narrow+scholarly_communication+open_science+research_integrity+insufficient_payloadconsensus · insufficient_payload
1
citations
aboutno affunlabeled
A pre and post data warehouse cleaning technique.
Timothy E. Ohanekwu
2002· article· en· Scholarship at UWindsor (University of Windsor)· Decision Sciences
distilled prediction:candidate · insufficient_payloadconsensus · none
1
citations
affno abstractunlabeled
Data Profiling Tools
Ziawasch Abedjan, Lukasz Golab, Felix Naumann, Thorsten Papenbrock
2019· book-chapter· en· Synthesis lectures on data management· Decision Sciences
distilled prediction:candidate · metaepi_narrow+scholarly_communication+open_science+insufficient_payloadconsensus · open_science+insufficient_payload
1
citations
affunlabeled
CERTEM
Tommaso Teofili, Donatella Firmani, Nick Koudas, Paolo Merialdo, Divesh Srivastava
2022· article· en· Proceedings of the VLDB Endowment· Decision Sciences
distilled prediction:candidate · noneconsensus · none
1
citations
aboutno affunlabeled
Data Management and Data Administration
Peter Aiken, Mark L. Gillenson, Xihui Zhang, David Rafner
2013· book-chapter· en· IGI Global eBooks· Decision Sciences
distilled prediction:candidate · metaepi_narrow+scholarly_communication+open_science+insufficient_payloadconsensus · open_science
1
citations
affunlabeled
Trusted Data in IBM's Master Data Management
Przemyslaw Pawluk, Jarek Gryz, Stephanie Hazlewood, Paul van Run
2011· article· en· Databases, Knowledge, and Data Applications· Decision Sciences
distilled prediction:candidate · metaepi_narrow+open_science+insufficient_payloadconsensus · open_science+insufficient_payload
1
citations
affvenueunlabeled
Measuring Data Re-Use In OpenAlex by Researchers, Institutions, and Countries
Geoff Krause, Timothy D. Bowman, Philippe Mongeon, Domenic Rosati, Michael Smit
2024· article· en· Proceedings of the Annual Conference of CAIS / Actes du congrès annuel de l ACSI· Decision Sciences
distilled prediction:candidate · metaresearch+scholarly_communicationconsensus · scholarly_communication
1
citations
affaboutunlabeled
Evaluating PPRL Vs Clear Text Linkage with Real-World Data
Michael Jarrett, Brent Hills, Yinshan Zhao, Adrian Brown, Sean Randall, James Boyd +2 more
2020· article· en· International Journal for Population Data Science· Decision Sciences
distilled prediction:candidate · metaresearch+scholarly_communication+open_scienceconsensus · none
1
citations
fundno affunlabeled
Data catalog tools: A systematic multivocal literature review
Marco Tonnarelli, Indika Kumara, Stefan Driessen, Damian A. Tamburri, Willem‐Jan van den Heuvel, Patrick Oor
2025· article· en· Journal of Systems and Software· Decision Sciences
distilled prediction:candidate · metaresearchconsensus · none
1
citations
fundno affunlabeled
RESTORE: Automated Regression Testing for Datasets
Lei Zhang, S.T. Howard, Tom Montpool, Jessica Moore, Krittika Mahajan, Andriy Miranskyy
2019· preprint· en· Maryland Shared Open Access Repository (USMAI Consortium)· Decision Sciences
distilled prediction:candidate · metaresearch+metaepi_narrow+scholarly_communication+open_scienceconsensus · open_science
1
citations
affunlabeled
Database Repairs and Analytic Tableaux
Leopoldo Bertossi, Camilla Schwind
2002· preprint· en· ArXiv.org· Decision Sciences
distilled prediction:candidate · metaepi_narrow+insufficient_payloadconsensus · insufficient_payload
1
citations
affno abstractunlabeled
Data quality in internet time, space, and communities (panel session)
Paul L. Bowen, James D. Funk, Matthias Jarke, Yang W. Lee, Yair Wand
2000· article· en· International Conference on Information Systems· Decision Sciences
distilled prediction:candidate · scholarly_communication+insufficient_payloadconsensus · insufficient_payload
1
citations

How this was built: Screen · Findings · About