MétaCan
Menu
Cohort builder

4,299,418 works, Canadian by any of four routes.

Every filter state is a URL; the URL is the query; the query is citable via /q/⟨hash⟩. The page, the API and the export parse the same parameters.

The current cohort, streamed from the database: every work column, the machine labels, the provisional scores, and the per-row validation status. Exports are capped at 100,000 rows. Mints a permanent /q/ link for this exact query. The same filters always produce the same link, whoever asks.

Search term
Author
Year range
Sort
Language
Type
Field
Venue
Topic
Text and Document Classification Technologies
Retraction
Abstract
Evidence source
Study design
Label agreement
Label status

Direct Codex and Gemma labels are unvalidated and sparse. Distilled predictions cover the full frame and are also unvalidated. Choose the evidence source explicitly; absence of a direct label is never a negative label.

affaffiliation
fundfunder
venuejournal
aboutaboutness

The four routes compose: require the funder route and exclude affiliation to get the funder-only stratum no affiliation-based frame ever sees.

374 results · 1 filter active ·
Results by year
20002025
Publication date
Categories
Machine labels · sparse coverage
Evidence
Language
Type
Citations
An unlabeled work is unknown, not a negative. Label coverage is reported on every query.
374 works in the cohort · of 4,299,418page 5 of 8

Labels cover 2 of 374 works in this cohort. The rest are unlabeled, which is not a negative label: the label table is sparse today and grows as labeling rounds land.

Distilled predictions cover 374 of 374 works in this cohort. Predictions are machine_predicted_unvalidated. The Gemma side is a direct model label for every work (title-only); the Codex side is a distilled, calibrated classifier. Candidate is the union; consensus is the intersection.

affunlabeled
How Well Do We Know Bernoulli
Giorgio Maria Di Nunzio, Alessandro Sordoni
2012· book-chapter· en· Research Padua Archive (University of Padua)· Computer Science
machine prediction:candidate · noneconsensus · none
3
citations
affunlabeled
Banking Order Classification and Information Extraction
Veli Oğuzalp Bakır, İlhan Cağatay, Melih Güven, Murat Koras, Mehmet Gönen, Barış Akgün
2022· article· en· 2022 30th Signal Processing and Communications Applications Conference (SIU)· Computer Science
machine prediction:candidate · noneconsensus · none
3
citations
aboutno affunlabeled
Study of Recommendation System
Shivani Parab, Saloni Pawar
2022· article· en· International Journal of Advanced Research in Science Communication and Technology· Computer Science
machine prediction:candidate · noneconsensus · none
2
citations
affunlabeled
Classification in Terms of Basic Concepts
Rick Szostak
2013· article· en· Advances in Classification Research Online· Computer Science
machine prediction:candidate · noneconsensus · none
2
citations
affunlabeled
Using self-supervised word segmentation in Chinese information retrieval
Fuchun Peng, Dale Schuurmans, Nick Cercone, Stephen Robertson
2002· article· en· Proceedings of the 25th annual international ACM SIGIR conference on Research and development in information retrieval - SIGIR '02· Computer Science
machine prediction:candidate · noneconsensus · none
2
citations
affno abstractunlabeled
Binary Text Representation for Feature Selection
Nguyen Lang, Ibrahim Zincir, A. Nur Zincir‐Heywood
2020· book-chapter· en· Advances in intelligent systems and computing· Computer Science
machine prediction:candidate · noneconsensus · none
2
citations
affno abstractunlabeled
Support Vector Machine for String Vectors
Malrey Lee, Taeho Jo
2006· book-chapter· en· Lecture notes in control and information sciences· Computer Science
machine prediction:candidate · noneconsensus · none
2
citations

How this was built: Screen · Findings · About