MétaCan
Menu
Cohort builder

4,299,418 works, Canadian by any of four routes.

Every filter state is a URL; the URL is the query; the query is citable via /q/⟨hash⟩. The page, the API and the export parse the same parameters.

The current cohort, streamed from the database: every work column, the machine labels, the provisional scores, and the per-row validation status. Exports are capped at 100,000 rows. Mints a permanent /q/ link for this exact query. The same filters always produce the same link, whoever asks.

Search term
Author
Year range
Sort
Language
Type
Field
Venue
Topic
Web Data Mining and Analysis
Retraction
Abstract
Evidence source
Study design
Label agreement
Label status

Direct Codex and Gemma labels are unvalidated and sparse. Distilled predictions cover the full frame and are also unvalidated. Choose the evidence source explicitly; absence of a direct label is never a negative label.

affaffiliation
fundfunder
venuejournal
aboutaboutness

The four routes compose: require the funder route and exclude affiliation to get the funder-only stratum no affiliation-based frame ever sees.

630 results · 1 filter active ·
Results by year
20002025
Publication date
Categories
Machine labels · sparse coverage
Evidence
Language
Type
Citations
An unlabeled work is unknown, not a negative. Label coverage is reported on every query.
630 works in the cohort · of 4,299,418page 4 of 13

Labels cover 3 of 630 works in this cohort. The rest are unlabeled, which is not a negative label: the label table is sparse today and grows as labeling rounds land.

Distilled predictions cover 630 of 630 works in this cohort. Predictions are machine_predicted_unvalidated. The Gemma side is a direct model label for every work (title-only); the Codex side is a distilled, calibrated classifier. Candidate is the union; consensus is the intersection.

affunlabeled
Topic-oriented collaborative crawling
Chiasen Chung, Charles L. A. Clarke
2002· article· en· Computer Science
machine prediction:candidate · noneconsensus · none
9
citations
affunlabeled
Mapping the Internet
Arman Danesh, Ljiljana Trajković, S. H. RUBIN, Matt Smith
2002· article· en· Computer Science
machine prediction:candidate · noneconsensus · none
8
citations
affunlabeled
An Automated Approach to Identifying Corporate Editing
Veniamin Veselovsky, Dipto Sarkar, Jennings Anderson, Robert Soden
2022· article· en· Proceedings of the International AAAI Conference on Web and Social Media· Computer Science
machine prediction:candidate · noneconsensus · none
8
citations
affunlabeled
Capturing the Long Tail of Sensor Web
Steve Liang, James Badger, Rohana Rezel, Shawn Chen, Chih‐Yuan Huang, Ren-Yu Li
2010· article· en· Computer Science
machine prediction:candidate · noneconsensus · none
8
citations
affunlabeled
Extracting relational data from HTML repositories
Ruth Yuee Zhang, Laks V. S. Lakshmanan, Ruben H. Zamar
2004· article· en· ACM SIGKDD Explorations Newsletter· Computer Science
machine prediction:candidate · noneconsensus · none
8
citations
afffundno abstractunlabeled
Indexing Rich Internet Applications Using Components-Based Crawling
Ali Moosavi, Salman Hooshmand, Sara Baghbanzadeh, Guy-Vincent Jourdan, Gregor von Bochmann, Iosif Viorel Onut
2014· book-chapter· en· Lecture notes in computer science· Computer Science
machine prediction:candidate · noneconsensus · none
6
citations
affunlabeled
Standards opportunities around data-bearing Web pages
David R. Karger
2013· article· en· Philosophical Transactions of the Royal Society A Mathematical Physical and Engineering Sciences· Computer Science
machine prediction:candidate · noneconsensus · none
6
citations
affno abstractunlabeled
Wrapping HTML Tables into XML
Shijun Li, Mengchi Liu, Zhiyong Peng
2004· book-chapter· en· Lecture notes in computer science· Computer Science
machine prediction:candidate · noneconsensus · none
6
citations

How this was built: Screen · Findings · About