MétaCan
Menu
Cohort builder

4,299,418 works, Canadian by any of four routes.

Every filter state is a URL; the URL is the query; the query is citable via /q/⟨hash⟩. The page, the API and the export parse the same parameters.

The current cohort, streamed from the database: every work column, the machine labels, the provisional scores, and the per-row validation status. Exports are capped at 100,000 rows. Mints a permanent /q/ link for this exact query. The same filters always produce the same link, whoever asks.

Search term
Author
Year range
Sort
Language
Type
Field
Venue
Topic
Software Engineering Research
Retraction
Abstract
Evidence source
Study design
Label agreement
Label status

Direct Codex and Gemma labels are unvalidated and sparse. Distilled predictions cover the full frame and are also unvalidated. Choose the evidence source explicitly; absence of a direct label is never a negative label.

affaffiliation
fundfunder
venuejournal
aboutaboutness

The four routes compose: require the funder route and exclude affiliation to get the funder-only stratum no affiliation-based frame ever sees.

3,468 results · 1 filter active ·
Results by year
20002025
Publication date
Categories
Machine labels · sparse coverage
Evidence
Language
Type
Citations
An unlabeled work is unknown, not a negative. Label coverage is reported on every query.
3,468 works in the cohort · of 4,299,418page 1 of 70

Labels cover 10 of 3,468 works in this cohort. The rest are unlabeled, which is not a negative label: the label table is sparse today and grows as labeling rounds land.

Distilled predictions cover 3,468 of 3,468 works in this cohort. Predictions are machine_predicted_unvalidated. The Gemma side is a direct model label for every work (title-only); the Codex side is a distilled, calibrated classifier. Candidate is the union; consensus is the intersection.

affunlabeled
Software: Practice and Experience
Frédéric Gervais, Benoît Fraikin
2006· paratext· en· Software Practice and Experience· Computer Science
machine prediction:candidate · noneconsensus · none
1,222
citations
afffundunlabeled
Who should fix this bug?
John Anvik, Lyndon Hiew, Gail C. Murphy
2006· article· en· Computer Science
machine prediction:candidate · noneconsensus · none
921
citations
affunlabeled
What makes a good bug report?
Nicolas Bettenburg, Sascha Just, Adrian Schröter, Cathrin Weiß, Rahul Premraj, Thomas Zimmermann
2008· article· en· Computer Science
machine prediction:candidate · noneconsensus · none
603
citations
affunlabeled
UMLDiff
Zhenchang Xing, Eleni Stroulia
2005· article· en· Computer Science
machine prediction:candidate · noneconsensus · none
369
citations
afffundno abstractunlabeled
A field study of API learning obstacles
Martin P. Robillard, Robert DeLine
2010· article· en· Empirical Software Engineering· Computer Science
machine prediction:candidate · noneconsensus · none
352
citations
affno abstractunlabeled
Curating GitHub for engineered software projects
Nuthan Munaiah, Steven Kroh, Craig Cabrey, Meiyappan Nagappan
2017· article· en· Empirical Software Engineering· Computer Science
machine prediction:candidate · noneconsensus · none
345
citations
affno abstractunlabeled
Latent Dirichlet Allocation
Joshua Charles Campbell, Abram Hindle, Eleni Stroulia
2015· book-chapter· en· Elsevier eBooks· Computer Science
machine prediction:candidate · noneconsensus · none
344
citations
affno abstractunlabeled
Do developers update their library dependencies?
Raula Gaikovina Kula, Daniel M. Germán, Ali Ouni, Takashi Ishio, Katsuro Inoue
2017· article· en· Empirical Software Engineering· Computer Science
machine prediction:candidate · metaresearchconsensus · none
335
citations
affno abstractunlabeled
Reporting Experiments in Software Engineering
Andreas Jedlitschka, Marcus Ciolkowski, Dietmar Pfahl
2007· book-chapter· en· Computer Science
machine prediction:candidate · metaresearchconsensus · none
330
citations
affunlabeled
A survey of software learnability
Tovi Grossman, George Fitzmaurice, Ramtin Attar
2009· article· en· Computer Science
machine prediction:candidate · noneconsensus · none
318
citations
affunlabeled
The promises and perils of mining git
Christian Bird, Peter C. Rigby, Earl T. Barr, David J. Hamilton, Daniel M. Germán, Prémkumar Dévanbu
2009· article· en· Computer Science
machine prediction:candidate · metaresearchconsensus · none
310
citations
affunlabeled
Mylar
Mik Kersten, Gail C. Murphy
2005· article· en· Computer Science
machine prediction:candidate · noneconsensus · none
296
citations
affunlabeled
On the naturalness of software
Abram Hindle, Earl T. Barr, Mark Gabel, Zhendong Su, Prémkumar Dévanbu
2016· article· en· Communications of the ACM· Computer Science
machine prediction:candidate · stsconsensus · none
291
citations
afffundunlabeled
Concern graphs
Martin P. Robillard, Gail C. Murphy
2002· article· en· Computer Science
machine prediction:candidate · noneconsensus · none
289
citations

How this was built: Screen · Findings · About