MétaCan
Menu
Cohort builder

4,299,418 works, Canadian by any of four routes.

Every filter state is a URL; the URL is the query; the query is citable via /q/⟨hash⟩. The page, the API and the export parse the same parameters.

The current cohort, streamed from the database: every work column, the machine labels, the provisional scores, and the per-row validation status. Exports are capped at 100,000 rows. Mints a permanent /q/ link for this exact query. The same filters always produce the same link, whoever asks.

Search term
Author
Year range
Sort
Language
Type
Field
Venue
Empirical Software Engineering
Topic
Retraction
Abstract
Evidence source
Study design
Label agreement
Label status

Direct Codex and Gemma labels are unvalidated and sparse. Distilled predictions cover the full frame and are also unvalidated. Choose the evidence source explicitly; absence of a direct label is never a negative label.

affaffiliation
fundfunder
venuejournal
aboutaboutness

The four routes compose: require the funder route and exclude affiliation to get the funder-only stratum no affiliation-based frame ever sees.

371 results · 1 filter active ·
Results by year
20002025
Publication date
Categories
Machine labels · sparse coverage
Evidence
Language
Type
Citations
An unlabeled work is unknown, not a negative. Label coverage is reported on every query.
371 works in the cohort · of 4,299,418page 2 of 8

Labels cover 1 of 371 works in this cohort. The rest are unlabeled, which is not a negative label: the label table is sparse today and grows as labeling rounds land.

Distilled predictions cover 371 of 371 works in this cohort. Predictions are machine_predicted_unvalidated. The Gemma side is a direct model label for every work (title-only); the Codex side is a distilled, calibrated classifier. Candidate is the union; consensus is the intersection.

affno abstractunlabeled
The evolution of Java build systems
Shane McIntosh, Bram Adams, Ahmed E. Hassan
2011· article· en· Empirical Software Engineering· Computer Science
machine prediction:candidate · noneconsensus · none
73
citations
affno abstractunlabeled
Recommending reference API documentation
Martin P. Robillard, Yam Bahadur Chhetri
2014· article· en· Empirical Software Engineering· Computer Science
machine prediction:candidate · noneconsensus · none
67
citations
affno abstractunlabeled
Studying and detecting log-related issues
Mehran Hassani, Weiyi Shang, Emad Shihab, Nikolaos Tsantalis
2018· article· en· Empirical Software Engineering· Computer Science
machine prediction:candidate · noneconsensus · none
64
citations
affno abstractunlabeled
Qualitative research in software engineering
Tore Dybå, Rafael Prikladnicki, Kari Rönkkö, Carolyn Seaman, Jonathan Sillito
2011· article· en· Empirical Software Engineering· Computer Science
machine prediction:candidate · metaresearchconsensus · none
59
citations
affno abstractunlabeled
Coevolution of variability models and related software artifacts
Leonardo Passos, Leopoldo Teixeira, Nicolas Dintzner, Sven Apel, Andrzej Wąsowski, Krzysztof Czarnecki +2 more
2015· article· en· Empirical Software Engineering· Computer Science
machine prediction:candidate · noneconsensus · none
57
citations
affno abstractunlabeled
On the unreliability of bug severity data
Yuan Tian, Nasir Ali, David Lo, Ahmed E. Hassan
2015· article· en· Empirical Software Engineering· Computer Science
machine prediction:candidate · metaresearchconsensus · none
51
citations
affno abstractunlabeled
An empirical study of software release notes
Surafel Lemma Abebe, Nasir Ali, Ahmed E. Hassan
2015· article· en· Empirical Software Engineering· Computer Science
machine prediction:candidate · metaresearchconsensus · none
51
citations
affno abstractunlabeled
Studying high impact fix-inducing changes
Ayşe Tosun, Emad Shihab, Yasukata Kamei
2015· article· en· Empirical Software Engineering· Computer Science
machine prediction:candidate · metaresearchconsensus · none
46
citations
affno abstractunlabeled
License usage and changes: a large-scale study on gitHub
Christopher Vendome, Gabriele Bavota, Massimiliano Di Penta, Mario Linares‐Vásquez, Daniel M. Germán, Denys Poshyvanyk
2016· article· en· Empirical Software Engineering· Computer Science
machine prediction:candidate · bibliometricsconsensus · none
45
citations
afffundno abstractunlabeled
Bugs in large language models generated code: an empirical study
Florian Tambon, Arghavan Moradidakhel, Amin Nikanjam, Foutse Khomh, Michel C. Desmarais, Giuliano Antoniol
2025· article· en· Empirical Software Engineering· Computer Science
machine prediction:candidate · noneconsensus · none
44
citations

How this was built: Screen · Findings · About