MétaCan
Menu
Cohort builder

4,299,418 works, Canadian by any of four routes.

Every filter state is a URL; the URL is the query; the query is citable via /q/⟨hash⟩. The page, the API and the export parse the same parameters.

The current cohort, streamed from the database: every work column, the machine labels, the provisional scores, and the per-row validation status. Exports are capped at 100,000 rows. Mints a permanent /q/ link for this exact query. The same filters always produce the same link, whoever asks.

Search term
Author
Year range
Sort
Language
Type
Field
Venue
Topic
Software Engineering Research
Retraction
Abstract
Evidence source
Study design
Label agreement
Label status

Direct Codex and Gemma labels are unvalidated and sparse. Distilled predictions cover the full frame and are also unvalidated. Choose the evidence source explicitly; absence of a direct label is never a negative label.

affaffiliation
fundfunder
venuejournal
aboutaboutness

The four routes compose: require the funder route and exclude affiliation to get the funder-only stratum no affiliation-based frame ever sees.

3,468 results · 1 filter active ·
Results by year
20002025
Publication date
Categories
Machine labels · sparse coverage
Evidence
Language
Type
Citations
An unlabeled work is unknown, not a negative. Label coverage is reported on every query.
3,468 works in the cohort · of 4,299,418page 12 of 70

Labels cover 10 of 3,468 works in this cohort. The rest are unlabeled, which is not a negative label: the label table is sparse today and grows as labeling rounds land.

Distilled predictions cover 3,468 of 3,468 works in this cohort. Predictions are machine_predicted_unvalidated. The Gemma side is a direct model label for every work (title-only); the Codex side is a distilled, calibrated classifier. Candidate is the union; consensus is the intersection.

afffundunlabeled
CLEVER
Mathieu Nayrolles, Abdelwahab Hamou‐Lhadj
2018· article· en· Computer Science
machine prediction:candidate · insufficient_payloadconsensus · none
50
citations
affunlabeled
Natural Software Revisited
Musfiqur Rahman, Dharani Palani, Peter C. Rigby
2019· article· en· Computer Science
machine prediction:candidate · noneconsensus · none
49
citations
affno abstractunlabeled
Taupe : Visualizing and analyzing eye-tracking data
Benoît De Smet, Lorent Lempereur, Zohreh Sharafi, Yann‐Gaël Guéhéneuc, Giuliano Antoniol, Naji Habra
2012· article· en· Science of Computer Programming· Computer Science
machine prediction:candidate · noneconsensus · none
49
citations
affunlabeled
A Study of C/C++ Code Weaknesses on Stack Overflow
Haoxiang Zhang, Shaowei Wang, Heng Li, Tse-Hsun Chen, Ahmed E. Hassan
2021· article· en· IEEE Transactions on Software Engineering· Computer Science
machine prediction:candidate · noneconsensus · none
48
citations
affunlabeled
The evolution of ANT build systems
Shane McIntosh, Bram Adams, Ahmed E. Hassan
2010· article· en· Computer Science
machine prediction:candidate · noneconsensus · none
48
citations
affunlabeled
Fishtail
Nicholas Sawadsky, Gail C. Murphy
2011· article· en· Computer Science
machine prediction:candidate · insufficient_payloadconsensus · none
47
citations
affno abstractunlabeled
Studying high impact fix-inducing changes
Ayşe Tosun, Emad Shihab, Yasukata Kamei
2015· article· en· Empirical Software Engineering· Computer Science
machine prediction:candidate · metaresearchconsensus · none
46
citations
affno abstractunlabeled
Mining API usage scenarios from stack overflow
Gias Uddin, Foutse Khomh, Chanchal K. Roy
2020· article· en· Information and Software Technology· Computer Science
machine prediction:candidate · noneconsensus · none
46
citations
affunlabeled
VISUALIZING THE EVOLUTION OF SOFTWARE USING SOFTCHANGE
Daniel M. Germán, Abram Hindle
2006· article· en· International Journal of Software Engineering and Knowledge Engineering· Computer Science
machine prediction:candidate · noneconsensus · none
46
citations
affno abstractunlabeled
License usage and changes: a large-scale study on gitHub
Christopher Vendome, Gabriele Bavota, Massimiliano Di Penta, Mario Linares‐Vásquez, Daniel M. Germán, Denys Poshyvanyk
2016· article· en· Empirical Software Engineering· Computer Science
machine prediction:candidate · bibliometricsconsensus · none
45
citations
affunlabeled
*J
Bruno Dufour, Laurie Hendren, Clark Verbrugge
2003· article· en· Computer Science
machine prediction:candidate · insufficient_payloadconsensus · none
45
citations
affunlabeled
The value of design rationale information
Davide Falessi, Lionel Briand, Giovanni Cantone, Rafael Capilla, Philippe Kruchten
2013· article· en· ACM Transactions on Software Engineering and Methodology· Computer Science
machine prediction:candidate · noneconsensus · none
45
citations
affunlabeled
Why did this code change
Sarah Rastkar, Gail C. Murphy
2016· article· en· Computer Science
machine prediction:candidate · noneconsensus · none
45
citations

How this was built: Screen · Findings · About