MétaCan
Menu
Retour à la cohorte
Enregistrement W4393745689 · doi:10.5281/zenodo.5907847

ApacheJIT: A Large Dataset for Just-In-Time Defect Prediction

2022· dataset· en· W4393745689 sur OpenAlexaff
Hossein Keshavarz, Meiyappan Nagappan

Notice bibliographique

RevueZenodo (CERN European Organization for Nuclear Research) · 2022
Typedataset
Langueen
DomaineEngineering
ThématiqueIndustrial Vision Systems and Defect Detection
Établissements canadiensUniversity of Waterloo
Organismes subventionnairesnon disponible
Mots-clésComputer science

Résumé

récupéré en direct d'OpenAlex

<strong>ApacheJIT: A Large Dataset for Just-In-Time Defect Prediction</strong> This archive contains the <strong>ApacheJIT</strong> dataset presented in the paper "ApacheJIT: A Large Dataset for Just-In-Time Defect Prediction" as well as the replication package. The paper is submitted to <strong>MSR 2022 Data Showcase Track</strong>. The datasets are available under directory <em>dataset</em>. There are 4 datasets in this directory. <strong>apachejit_total.csv</strong>: This file contains the entire dataset. Commits are specified by their identifier and a set of commit metrics that are explained in the paper are provided as features. Column <em>buggy</em> specifies whether or not the commit introduced any bug into the system. <strong>apachejit_train.csv</strong>: This file is a subset of the entire dataset. It provides a balanced set that we recommend for models that are sensitive to class imbalance. This set is obtained from the first 14 years of data (2003 to 2016). <strong>apachejit_test_large.csv</strong>: This file is a subset of the entire dataset. The commits in this file are the commits from the last 3 years of data. This set is not balanced to represent a real-life scenario in a JIT model evaluation where the model is trained on historical data to be applied on future data without any modification. <strong>apachejit_test_small.csv</strong>: This file is a subset of the test file explained above. Since the test file has more than 30,000 commits, we also provide a smaller test set which is still unbalanced and from the last 3 years of data. In addition to the dataset, we also provide the scripts using which we built the dataset. These scripts are written in Python 3.8. Therefore, Python 3.8 or above is required. To set up the environment, we have provided a list of required packages in file <em>requirements.txt</em>. Additionally, one filtering step requires GumTree [1]. For Java, GumTree requires Java 11. For other languages, external tools are needed. Installation guide and more details can be found here. The scripts are comprised of Python scripts under directory <em>src</em> and Python notebooks under directory <em>notebooks</em>. The Python scripts are mainly responsible for conducting GitHub search via GitHub search API and collecting commits through PyDriller Package [2]. The notebooks link the fixed issue reports with their corresponding fixing commits and apply some filtering steps. The bug-inducing candidates then are filtered again using <em>gumtree.py</em> script that utilizes the GumTree package. Finally, the remaining bug-inducing candidates are combined with the clean commits in the <em>dataset_construction</em> notebook to form the entire dataset. More specifically, <em>git_token</em> handles GitHub API token that is necessary for requests to GitHub API. Script <em>collector</em> performs GitHub search. Tracing changed lines and git annotate is done in <em>gitminer</em> using PyDriller. Finally, <em>gumtree</em> applies 4 filtering steps (number of lines, number of files, language, and change significance). References: <strong>1. GumTree</strong> https://github.com/GumTreeDiff/gumtree Jean-Rémy Falleri, Floréal Morandat, Xavier Blanc, Matias Martinez, and Martin Monperrus. 2014. Fine-grained and accurate source code differencing. In ACM/IEEE International Conference on Automated Software Engineering, ASE ’14,Vasteras, Sweden - September 15 - 19, 2014. 313–324 <strong>2. PyDriller</strong> https://pydriller.readthedocs.io/en/latest/ Davide Spadini, Maurício Aniche, and Alberto Bacchelli. 2018. PyDriller: Python Framework for Mining Software Repositories. In Proceedings of the 2018 26th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering(Lake Buena Vista, FL, USA)(ESEC/FSE2018). Association for Computing Machinery, New York, NY, USA, 908–911

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,002
score de la tête « metaresearch » (Gemma)0,001
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesMéta-épidémiologie (sens strict), Études des sciences et des technologies, Charge utile insuffisante (le modèle a refusé de juger)
Catégories consensuellesCharge utile insuffisante (le modèle a refusé de juger)
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: Sans objet
GenreSignal candidat: Jeu de données · Signal consensuel: Jeu de données
Score de désaccord entre enseignants0,060
Score d'incertitude au seuil1,000

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0020,001
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0000,000
Bibliométrie0,0010,001
Études des sciences et des technologies0,0010,000
Communication savante0,0010,000
Science ouverte0,0010,001
Intégrité de la recherche0,0000,001
Charge utile insuffisante (le modèle a refusé de juger)0,0640,004

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,037
Tête enseignante GPT0,251
Écart entre enseignants0,214 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; les deux têtes enseignantes s’accordent sur ce qui est montré ici.

Devis d'étudeSans objet
Domainenon disponible
GenreJeu de données

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations1
Publié2022
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueZenodo (CERN European Organization for Nuclear Research)Même sujetIndustrial Vision Systems and Defect DetectionTravaux en français237 207