MétaCan
Menu
Back to cohort
Record W2783444782 · doi:10.1109/icdim.2017.8244675

Code style analytics for the automatic setting of formatting rules in IDEs: A solution to the Tabs vs. Spaces Debate

2017· article· en· W2783444782 on OpenAlexaff
Eric Torunski, M. Omair Shafiq, Anthony Whitehead

Bibliographic record

Venuenot available
Typearticle
Languageen
FieldComputer Science
TopicSoftware Engineering Research
Canadian institutionsCarleton University
Fundersnot available
KeywordsDisk formattingComputer scienceSource codeCode reviewCode (set theory)Source lines of codeStyle (visual arts)AnalyticsProgramming languageSoftwareClass (philosophy)Static program analysisKPI-driven code analysisSoftware engineeringWorld Wide WebDatabaseSoftware developmentArtificial intelligenceOperating systemSet (abstract data type)

Abstract

fetched live from OpenAlex

The use of code style is very important since it conveys meaning as well as intent of source code. Developers are used to reading code according to their preferred style but those guidelines of proper style vary among software teams, and even different companies. Code style decisions are typically made by managers of software developers, but we would like to investigate how common the different variations of code style are. There are also automated tools to convert code style in a file, however the tools must be configured manually. In this paper, we present a tool for the collection and analysis of code style metrics. We demonstrate the feasibility of scanning existing source code to automatically generate the code style rules for existing tools. We also look at the results of our data mining to look at trends in source code. We perform a quantitative analysis on source code for questions like: How many functions are in a class, on average? How many lines of code are in a method, on average? We also present graphs of the distribution of these data, as well look at special cases of outliers.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.002
metaresearch head score (Gemma)0.003
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.750
Threshold uncertainty score0.430

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0020.003
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0020.001
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.033
GPT teacher head0.315
Teacher spread0.282 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations19
Published2017
Admission routes1
Has abstractyes

Explore more

Same topicSoftware Engineering ResearchFrench-language works237,207