Syntactic multilingual probing of pre-trained language models of code
Bibliographic record
Abstract
Pre-trained language models (PLMs) have demonstrated remarkable abilities in coding tasks, establishing themselves as a state-of-the-art technique in machine learning for code. However, due to their deep neural network-based structure, PLMs function as black-box systems, making it crucial to understand the types of information they actually learn. Recent studies indicate that PLMs possess cross-lingual capabilities, allowing them to generalize to unseen programming languages and outperform monolingual models when trained in a multilingual setting. Nonetheless, the reasons behind these cross-lingual abilities remain largely uncharted and remain open questions. In this paper, we explore this phenomenon through a syntactic perspective. Specifically, we build on our prior work, the AST-Probe, a probing methodology that evaluates whether a PLM encodes the complete grammatical structure of a programming language. This probe identifies a syntactic subspace within the PLM’s vector representations, which is then used to reconstruct ASTs. We extend this approach in two ways. First, we conducted experiments on eight programming languages and eight PLMs and found that: (1) this syntactic structure can be extracted in all cases, (2) CodeBERT and GraphCodeBERT excel at encoding ASTs, and (3) syntactic knowledge resides in the middle layers of all PLMs, with a distribution that is independent of the programming language. Secondly, we mathematically adapt the AST-Probe to a multilingual setting and apply it to CodeBERT. Our findings provide evidence that CodeBERT learns cross-lingual representations of programming languages syntax.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".