An Exception to the Rule? Lone French Nouns in Tunisian Arabic
Bibliographic record
Abstract
Reports on language mixing involving Arabic often qualify that language as resistant to constraints operating on other language pairs. But many fail to situate the purported violations with respect to recipient and donor languages, making it impossible to ascertain whether these are exceptional code-switches or (nonce) borrowings; isolated cases or robust patterns. We address these issues through variationist analysis of Tunisian Arabic/French bilingual discourse. Focusing on conflict sites that reveal which grammar is operative when the other language is accessed, we compare quantitatively the behavior of lone French-origin nouns in Arabic with their counterparts in both donor and recipient languages. Despite a higher order community resistance to morphological inflection of other-language items, results show treatment of French nouns to be consistent with the (variable) grammar of Arabic and different from that of French. Applying the same accountable methodology to the contentious French det+n sequences (“constituent insertions”) shows that most are integrated in the same way as their lone counterparts. These too are treated as (compound) borrowings, largely motivated by the semantic imperative of expressing plurality while eschewing inflection. As borrowings, they do not constitute exceptions to code-switching constraints, confirming that the status of mixed items cannot be determined in isolation; they must be contextualized with respect to the remainder of the bilingual system, including donor, recipient, and other mixed-language elements.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".