Urdu as a First Language: The Impact of Script on Reading in the L1 and English as a Second Language
Bibliographic record
Abstract
Urdu is a classic example of digraphia, a linguistic situation in which different scripts are used to write the same language (Ahmad, 2011). The analysis of orthographic practices of reading and writing Urdu in Arabic script versus in Urdu script reveals that Muslims learning to read the Quran in Western countries are not explicitly aware of the features specific to Urdu script. The present study examined awareness of script similarity; suggesting that bilinguals who read both scripts, Urdu and Arabic would have an advantage in acquiring L1 through the same scripts. Fifty Canadian bilingual children (6-10 years) were tested for language ability, cognitive and phonological processing skills in two languages: Urdu their L1 and English their L2. In contrast to English, Urdu was written in an adapted version of Arabic script. Groups were created based on whether they were above or below the standardized mean on the Urdu and/or Arabic measures. A binary logistic regression showed that there was a significant difference between readers who are good and/or poor at one or both languages, in terms of L2 decoding. The correlations between L1 and L2 phonological awareness showed that phonological awareness is not a language specific mechanism (Comeau et al., 1999). These skills predict word recognition cross-linguistically as a result of the linguistic interdependence between L1 and L2. Recent research on L2 literacy development suggests the need to examine transfer of literacy skills on a case-by-case basis for each language, based on similarities and differences between L1 and L2 scripts, particularly if the readers show low levels of literacy in one script, in this case their L1 (Genesse et al., 2006).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".