Mapping the intellectual landscape of heritage language transmission: a bibliometric review of strategies and sociocultural dynamics within families (2014–2024)
Bibliographic record
Abstract
This review provides a bibliometric examination of the academic literature pertaining to Heritage Language Transmission in Strategies and Sociocultural Dynamics within Familial Contexts, employing a dataset comprising 543 entries sourced from the Web of Science Core Collection during the period from 2014 to 2024. The primary objective of this analysis was to elucidate the intellectual framework, thematic developments, and collaborative dynamics that typify this expanding area of study. The software CiteSpace (version 6.3.R1) was utilized to perform co-citation, co-authorship, and keyword co-occurrence analyses. A refined corpus of 529 documents yielded a network consisting of 214 nodes and 535 connections, thereby providing insights into the interrelatedness of prominent authors, institutions, and thematic clusters. The results indicated a significant increase in academic production, with a notable acceleration observed post-2020. Central research themes encompassed family language policy, bilingual development in children, translanguaging phenomena, emotional dynamics, and identity negotiation processes. High-impact contributors, including Natalia Meir, Tanja Kupisch, and Johanne Paradis, were identified based on metrics such as citation frequency, degree of influence, and centrality within the network. At the institutional level, the University of Toronto, Bar Ilan University, and UiT the Arctic University of Tromsø emerged as pivotal canters of scholarly influence. On a national scale, the United States was positioned as the leader in terms of academic output and network centrality, followed by Germany, Canada, and England. The keyword analysis underscored “acquisition,” “heritage language,” and “family language policy” as predominant constructs, while emerging terminology such as “translanguaging,” “input quantity,” and “language beliefs” indicated potential new trajectories for research. This bibliometric review furnishes a detailed cartography of the heritage language transmission research domain and posits a data-driven basis for forthcoming inquiries, particularly those that focus on culturally embedded language practices, family-oriented adaptation strategies, and policy-driven interventions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.012 | 0.047 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.123 | 0.183 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.005 | 0.006 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".