Heterogeneity of Taiwan’s indigenous population: possible relation to prehistoric Mongoloid dispersals
Bibliographic record
Abstract
Taiwan's 9 indigenous tribes (Tsou, Bunun, Paiwan, Rukai, Atayal, Saisiat, Ami, Puyuma, Yami) are highly homogeneous within each tribe, but diversified among the different tribes due to long-term isolation, most probably since Taiwan became an island about 12,000 years ago. Homogeneity of each tribe is evidenced by many HLA-A,B,C alleles having the world's highest ever reported frequencies, e.g. A24 (86.3%), A26 (18.8%), Cw10 (36.8%), Cw7 (66%), Cw8 (32.1%), B13 (27.9%), B62 (37.4%), B75 (18%), B39 (53.5%), B60 (33.3%), and B48 (24%). Also, all of these tribes have HLA class I haplotype frequencies greater than 10%, with A24-Cw7-B39 in Saisiat (44.5%) being the highest, suggesting Taiwan's indigenous tribes are probably the most homogeneous ( the "purest") population in the world. A24-Cw8-B48, A24-Cw10-B60 and A24-Cw9-B61 found common to many Taiwan indigenous tribes, have also been observed in Maori, Papua New Guinea Highlanders, Orochons, Mongolians, Inuit, Japanese, Man, Buryat, Yakut, Tlingit, Tibetans and Thais. These findings suggest Taiwan's indigenous groups are more or less genetically related to both northern and southern Asians. Principal component analysis and the phylogenetic tree (using the neighbor-joining method) showed close relationship between the indigenous groups and Oceanians. This relationship supports the hypothesis that Taiwan was probably on the route of prehistoric Mongoloid dispersals that most likely took place along the coastal lowland of the Asian continent (which is under the sea today). Cultural anthropology also suggests a relationship between Taiwan's indigenous tribes and southern Asians and to a lesser extent, northern Asians. However, the indigenous groups show little genetic relationship to current southern and northern Han Chinese.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".