Conserving the Simalungun Language Maintenance through Demographic Community: The Analysis of Taboo Words Across Times
Bibliographic record
Abstract
The research was intended to describe the use of Simalungun taboo words across times in Simalungun (1930-2021). The language of Simalungun is spoken by people living outside the district of Simalungun, North of Sumatera and other people. This research was carried out in a multi-case descriptive qualitative design. Descriptual qualitative research design was defined as a social science research approach that emphasised the collection, use of inductive thinking and understanding of descriptive data in natural environments. While multi case is defined as a study which is using two or more subjects, settings, or depositories of data (Bogdan & Biklen, 1982). Documentation, interviews and observations of participants were used to collect data on linguistic taboos. The data sources were collected from 45 informants of different ages (1930-2021) and sexes who reside in Pematangsiantar, Pematangraya and Saribudolok. After having analyzed the collected data, the research finding showed that there were 62 words out of 106 the taboo words of ten categories: sexual organ, sexual activity, cursing, swearing, calling people, action, disease, dwelling ghost and name of God which were used stably across time (from 1930 to 2021) in Simalungun are 62 words, out of 106 words.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.006 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.004 | 0.006 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".