Ihan legit asiaa : Finglishin käytön tarkastelua Suomi24-foorumilla trenditutkimuksen keinoin
Bibliographic record
Abstract
Tässä tutkielmassa pohdin finglishin esiintyvyyttä foorumikeskusteluissa eli sitä, kuinka usein verkkokeskusteluissa kirjoitetaan yhden kommentin sisällä suomea ja englantia ”sekaisin”. Tutkielman aineisto on kerätty Suomi24-foorumin korpuksesta, joka kattaa kaikki Suomi24-foorumin keskustelut vuosilta 2001–2020. Varsinainen aineisto koostuu virkkeistä, joissa on mainittu joko binge, legit tai outfit. Aineiston pohjalta analysoin valittujen sanojen sopeutumista suomen kieleen, millaisissa konteksteissa sanat esiintyvät sekä sanojen yleisyyttä korpuksessa. \n \nTutkielman taustalla on pohdintaa siitä, mitä on verkkokieli ja mitä finglish tarkoittaa. Lähestyn korpuslähtöistä tutkimustani sekä määrällisesti että laadullisesti analysoiden ottaen huomioon sosiolingvistiikan näkökulmia. Tutkimuksessani yhdistyy trendi- ja tapaustutkimuksen piirteitä. Hyödynnän aineiston analysoinnissa Maria Vilkunan teosta Suomen lauseopin perusteet (1996), verkkoversiota Iso Suomen Kielioppi -teoksesta (2008) sekä omaa kielitajuani. Kolmen tapauksen tuloksista on mahdollista saada osviittaa finglishin käytön suosiosta, sillä oletuksena on jokaisen esiintymän taustalla olevan eri ihminen. \n \nTutkimustulokset osoittavat, että lähes 80 prosentissa kaikista aineiston esiintymistä hakusana esiintyy sellaisenaan eikä mukaudu suomen kielioppiin. Finglish ei siis osoittaudu kovin suosituksi kielimuodoksi ainakaan Suomi24-foorumilla vuosina 2001–2020. Tähän saattaa vaikuttaa hakusanat tai niiden vähyys. Tulevaisuudessa tutkimusta voisi kehittää aineiston kannalta relevantimmaksi sekä metodin kannalta tarkemmaksi.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.669 | 0.008 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".