Den fjärde kvadrantens dilemma : Kunskapsbedömning i en föränderlig historiekultur
Bibliographic record
Abstract
The purpose of this dissertation is to explore how history teachers in Swedish upper-secondary school perceive voluntary external tests in history. The study focuses on history teachers' perceptions of 1) the design of the tests, 2) the implementation of the tests, 3) the use of the tests and 4) the impact of the tests on their assessment practice. The overall aim is to investigate how historical culture, as it is expressed in the school subject of history, changes over time. The analytical focus for the study has been to explore the significance of knowledge assessment in general, and voluntary external tests in particular, for this change. The temporal focus for the study has been the 2011 Swedish school policy reforms manifested in the curriculum for upper-secondary school (Gy11). The theoretical framework combines theory of history didactics and theory of assessment. Historical culture and assessment culture are the central concepts of the study. The overall hypothesis of the study is that the relationship between the school's historical-cultural dimension and the assessment-cultural dimension contains a latent tension referred to as the dilemma of the fourth quadrant. The empirical material was collected from qualitative interviews with eight history teachers. The first sub-study included four teachers during the year 2009 who had used a voluntary external test named the History Teachers' Test (HLP). The second sub study included four teachers during the year 2016 who had used a voluntary external test named the Course Test in History (KP). The results show that history teachers perceive and handle the external tests in different ways. One possible interpretation of this difference is that the tests respond to different needs or problems. In the study, two such problem areas have emerged, the equivalence problem and the alignment problem. A conclusion based on the empirical results is that the HLP-teachers use the test results for a summative purpose to deal with an equivalence problem, while the KP-teachers use the test tools for a formative purpose to deal with an alignment problem. The results also show that there is strong a connection between the school's historical-cultural dimension and assessment-cultural dimension, manifested in the history teachers' different ways of perceiving and using voluntary external tests.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.012 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.005 | 0.012 |
| Scholarly communication | 0.014 | 0.006 |
| Open science | 0.001 | 0.005 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".