When words matter: evaluating the quality of open educational resources through lexicon.
Bibliographic record
Abstract
There is a general agreement on the advantages of open education in e-learning in terms of inclusiveness and equity (Wiley, 2010). However, there is a relatively less developed debate about the quality evaluation of open educational resources (OERs), accessible for free and without spatial or time limits. Usually the OERs’ quality is assessed on structural features, i.e. accessibility, usability, learning goals, possibility for the learner to assess his/her progress during learning (McGill et al. 2013; Ehlers & Joosten, 2009). When it comes to the content of OERs, this is considered mainly at the macro-level (i.e. text cohesion and coherence, cultural differences among potential users, gender differences), and not at the micro-level of its lexical components. Conversely, one of the most debated issues within the quality of education, both traditional face-to-face education and distance technologyenhanced education, is on how the educational mediation takes place in order to develop transversal competences and basic skills linked to the use of language (Marconi, 1997). W. Nagy, expert in vocabulary development, asks, “Which words should a teacher teach?” (2011) and, more in general, then he considers how, irrespectively of the disciplinary contents, words that compose the educational \nmessage do have an intrinsic value and should descend from an intentional choice in the instructional design phase. \nMeasures related to text readability are generally based on the frequency of a word in a corpus of sufficiently wide dimensions. Texts with rare words are more difficult to understand than those that contain common words. However, the emphasis on the use of these tools to study the adequacy of textbooks and learning materials has been criticized (Davison - Green, 1988), as not always common words are easy to define (e.g. the article “the” it is very common but with difficult to define) whereas rare words in a written texts can be easier as, for instance, they are common in the spoken language (e.g. “t-shirt” or “fireman”). Building a structure of semantic relationships between words (similarities/oppositions, inclusion/exclusion, semantic fields) is instead one of the ways to help memorization, and it makes likely the passage from receptive to productive lexicon in learners. On these premises, it can be envisaged a set of criteria to evaluate the appropriateness of an OER text with respect to its learning goals. \nThe paper discusses the assumptions to evaluate the lexical aspects of OERs’ instructional messages, and presents the first results obtained from an exploratory study carried out on a set of OERs’ in Italian language on a variety of contents.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".