Authoring Tools for Designing Intelligent Tutoring Systems: a Systematic Review of the Literature
Bibliographic record
Abstract
Authoring tools have been broadly used to design Intelligent Tutoring Systems (ITS). However, ITS community still lacks a current understanding of how authoring tools are used by non-programmer authors to design ITS. Hence, the objective of this work is to review how authoring tools have been supporting ITS design for non-programmer authors. In order to meet our goal, we conduct a Systematic Literature Review (SLR) to identify the primary studies on the use of ITS authoring tools, following a pre-defined review protocol. Among the 4622 papers retrieved from seven digital libraries published from 2009 to June 2016, 33 papers are finally included after applying our exclusion and inclusion criteria. We then identify the main ITS components authored, the ITS types designed, the features used to facilitate the authoring process, the technologies used to develop authoring tools and the time at which authoring occurs. We also look for evidence of the benefits of ITS authoring tools. In summary, the main findings of this work are: (1) there is empirical evidence of the benefits (i.e., mainly in terms of effectiveness, efficiency, quality of authored artifacts, and usability) of using ITS authoring tools for non-programmer authors, specially to aid authoring of learning content and to support authoring of model-tracing/cognitive and example-tracing tutors; 2) domain and pedagogical models have been much more targeted by authoring tools; (3) several ITS types have been authored, with an emphasis on model-tracing/cognitive and example-tracing tutors; (4) besides providing features for authoring all four ITS components, current authoring tools are also presenting general features (e.g., view learners’ statistics and reuse tutor design) to create broader authoring tools; (5) a great diversity of technologies, which include AI techniques, software solutions and distributed technologies, are used to develop ITS authoring tools; and (6) authoring tools have been mainly targeting ITS design before students’ instruction, but works are also addressing authoring during and/or post-instruction relying both on human and artificial intelligence. We conclude this work by showing several promising research opportunities that are quite important and interesting but underexplored in current research and practice.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.039 | 0.127 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.006 | 0.005 |
| Bibliometrics | 0.026 | 0.022 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.005 | 0.007 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".