Automatic modeling of student characteristics with interaction and physiological data using machine learning: A review
Bibliographic record
Abstract
Student characteristics affect their willingness and ability to acquire new knowledge. Assessing and identifying the effects of student characteristics is important for online educational systems. Machine learning (ML) is becoming significant in utilizing learning data for student modeling, decision support systems, adaptive systems, and evaluation systems. The growing need for dynamic assessment of student characteristics in online educational systems has led to application of machine learning methods in modeling the characteristics. Being able to automatically model student characteristics during learning processes is essential for dynamic and continuous adaptation of teaching and learning to each student's needs. This paper provides a review of 8 years (from 2015 to 2022) of literature on the application of machine learning methods for automatic modeling of various student characteristics. The review found six student characteristics that can be modeled automatically and highlighted the data types, collection methods, and machine learning techniques used to model them. Researchers, educators, and online educational systems designers will benefit from this study as it could be used as a guide for decision-making when creating student models for adaptive educational systems. Such systems can detect students' needs during the learning process and adapt the learning interventions based on the detected needs. Moreover, the study revealed the progress made in the application of machine learning for automatic modeling of student characteristics and suggested new future research directions for the field. Therefore, machine learning researchers could benefit from this study as they can further advance this area by investigating new, unexplored techniques and find new ways to improve the accuracy of the created student models.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".