Analysis and evaluation of college entrance examination questions based on clustering algorithm
Bibliographic record
Abstract
This paper analyzes and evaluates high school examination questions based on machine learning.The study first introduces Bloom's classification method and constructs a categorized dataset of high school exam questions according to three steps of data collection, data annotation and data analysis.Then an automatic assessment model (WoBERT-CNN) based on WoBERT and Text-CNN is designed.The semantic similarity of word vector mapping is used to label the cases for determination, the improved WoBERT encoder is used to represent the text in word vectors, Text-CNN is used as a text classifier to extract the textual semantic features, and the features are integrated and screened, so as to realize the automatic classification of the cases in Bloom's taxonomy.Finally, based on the deep representation framework, the text information of the test questions is deeply mined and utilized to establish the relationship between the text of the test questions and the actual difficulty, and to realize the difficulty prediction of the test questions.The classification accuracy of the WoBERT-CNN model reaches more than 92%.The prediction error range of the H-MIDP model on the score rate of the test questions is between 1.3% and 3.2%, which is not too far from the real value.In conclusion, the automatic assessment model and difficulty prediction model designed in this paper can be applied in the analysis and evaluation of high school test questions, helping the high school test paper proposition and talent cultivation strategy.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".