Analysis and evaluation of college entrance examination questions based on clustering algorithm
Notice bibliographique
Résumé
This paper analyzes and evaluates high school examination questions based on machine learning.The study first introduces Bloom's classification method and constructs a categorized dataset of high school exam questions according to three steps of data collection, data annotation and data analysis.Then an automatic assessment model (WoBERT-CNN) based on WoBERT and Text-CNN is designed.The semantic similarity of word vector mapping is used to label the cases for determination, the improved WoBERT encoder is used to represent the text in word vectors, Text-CNN is used as a text classifier to extract the textual semantic features, and the features are integrated and screened, so as to realize the automatic classification of the cases in Bloom's taxonomy.Finally, based on the deep representation framework, the text information of the test questions is deeply mined and utilized to establish the relationship between the text of the test questions and the actual difficulty, and to realize the difficulty prediction of the test questions.The classification accuracy of the WoBERT-CNN model reaches more than 92%.The prediction error range of the H-MIDP model on the score rate of the test questions is between 1.3% and 3.2%, which is not too far from the real value.In conclusion, the automatic assessment model and difficulty prediction model designed in this paper can be applied in the analysis and evaluation of high school test questions, helping the high school test paper proposition and talent cultivation strategy.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,003 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,001 | 0,001 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».