Functional near-infrared spectroscopy of medical students answering various item types
Bibliographic record
Abstract
Traditionally, the effect of assessment item types including true/false questions (TFQs), multiple-choice questions (MCQs), short answer questions (SAQs), and case scenario questions (CSQs) is examined through psychometric qualities or student interviews. However, brain activity while answering such questions or items remains unknown. Functional near-infrared spectroscopy (fNIRS) can be used to safely measure cerebral cortex hemodynamic response during various tasks. Hence, this fNIRS study aimed to determine differences in frontotemporal cortex activity as medical students answered TFQs, MCQs, SAQs, and CSQs.In total, 24 medical students (13 males and 11 females) were recruited in this study during their mid-psychiatry posting. Oxy-hemoglobin and deoxy-hemoglobin levels in the frontal and temporal regions were measured with a 52-channel fNIRS system. Participants answered 9-18 trials under each of the four types of tasks that were based on their psychiatry curriculum during fNIRS measurements. The area under the oxy-hemoglobin curve (AUC) for each participant and each item type was derived. Repeated measures ANOVA with post-hoc Bonferroni-corrected pairwise comparisons were used to determine differences in oxy-hemoglobin AUC between TFQs, MCQs, SAQs, and CSQs.Oxy-hemoglobin AUC was highest during the CSQs, followed by SAQs, MCQs, and TFQs in both the frontal and temporal regions. Statistically significant differences between different types of items were observed in oxy-hemoglobin AUC of the frontal region (p ≤ 0.001). Oxy-hemoglobin AUC in the frontal region was significantly higher during the CSQs than TFQ (p = 0.005) and during the SAQ than TFQ (p = 0.025). Although the percentage of correct responses was significantly lower in MCQ than in the other item types, there was no correlation between the percentage of correct response and oxy-hemoglobin AUC in both regions for all four item types (p > 0.05).CSQs and SAQs elicited greater hemodynamic response than MCQs and TFQs in the prefrontal cortex of medical students. This suggests that more cognitive skills may be required to answer CSQs and SAQs.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".