Improving student's problem-solving ability as well as conceptual understanding without sacrificing the physics content of a class
Bibliographic record
Abstract
Four sections of introductory physics for physical scientists and engineers (about 180 students each) are compared. One section, treatment group, was organized so that students worked to learn the classical ideas connecting forces and motion over the first 6 weeks of the 10 week quarter and then used the final 4 weeks to apply those principles to algebraically complicated problems. The other sections learned ideas at essentially the same time as calculations over the entire 10 weeks of the quarter. The treatment group and one of the control sections were taught by the same instructor, had identical curricular materials and this instructor was blind to the comparison measure, the final exam. After controlling for GPA as well as for incoming conceptual understanding, the treatment group was found (with greater than 99% confidence) to perform better on the final exam than the control group taught by the same instructor and, by a similar measure, the treatment group performed significantly better than any other section. The treatment group also had higher conceptual learning gains and so should be better prepared for later learning.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".