Enhancing Mathematical Reasoning Through Autonomously Learning Knowledge
Bibliographic record
Abstract
Enabling machines to solve mathematical problems is a vital endeavor in developing intelligence that emulates human-like thinking and reasoning. However, most existing approaches focus on reconstructing human comprehension of problems, which are still far from enough since they neglect the fundamental human ability to learn knowledge from experiences. In this article, we focus on empowering models with the cognitive capacity to autonomously learn knowledge from mathematical problem-solving. We first propose a Cognitive Solver (CogSolver) that contains an intelligent BRAIN-ARM framework as the cognitive structure and operates the knowledge learning process in Store-Apply-Update steps inspired by two cognitive science theories. The BRAIN system stores three basic types of mathematical knowledge, and the ARM system applies them organically in answer reasoning process. After solving problems, the BRAIN updates its stored knowledge based on the ARM's feedback, with knowledge filters to eliminate redundancies and foster a more rational knowledge base. Our CogSolver carries out the above three steps iteratively, emulating a more human-like behavior. Furthermore, in order to overcome knowledge forgetting during the learning process, we extend CogSolver to CogSolver+ by incorporating an essential knowledge Recall mechanism, which is inspired by another prominent cognitive theory. We first discuss and fuse three crucial factors in simulating human memory replay. Then, we propose a influenced-based method with a theoretical guarantee of efficiency to consolidate the updated knowledge. Experiments on three math word problem benchmarks demonstrate the improvements of our CogSolver and CogSolver+ in answer reasoning and clearly illustrate how they acquire knowledge, leading to superior interpretability.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".