ARTIFICIAL INTELLIGENCE A Mechanical Solution of Schubert's Steamroller by Many-Sorted Resolution
Bibliographic record
Abstract
We demonstrate the advantage of using a many-sorted resolution calculus by a mechanical solulion of automated a challenge theorem problem. provers This before. problem Our known solution as clearly 'Schubert's demonstrates Steamroller ' the power had of been a many-so unsolvedt ed bY resolution calculus. The proposed method is applicable to all resolution-based inference systems. In 1978, problem I. Schubert's Problem Schubert of the University of Alberta set up the following challenge Wolves, foxes, birds, caterpillars, and snails are animals, and there are some of each of them. Also there are some grains, and grains are plants. Every animal either likes to eat all plants or all animals much smaller than itself that like to eat some plants. Gaterpillars and snails are much smaller than birds, which are much: ' smaller than foxes, which in turn are much smaller than wolves. Wolves do not like to eat foxes or grains, while birds like to eat caterpillars but not snails. Caterpillars and snails like to eat some plants. Therefore there is an animal that likes to eat a grain-eating animal. This problem became well known since in spite of its apparent simplicity it turned out to be too hard for existing theorem provers because the search space is just too big. Using the following predicates as abbreviations: A(x): x is an animal, W(x): x is a wolf, F(x): x is a fox, B(x): x is a bird, C(x): x is a caterpillar, S(x): x is a snail, G(x): x is a grain, P(x): x is a plant, M(xy): x is much smaller than y, E(xy): x likes to eat y,
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.006 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.002 | 0.004 |
| Scholarly communication | 0.003 | 0.005 |
| Open science | 0.002 | 0.004 |
| Research integrity | 0.001 | 0.004 |
| Insufficient payload (model declined to judge) | 0.013 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".