What's the best place for an AI conference, Vancouver or ______: Why\n completing comparative questions is difficult
Bibliographic record
Abstract
Although large neural language models (LMs) like BERT can be finetuned to\nyield state-of-the-art results on many NLP tasks, it is often unclear what\nthese models actually learn. Here we study using such LMs to fill in entities\nin human-authored comparative questions, like ``Which country is older, India\nor ______?'' -- i.e., we study the ability of neural LMs to ask (not answer)\nreasonable questions. We show that accuracy in this fill-in-the-blank task is\nwell-correlated with human judgements of whether a question is reasonable, and\nthat these models can be trained to achieve nearly human-level performance in\ncompleting comparative questions in three different subdomains. However,\nanalysis shows that what they learn fails to model any sort of broad notion of\nwhich entities are semantically comparable or similar -- instead the trained\nmodels are very domain-specific, and performance is highly correlated with\nco-occurrences between specific entities observed in the training set. This is\ntrue both for models that are pretrained on general text corpora, as well as\nmodels trained on a large corpus of comparison questions. Our study thus\nreinforces recent results on the difficulty of making claims about a deep\nmodel's world knowledge or linguistic competence based on performance on\nspecific benchmark problems. We make our evaluation datasets publicly available\nto foster future research on complex understanding and reasoning in such models\nat standards of human interaction.\n
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.005 | 0.001 |
| Scholarly communication | 0.006 | 0.004 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.109 | 0.032 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".