Bibliographic record
Abstract
I am scouring the Internet with my six-year-old daughter for a map of Italy to include in her Geography project.The flag, also required for the project, was easy to find.But the map has to be clear, and to me, must pay due respect to the southernmost tip-Sicily.We find it, much to my daughter's delight, and I point out where her Nonni are from.The village is not listed, but the province-Siracusa-is.I then show her where Zio Carmelo, her great-uncle, lives-"way up north in Veneto.""In a town very close to Venice," I add.Her eyes light up with recognition; I know she is picturing the opening credits of the children's program Are We There Yet? in which the gondolier forcefully declares, "Buongiorno!"Then we get to work on the meat of the assignment-listing seven things about the country, some of which have been suggested by her grade one teacher in the instruction sheet: Surrounding Bodies of Water, Population, Capital City, and Languages Spoken.We add "Monuments," "Museums and Works of Art," and "Food."The last reminds me that I
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.009 | 0.002 |
| Scholarly communication | 0.005 | 0.005 |
| Open science | 0.001 | 0.004 |
| Research integrity | 0.002 | 0.004 |
| Insufficient payload (model declined to judge) | 0.101 | 0.076 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".