Bibliographic record
Abstract
Near the end of his life, Rogers Hornsby published an article in a men's magazine titled “You've Got to Cheat to Win in Baseball.” Hornsby wrote, “I've been in pro baseball since 1914 and I've cheated, or watched someone on my team cheat, in practically every game. You've got to cheat.” Hornsby's confession was not a complete surprise. Revered as one of baseball's supreme hitters, he was also reputedly meaner than Ty Cobb, and absolutely devoted to winning at all costs. But Hornsby was not the only confessed cheater among the all-time greats, some with more likable reputations. Hank Greenberg described a season when he was managed by an expert sign-stealer. “I loved that. I was the greatest hitter in the world when I knew what kind of pitch was coming up.” It is a truth universally acknowledged that in matters of supreme importance people cheat. Baseball matters; you might say the rest follows. Today, when we talk about cheating in baseball, we automatically think of steroids, but it's important to understand that cheating is bigger than steroids, and in fact, as Thomas Boswell writes, cheating is baseball's oldest profession. Even performance-enhancing drugs (PEDs) have a venerable history. Pud Galvin, one of the nineteenth century's greatest pitchers, downed an elixir of monkey testosterone before an 1889 game against Boston. Galvin won and drew a favorable comment from the Washington Post : “If there still be doubting Thomases who concede no virtue of the elixir, they are respectfully referred to Galvin's record in yesterday's Boston–Pittsburgh game. It is the best proof yet furnished of the value of the discovery.”
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.003 | 0.002 |
| Scholarly communication | 0.006 | 0.004 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.038 | 0.013 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".