Bibliographic record
Abstract
WHAT AND WHENA holistic, problem-first statistics course is one built by starting with a class of problem and using that to justify what methods to cover, rather than starting with a method and using that to justify what problems to apply to it.Such courses could be based on topics of broad and ongoing public concern.The topics of these proposed courses are data ethics and safety, political polling and demography, sports analytics, gambling and games of chance, and clinical trials.Through assignments, problem-first courses like these would allow statistics students to build a portfolio of work relevant to a target industry.The primary drawback of having courses that draw from disparate methods is that they work poorly as pre-requisites.This drawback is fixed by building problem-first classes for senior undergrads that have most of their pre-requisites already.Senior undergrads are also the group that can most benefit from industry-specific knowledge and the chance to build an industry-specific portfolio.Seniors are also the best prepared for messy, open-ended aspects of statistical work like data cleaning, report writing, and visualizations, all of which would benefit with a clear, central problem. GAMBLING IMPLEMENTATIONIn this poster, I propose an implementation of a gambling course.A problem like pricing sporting event wagers (e.g., +225 or 3.25 for a given soccer team to win given match) could be used as a motivation for a survey of logistic regression, ordinal logistic regression, and Monte Carlo simulations, each as attacks on the problem.The final deliverable could be a model that assigns prices to some future matches.Similarly, a bluffing game like Texas Hold'Em could be used as motivation to explore decision trees, game theory, and conditional probability.Identifying anomalous player behaviour opens to the door to hypothesis testing and distribution theory.The chi-squared test for independence is valuable in finding blackjack card counters, and the non-central chi-squared distribution can describe the behaviour of loaded dice.There would be a large introductory section on identifying, preventing, and treating compulsive gambling to drive home that this is not an endorsement of gambling.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".