MétaCan
Menu
Back to cohort

Practical and Efficient Estimates of One’s Accuracy in Darts

2011· article· en· W1601797757 on OpenAlexaff
Sudipta Sadhukhan, Zhihui Liu, James A. Hanley

Bibliographic record

VenueJournal of the Royal Statistical Society Series A (Statistics in Society) · 2011
Typearticle
Languageen
FieldVeterinary
TopicAnimal Behavior and Welfare Studies
Canadian institutionsMcGill University
Fundersnot available
KeywordsComputer scienceAlgorithmArithmeticMathematicsCalculus (dental)StatisticsEconometricsOrthodonticsMedicine

Abstract

fetched live from OpenAlex

We thank Tibshirani et al. (2011) for their most interesting essay. In addition to its innovative use of a personalized heat map to show the optimal strategy for throwing darts, it provides an engaging example for teaching several statistical concepts and techniques, such as fast Fourier transforms, the EM algorithm, Monte Carlo integration, importance sampling and the Metropolis–Hastings algorithm. It is a delightful blend of the applied and the theoretical, the algebraic and the graphical. It also continues the tradition of statisticians’ fascination with the imagery of marksmen (Turner and Hanley, 2010). In her chapter on metaphor and reality of target practice, Klein (1997) wrote of ‘men reasoning on the likes of target practice’ and described how this imagery has pervaded the thinking and work of natural philosophers and statisticians. Klein showed a frequency curve, by Yule, for 1000 shots from an artillery gun in American target practice. Pearson used it in his 1894 lectures on evolution; he decomposed the frequency curve into two chance distributions centred slightly to the right and left of the target, gave reasons why this might occur and used it to illustrate the interplay between random variation and natural selection. He also used it in Pearson (1900) in one of the illustrations of his test of goodness of fit. Since the optimal aiming spot in darts—and thus the heat map provided by the on-line applet—depends strongly on one’s accuracy, much of Tibshirani et al. (2011) is devoted to the challenge of estimating the (co)variance parameter(s) that describes this accuracy. All the estimators rely on the data generated by throwing n darts, aiming each time at the centre of the board, i.e. the double bull’s-eye, and recording the result for each throw. Tibshirani et al. (2011) noted that they would lose considerable information by not measuring the actual locations where the darts land but considered this to be too time consuming and error prone. Instead, they chose the individual scores produced by the throws (the 44 possible scores are 0–22, 24–28, 30, 32–34, 36, 38–40, 42, 45, 48, 50, 51, 54, 57 and 60). Based on n = 100 throws by authors 1 and 2, assuming the simplest variance model (equal, uncorrelated vertical and horizontal Gaussian errors), their standard deviations were estimated to be σ^=64.6 and σ^=26.9 respectively (the applet gives σ^ to two decimal places). We write to provide a measure of the statistical precision of these estimates of accuracy (for example, we calculate that the 95% limits to accompany the reported point estimate 64.6 derived from 100 scores are approximately 56 and 75). More importantly, we show that more precise estimates of σ can often be achieved with the same number of throws (or the same precision with fewer throws) if we use a simpler yet more informative version of the result from each throw. Here we focus on the simplest variance model. The low information content of the scores with respect to σ is because many of them arise from throws that land at very different distances from the centre. For example, a score of 18 can arise from a throw that lands in one of four regions: double 9 (least accurate), outer single 18 (accurate), triple 6 (more accurate) or inner single 18 (most accurate). This ambiguity and loss of information are avoided if we simply record instead which of the seven ‘rings’ the throw lands in: 1, the double bulls-eye; 2, the single bull’s-eye; the rings formed by the 3, single bulls-eye and inner triple, 4, inner and outer triple, 5, outer triple and inner double, and 6, inner and outer double wires respectively; and 7, beyond the outer double wire (i.e. the throw misses the board), i.e. we need only to divide the dartboard into seven rings according to their distance to the centre. To quantify how much information is conserved if the raw location data are reduced to ‘ring’ data and ‘score’ data, we can measure the relative efficiency of these two latter methods of data recording. Since the log-likelihood is more symmetric in log (σ) than in σ, each panel in Fig. 1 shows the log-likelihood, but on a for σ. Log-likelihoods, and amounts of information regarding log(σ), using (expected) results of n = 50 throws, if we record the actual locations (,I log (σ) = 4n), or reduce them to the seven possible rings (–––), or the 44 possible scores (……) (log-likelihoods, relocated to equal 0 at σ^MLE⁠, are plotted against σ, but with a log-scale for the horizontal axis, with corresponding amounts of information regarding log(σ); expected frequencies (scaled to sum to 1000) in the seven rings are shown in grey in the background; for low values of σ ((a)–(c)) the simpler ring data provide the same amount of information as the score data (the log-likelihoods overlap); for larger σ-values ((d)–(f)) they provide a greater amount of information): (a) locations 200, rings 128, scores 128; (b) locations 200, rings 134, scores 134; (c) locations 200, rings 54, scores 54; (d) locations 200, rings 108, scores 31; (e) locations 200, rings 171, scores 90; (f) locations 200, rings 169, scores 125 The log-likelihood function and the three amounts of (Fisher) information are based on a sample size n = 50, as suggested by the authors. The expected amount of information concerning log (σ) contained in the raw location values can be shown analytically to be 4n, or 200 in our example. We calculated the corresponding information for the competitors by using the expected (multinomial) frequencies. Fig.1 shows that the ring data are often much more (and never less) informative than the score data. This difference in information is greatest when the player is moderately accurate: as is seen in Fig. 1(d) and Fig. 1(e) we can obtain the same amount of information about log (σ) by using ring data on 26 (=50×90/171) throws or score data on 50 throws. This difference is least when the results from the two data recording systems overlap considerably, i.e. if most of the throws are in or close to one of the two bulls-eye regions (Figs 1(a)–1(c), where curves shown with dotted and broken lines are virtually indistinguishable), or if a large percentage of throws fall outside the board (Fig. 1(f)). Fig.1 can be used to provide a confidence interval to accompany (for example) the reported σ^=64.6 based on Tibshirani’s 100 scores (see Fig. 1(e)). If this estimate had been based on the detailed locations for n = 100 throws, SE{log(σ^)} would have been approximately (1/4n)1/2=1/4001/2, the multiplicative margin of error for a 95% confidence interval would be approximately exp(1.96SE) = 1.1 and so the confidence interval for σ would be approximately from 64.6/1.1 to 64.6×1.1, or 59–71. However, since they were in fact based on scores, with an efficiency of only 90/200=0.45, SE{log(σ^)} is approximately {1/(4n×0.45)}1/2=(1/180)1/2, the multiplicative margin of error for a 95% confidence interval is approximately 1.16 and so the limits are approximately from 64.6/1.16 to 64.6×1.16, or 56–75. For σ-values in this range, the information content of the ring data is 2×171/400ths, or 85.5% that of the full location data. Others may wish to explore what additional data could be used to recover more of the information about the more complex variance structures that were considered by Tibshirani et al. (2011). Again, we salute Tibshirani et al. (2011) for their readable modern essay and for maintaining a statistical focus on marksmanship, yet using less dangerous missiles than those studied by statisticians of centuries past.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.046
metaresearch head score (Gemma)0.334
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.046
Threshold uncertainty score0.245

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0460.334
Meta-epidemiology (narrow)0.0030.002
Meta-epidemiology (broad)0.0040.003
Bibliometrics0.0060.004
Science and technology studies0.0020.009
Scholarly communication0.0060.014
Open science0.0060.012
Research integrity0.0040.011
Insufficient payload (model declined to judge)0.0100.003

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.085
GPT teacher head0.356
Teacher spread0.271 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2011
Admission routes1
Has abstractyes

Explore more

Same venueJournal of the Royal Statistical Society Series A (Statistics in Society)Same topicAnimal Behavior and Welfare StudiesFrench-language works237,207