Correlation perception in scatterplots is invariant to dot size
Bibliographic record
Abstract
To determine the extent to which the perception of Pearson correlation r in a scatterplot depends on its appearance, we examined the effect of dot size on discrimination and magnitude estimation. Scatterplots were formed of 48 solid black dots on a white background, with axes 6.5 cm by 6.5 cm, and distribution standard deviations 1.3 cm. Observers (N=18) were tested via a within-observer design involving five conditions: dot diameters of 1 mm, 3 mm, 5 mm, 8 mm, and a mix of these sizes. Viewing distance was 67 cm. The methodology was that of Rensink and Baldridge (2010). In the discrimination task, observers were asked to select the plot with the higher perceived correlation; the just noticeable difference (JND) was measured at three base correlations (0.3, 0.6, 0.9). In the magnitude estimation task, observers adjusted a test plot until its perceived correlation was midway between those of two reference plots. The discrimination task was sandwiched between two sets of estimation tasks. All conditions were counterbalanced by base correlations and dot size, using a Latin square. The resulting JNDs were slightly higher than those reported by Rensink and Baldridge (2010) and Rensink (2017), but were still strongly linear functions of correlation (R^2=0.97); the Fechner assumption of equal perceived difference for each JND was also supported. Importantly, neither discrimination nor estimation were significantly affected by dot size. This further supports the proposal (Rensink, 2017) that perceived correlation in scatterplots is based not on the physical appearance of the scatterplot, but on a more abstract quantity, such as the shape of the probability density function derived from the locations of the dots in the image.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.031 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".