Covering: Mutable Characteristics and Perceptions of (Masculine) Voice in the U.S. Supreme Court
Bibliographic record
Abstract
The emphasis on “fit” as a hiring criterion has raised the spectrum of a new form of subtle discrimination (Yoshino 1998; Bertrand and Duflo 2016). Under complete markets, correlations between employee characteristics and outcomes persist only if there exists animus for the marginal employer (Becker 1957), but who is the marginal employer for mutable characteristics? Using data on 1,901 U.S. Supreme Court oral arguments between 1998 and 2012, we document that voice-based snap judgments based on lawyers’ identical introductory sentences, “Mr. Chief Justice, (and) may it please the Court?”, predict court outcomes. The connection between vocal characteristics and court outcomes is specific only to perceptions of masculinity and not other characteristics, even when judgment is based on less than three seconds of exposure to a lawyer’s speech sample. Consistent with employers irrationally favoring lawyers with masculine voices, perceived masculinity is negatively correlated with winning and the negative correlation is larger in more masculine-sounding industries. The first lawyer to speak is the main driver. Among these petitioners, males below median in masculinity are 7 percentage points more likely to win in the Supreme Court. Justices appointed by Democrats, but not Republicans, vote for lessmasculine men. Female lawyers are also coached to be more masculine and women’s perceived femininity predict court outcomes. Republicans, more than Democrats, vote for more feminine-sounding females. A de-biasing strategy is tested and shown to reduce evaluators’ tendency to perceive masculine voices as more likely to win. Perceived masculinity explains 3-10% additional variance compared to the current best prediction model of Supreme Court votes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".