Consumer Arbitrations with The American Arbitration Association 2009 to Present
Bibliographic record
Abstract
This website reposts data previously posted on the American Arbitration Association's website: https://www.adr.org/consumer. The AAA describes the data on their website: "The AAA maintains an online Consumer Arbitration Statistics report based on consumer cases filed with the AAA for at least the last five years. This report is made available pursuant to state statutes such as the California Code of Civil Procedure §1281.96 and Maryland Commercial Law §§ 14-3901 to 3905 and updated quarterly, as required by law." In practice, each time the AAA adds the latest quarter of data, it takes down the earliest quarter. This website aids researchers by retaining the data each quarter in exactly the format in which it was originally posted. We are retaining and posting this data because we have found it useful in trying to understand the effect of mandates for consumers to arbitrate. Caveats are in order. A first limitation of the data is the absence of access to the underlying materials, which are held privately. As the AAA explains, it does not independently verify what arbitrators report to it. A second problem is that coding errors can occur at both individual and aggregate levels. For example, when researching consumer arbitration between 2015 and 2016, we identified sixty-two cases in the set that were described as seeking the same amount ($607,525.40) and in which each consumer was listed as having received the same award ($585.71). AAA research staff responded to our inquiries, identified a computer coding error affecting these cases as well as other cases, and posted corrected data. But no red flags told other researchers that the data had been corrected. Thus, a vivid example of a potential error may be found through culling thousands of entries and then seeking clarification, but the general public has no systematic method of checking the accuracy of the data posted by AAA. Several authors have used this data, see, e.g., David Horton & Andrea Cann Chandrasekher, After the Revolution: An Empirical Study of Consumer Arbitration, 104 Geo. L.J. 57 (2015); Judith Resnik, Diffusing Disputes: The Public in the Private of Arbitration, the Private in Courts, and the Erasure of Rights, 124 Yale L.J. 2804 (2015). Another site that has usable AAA data is Level Playing Field (http://levelplayingfield.io), and there could be other sites as well. If you have ideas about or corrections to this data please email ylsarbitrationdataarchive@gmail.com. Thanks are due to the Yale Law Library for their help conceiving and constructing this website.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.036 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.004 | 0.007 |
| Science and technology studies | 0.005 | 0.001 |
| Scholarly communication | 0.009 | 0.005 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.003 | 0.005 |
| Insufficient payload (model declined to judge) | 0.297 | 0.166 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".