Bibliographic record
Abstract
As a licensing exam, the purpose of the bar exam is consumer protection–-ensuring that new lawyers have the minimum competencies required to practice law effectively. As critics point out, however, the exam, and particularly the multiple-choice question portion of the exam, has significant flaws because it assesses legal knowledge and analysis in an artificial and unrealistic context, and the closed-book format rewards the ability to memorize thousands of legal rules, a skill unrelated to law practice. This essay discusses how to improve the exam by changing its multiple-choice content and format. We use two law licensing exams to illustrate how bar examiners could utilize an open-book format and develop multiple-choice questions that assess a candidate’s ability to engage in legal reasoning and analysis without demanding unproductive memorization of so many detailed rules of law. The first example, the case file approach, is drawn from a 1983 California “Performance Test” in which test-takers received a case file and a series of multiple-choice questions testing the candidates’ ability to read, understand, and use cases to support their legal positions. The second example discusses the current licensing exam administered by The Law Society of Upper Canada (LSUC), an open-book multiple-choice exam that tests the use of doctrinal knowledge in the context of law practice. These two licensing exams demonstrate how we could re-structure the bar exam’s multiple-choice questions to measure legal analysis and reasoning skills as lawyers use those skills to represent clients. They also demonstrate that we can do a better job of testing some aspects of minimum competence, while still using a multiple-choice exam format.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".