“You Keep Using That Word”: Why Privacy Doesn’t Mean What Lawyers Think
Bibliographic record
Abstract
This article explores how the need to define privacy has impeded our ability to protect it in law. The meaning of “privacy” is notoriously hard to pin down. This article contends that the problem is not with the word “privacy,” but with the act of trying to pin it down. The problem lies with the act of definition itself and is particularly acute when the words in question have deep-seated and longstanding common-language meanings, such as liberty, freedom, dignity, and certainly privacy. If one wishes to determine what words like these actually mean to people, definition is the wrong tool to use. The exact wrong way to go about understanding privacy is by supplying one’s own definition; that is unscientific. Since words in a living language mean many things (e.g., what does “cool” mean?), the act of definition reduces the multiple meanings of the defined word to a specified meaning. Each increase in precision comes with a corresponding separation from some set of meanings that would have applied to the living, undefined version of the word. The resulting defined word may be more precise but is often crippled, isolated, and bereft of the connections and connotations that made it part of a rich and living language. Like Procrustes, who strapped his victims to a bed and then either lopped off their feet if they stuck out or stretched the person on a rack if they were too short, lawyers are specifically trained to stretch and cut words. Tools of definition are badly suited to determine what people mean when they say “privacy.” For example, the actual meaning of “privacy” might better be explored through the tools of linguistics or cultural anthropology than through the tool of legal definition. This article therefore recommends that lawyers should set aside the flawed tool of definition and pick up the tool of analogy when they ask what words like privacy mean. This article asks why privacy has been uniquely pressed by concerns about supposed imprecision. For example, we do not stop our search for “security” because of a supposed lack of definition of the word. If privacy must have a definition to be operationalized, it will remain be conveniently narrow. moribund. And if privacy requires narrowing to be operationalized, any operationalization will be conveniently narrow. “You keep using that word. I do not think it means what you think it means.”
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.025 | 0.072 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.021 | 0.054 |
| Scholarly communication | 0.019 | 0.040 |
| Open science | 0.002 | 0.009 |
| Research integrity | 0.015 | 0.034 |
| Insufficient payload (model declined to judge) | 0.009 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".