01
The affiliation gap
508,744 of 793,883 works (64%) in the metaresearch topic space have no raw affiliation strings in OpenAlex.
affiliation_gap · computed 2026-07-13T11:54:42Z
02
The topic route
Of 4,516 OpenAlex topics, 0 name metaresearch as a field. The 11 topics that do carry metaresearch content are scattered across 7 different OpenAlex fields.
topics · computed 2026-07-13T11:54:42Z
03
Polysemy defeats the lexicon
The single term reproducibility retrieves 43,392 Canadian works, of which only 0.8% fall in the metaresearch topic space. Keyword retrieval cannot separate the metaresearch sense from the everyday one.
polysemy · computed 2026-07-13T11:54:44Z
04
The language gap
French is 2.7% (395/14,873) of Canadian metaresearch in OpenAlex. A dedicated French lexicon finds 168 Canadian works, against 8,026 in English.
language_gap · computed 2026-07-13T11:54:43Z
05
Érudit is invisible to OpenAlex
Erudit matches 0 sources in OpenAlex, but its OAI-PMH endpoint is live and exposes 379 harvestable sets.
erudit · computed 2026-07-13T11:54:43Z
06
Capture-recapture is void here
Naive two-route capture-recapture estimates 467,541 Canadian metaresearch works, implying Canada produces 59% of the world's metaresearch against an observed 1.9%. The estimator is void here; we cut it rather than dress it up as a lower bound.
capture_recapture_fails · computed 2026-07-13T11:54:44Z
07
The three-model screen
Three frontier models (Opus 4.8, GPT-5.6 high, Grok 4.5) screened the same 5,600 works, drawn from the real 4.3M frame under a design whose seven strata PARTITION it (an earlier five-stratum design could not reach 12.9% of the frame at all; D22). Design-weighted base rates span 2.54% to 3.81% (1.5x). But the RATE is not the finding, the SETS are: of the 274 works ANY model called metaresearch, only 104 (38%) were called metaresearch by ALL THREE, and 117 (43%) rest on a SINGLE model's opinion; pairwise Jaccard on the in-scope sets is about 50%. THE FIELD'S BOUNDARY IS NOT A LINE THE MODELS SHARE; IT IS A REGION THEY EACH CUT DIFFERENTLY, and that result is STABLE across n = 1,000, 2,000 and 5,600 (unanimity 37%, 37%, 38%). A SECOND, PRETTIER CLAIM DID NOT SURVIVE: at n = 2,000 the models agreed markedly more on 'is this about research at all' (1.43x here) than on 'is it in scope' (1.51x here), and this project said so in capitals; at n = 5,600 the two spreads are within noise (ratio 1.06) and the claim is WITHDRAWN (D23). The largest tier confusion is OUT-vs-T2, every time, at every sample size: the adjacent traditions the inclusiveness criterion exists to protect. The deliverable is not a base rate. It is the disagreement dossier, the 274 works that mark the empirical boundary, each carrying all three models' stated reasons, and the criteria that have to be written against them.
three_model_screen · computed 2026-07-15T22:59:18Z
08
Swap the screener, move the answer
Swap which model is called 'the screener' and the base rate moves from 1.06% to 2.37%: a 2.2x spread, from 45,397 to 101,776 works in the frame. The two screeners agree on in/out for 98.4% of the frame (design-weighted), but that figure is dominated by the settled rejects: agreement falls to 95% inside the contested boundary. THE SCREENER-SWAP RANGE, NOT THE BINOMIAL CI ON EITHER MODEL ALONE, IS THE HONEST UNCERTAINTY ON THE FIELD'S SIZE.
agreement · computed 2026-07-15T22:52:08Z
09
The base rate
Screening 5,737 unfiltered Canadian works against the rubric puts metaresearch at 1.31% of Canadian research, implying ~56,206 works in the 4,299,418-work frame, sizing the field without a search strategy at all. The 95% CI on this screener's labels is 1.03-1.64%, but that is sampling error, NOT the uncertainty: swap the screener and the estimate lands outside it (finding 10). The machine-screener range in finding 10 is the honest uncertainty on field size, and only the human audit can narrow it.
base_rate · computed 2026-07-15T22:33:57Z
10
Base-rate robustness
The base rate's real bias is not recency but MISSING ABSTRACTS: 31.5% of the partition has none, and the screen finds 0.78% metaresearch there against 1.55% where an abstract exists (chi-square p = 0.023, robust to adjustment for year and language). A third of the frame is screened on its title alone. What this does NOT show, and an earlier draft wrongly claimed, is that the blindness is DIFFERENTIAL by tradition: the T2 no-abstract cell holds 4 works and the interaction is not significant (p = 0.141). That claim is withdrawn, as is 'Erudit's exact profile' (the stratum is 99% English and the works are NEWER, not older). The main effect is the finding.
- Source field:
partition - updated_date=2026-06-24
- Source field:
records_sent_to_screener - 6,202
- Source field:
records_silently_lost - 465
- Source field:
records_lost_pct - 7.5
- Source field:
pct_no_abstract_among_lost - 40.2
- Source field:
pct_no_abstract_among_labelled - 31.5
- Source field:
chisq_p_loss_bias - 0.00013
- Source field:
losses_are_biased - yes
- Source field:
works_2000_09 - 2,704
- Source field:
works_2020_25 - 1,195
- Source field:
old_to_new_ratio - 2.3
- Source field:
partition_skews_old - yes
- Source field:
base_rate_by_era_pct - 2000-09:
- 1.24
- 2010-19:
- 1.35
- 2020-25:
- 1.4
- Source field:
chisq_p_era - 0.906
- Source field:
era_events - 75
- Source field:
era_check_is_underpowered - yes
- Source field:
no_abstract_share_pct - 31.5
- Source field:
base_rate_no_abstract_pct - 0.78
- Source field:
base_rate_has_abstract_pct - 1.55
- Source field:
chisq_p_abstract - 0.023
- Source field:
abstract_effect_x - 2
- Source field:
t1_penalty_x - 1.4
- Source field:
t2_penalty_x - 3.6
- Source field:
t2_no_abstract_cell_count - 4
- Source field:
interaction_p - 0.141
- Source field:
differential_is_supported - no
- Source field:
differential_claim_withdrawn - yes
- Source field:
no_abstract_mean_year - 2013.7
- Source field:
has_abstract_mean_year - 2010.3
- Source field:
no_abstract_stratum_pct_english - 99
- Source field:
p_no_abstract_given_english_pct - 32
- Source field:
p_no_abstract_given_non_english_pct - 12.7
- Source field:
erudit_profile_claim_holds - no
- Source field:
residual_bias_direction - anti-conservative for coverage claims: the partition over-represents abstract-less works (31.5%), where the screen finds 2x less metaresearch, so 1.31% likely UNDER-states the frame's base rate, and the field is larger than the headline implies
- Source field:
supersedes - THREE retractions live here. (1) An earlier version tested ERA only, called the base rate robust, and published 3100/2900/1500 as percentages (a dplyr summarise() column-masking bug). (2) It then claimed the blindness is DIFFERENTIAL, T2 losing 3.6x against T1's 1.4x, and made that the proposal's whole answer to the inclusiveness criterion. The T2 no-abstract cell holds FOUR works and the interaction is not significant (p = 0.141). WITHDRAWN. (3) It claimed missing abstracts track 'older, non-English' records, 'Erudit's exact profile'. Backwards: they are NEWER, and the stratum is 99% ENGLISH. WITHDRAWN. See DEVIATIONS.md D4, D5, D6.
- Source field:
caveat - What survives is the MAIN EFFECT and only the main effect: a third of the frame is screened on its title alone and the screen finds half as much metaresearch there (p = 0.023, robust to adjustment for year and language). That is a real coverage problem and a reason to stratify the audit on abstract availability. It is NOT evidence of differential blindness by tradition, and this finding no longer says it is. Separately, the era check is underpowered (75 events, 3 strata) and cannot refute an era effect; it merely fails to detect one. And note D2: the harness silently dropped 465 records, non-randomly, on this very covariate.
base_rate_robustness · computed 2026-07-13T11:54:47Z
11
What the labels cannot tell us
Two limits on the machine labels, found by attacking the fixes. (A) The topic route's recall is 12% against screener A and 7% against screener B: finding 14's instrument (ii) scores filters against MACHINE labels, so it measures agreement with a machine, not accuracy, and finding 10 already showed that swings by a factor of two. The conclusion strengthens (the second screener thinks the route is WORSE) but 12% is not truth. (B) The pilot holds exactly 1 French in-scope work, so seeing 20 French positives needs ~1,620 coded French records against an audit budget of 1,000. FRENCH SENSITIVITY IS NOT ESTIMABLE IN AN OPENALEX-ONLY FRAME. That is an argument for the Erudit harvest, not against the French claim, but the pilot did not run that harvest, so the power is stated as a condition rather than a promise.
label_limits · computed 2026-07-13T11:54:49Z
12
Agent variance
Haiku, the model finding 13 budgets the entire screen on, lands near Sonnet's base rate (1.27% vs 1.06%) and agrees with it on 98.1% of the frame, but their in-scope SETS overlap 16% unweighted and 10% design-weighted (Sonnet-GPT: 54%/37%); of Sonnet's 58 positives Haiku agrees on 12. RATE AGREEMENT IS NOT SET AGREEMENT. Worse: agents of the SAME model on the SAME prompt disagree beyond chance in BOTH arms after conditioning on stratum (CMH p = 0.0056 and 0.015), with raw spreads 3.1x and 5.2x against 2.2x between models, and the agents' ordering FLIPS between arms. The noise inside one model is at least the size of the difference between models, and the pilot's own 40-agent screen is too underpowered to rule the same instability out (5 of 37 chunks found zero metaresearch; p = 0.113).
agent_variance · computed 2026-07-15T22:52:18Z
13
The abstract gap is structural
The screen's largest measured bias is the abstract gap: 23.3% of the frame (1,003,117 works) has NO ABSTRACT, and finding 11 showed the screen finds HALF as much metaresearch there. Cascading PubMed, Europe PMC and Crossref recovers 37.8% of a 500-work sample, cutting title-only exposure to ~14.5% of the frame. But I BUILT THE CASCADE AROUND CROSSREF as the discipline-agnostic rescue, and it recovered 2 abstracts against PubMed's 180: publishers do not deposit them, so THAT RESCUE DOES NOT EXIST (D15). The gap is therefore not a metadata failure a better index fixes; it is STRUCTURAL. Recovery is 91.2% for reviews against 6.2% for book chapters, 38.8% English against 15.4% French. So the tempting shortcut, 'just screen the works that have abstracts', is a SELECTION ON A COVARIATE THAT PREDICTS THE OUTCOME which would delete 61.6% of book chapters against 22.1% of articles, AND the works it deletes are exactly the works no cascade can rescue. Defensible only as a DECLARED exclusion with a measured cost, and the audit keeps a sampling floor in it.
abstract_cascade · computed 2026-07-13T11:55:17Z
14
A boolean over a four-state space
Joined to the Canadian frame by DOI, Retraction Watch records 143 works that OpenAlex does NOT flag as retracted, including 49 outright retractions. But the undercount is the smaller problem. 52 of these carry an EXPRESSION OF CONCERN, and OpenAlex HAS NO FIELD FOR ONE: `is_retracted` is a boolean over a state space with at least four values (retraction, expression of concern, correction, reinstatement), so it can express one and silently reports the rest as FALSE, which reads as 'fine'. Nor can a boolean carry WHY. This is finding 1's disease in a second schema: the canonical database cannot express the distinction the field turns on.
retraction_record · computed 2026-07-13T11:55:00Z
15
Topic-route recall
Scored against the rubric, the topic route retrieves 12% of Canadian metaresearch (95% CI 5.6-21.6%) at 60% precision: it misses 66 of 75. It fails because OpenAlex files a work by what it is about, and metaresearch about cardiology reads as cardiology: the field is invisible to topic retrieval precisely because it is about other fields.
topic_route_recall · computed 2026-07-13T11:54:47Z
16
Funder-route recall
The PRIMARY ESTIMAND leans on CA-FUND to rescue works whose affiliation is missing (finding 2: 64% have no affiliation string), and nothing had tested it. Tested against CIHR's own database of 44,190 funded projects, the result REFUTED THE HYPOTHESIS I WROTE BEFORE RUNNING IT: OpenAlex tags 178,133 frame works with CIHR, or 4.03 per grant, a plausible rate showing no under-tagging (DEVIATIONS.md D14). What the data DOES support needs no hypothesis of mine: 71.2% OF THE FRAME CARRIES NO FUNDER METADATA AT ALL, which is CA-FUND's hard ceiling, and 65.9% of Canadian-AFFILIATED works carry none either. Both clauses of the estimand rest on metadata that is mostly absent, which is why the frame is a union of four routes and why the audit must sample the works no route reached.
funder_route_recall · computed 2026-07-13T11:55:02Z
17
Canadian linkage
Affiliation finds 14,873 works; a further 3,964 are about Canada with no Canadian affiliation. NSERC has 9.2x SSHRC's linked works, so funder-based rules under-count the social sciences.
canadian_linkage · computed 2026-07-13T11:54:44Z
18
Trial linkage
The registry is the only REFERENCE STANDARD in this project not made of machine labels: ClinicalTrials.gov knows a Canadian trial happened independently of any pipeline, so it cannot be wrong in the pipeline's favour. Of 304 publications SPONSORS THEMSELVES reported as results of completed Canadian-located trials, the frame holds 160: a naive recall of 52.6%, which fell so close to the 44.5% THIS PROPOSAL OPENS WITH that it read as a replication. IT IS AN ARTIFACT, and running the disambiguation is the only thing that caught it: 137 of the 145 'misses' have NO CANADIAN AUTHOR (multi-site international trials with a Canadian SITE), and a frame of Canadian RESEARCH is CORRECT to exclude them. A trial with a Canadian site is not a publication with a Canadian author. Against the population the frame actually claims, recall is 95.2% (95% CI 90.8-97.9), and the real defect is 8 works OpenAlex holds WITH a Canadian author that the frame's own routes still missed. The frame is GOOD at this, the dramatic parallel was a coincidence between two different populations, and I had every incentive not to check. DEVIATIONS.md D16.
trial_linkage · computed 2026-07-13T11:55:20Z
19
Preprint coverage
The frame carries 156,086 preprints, and the tempting move was to assert that OpenAlex covers the preprint servers and skip the ingest. That is a COVERAGE CLAIM, and this project does not get to make one from inside the pipeline being claimed for: it is the topic route certifying its own recall (finding 12). Measured instead against the servers' OWN API, OpenAlex indexes 99.6% of the 705 bioRxiv and medRxiv preprints enumerated (3 missing). The claim survives measurement, so no separate preprint ingest is built.
preprint_coverage · computed 2026-07-13T11:55:18Z
20
The power of a human audit
The audit as first specified could not measure what it exists to measure. A simple random sample of the 4,243,096-record screened-out mass yields an expected 0.4 hits from the 600 records budgeted; seeing 20 would take 2,009 coder-hours against the 65 planned. Score-stratified oversampling rescues it only partly (the pilot shows 30 of 37 disputed works sit at the contested boundary), and it stays BLIND to the confident rejects. Which works those are is a MECHANISM, not a measured differential: the rubric says judge on the title alone with no abstract, and says T2 work may use none of the field's vocabulary, so a work with neither is rejected confidently and never sampled. (An earlier version cited a 3.6x figure here as if it were measured; it rested on four works and is withdrawn.) So recall is measured two other ways: against the 5,737 works that already carry rubric labels, and by KNOWN-ITEM recall on a venue reference set, because venue is an external criterion immune to the aboutness that defeats everything else.
audit_power · computed 2026-07-13T11:54:49Z
21
What screening costs
The full rubric over the 4,299,418-work frame with two screeners costs $6,567, not the $30,633 an earlier version of this script reported: the rubric is a system prompt sent once per CALL, and the pilot's own chunks batch 155 works per call, so it was overcharged 4.7x. That error was not cosmetic. It made me propose a cheap TRIAGE in front of the screen, which is a RETRIEVAL STEP in a project whose central finding is that retrieval destroys these maps. With the arithmetic right the triage is unnecessary: the full rubric over EVERY work in the frame, plus a second screener on a 20,000-record sample, costs $1,110 at the v1 instrument and $1,279 at the locked v3.1 instrument that will actually run. THE PREFILTER IS DELETED. And the compute is SELF-FUNDED: the call's CAD $4,000 covers, in its own words, 'travel, accommodation, and related participation costs', not inference and not coder wages (D30).
- Source field:
measured_from - pilot/screening/chunks/*.json (6-field), pilot/screening/haiku/p8_guided/chunk_*.json (8-field), docs/protocol/rubric.md; not asserted
- Source field:
chars_per_token_assumed - 4
- Source field:
tokens_per_work - 312
- Source field:
tokens_per_work_six_field_deviation - 256
- Source field:
payload_inflation_8_over_6 - 1.22
- Source field:
costed_at_the_rubric_payload_not_the_deviation - yes
- Source field:
tokens_rubric - 1,878
- Source field:
tokens_per_label - 37
- Source field:
batch_works_per_call - 155
- Source field:
frame_size - 4,299,418
- Source field:
grant_usd_approx - 2,900
- Source field:
naive_cost_charging_rubric_per_work_usd - 30,633
- Source field:
corrected_cost_two_screeners_usd - 6,567
- Source field:
cost_overstatement_x - 4.7
- Source field:
cost_full_rubric_whole_frame_cheap_usd - 1,094
- Source field:
cost_full_rubric_whole_frame_sonnet_usd - 3,283
- Source field:
cost_second_screener_on_sample_usd - 15
- Source field:
second_screener_n - 20,000
- Source field:
total_no_prefilter_usd - 1,110
- Source field:
tokens_rubric_v31 - 4,604
- Source field:
tokens_per_label_v31 - 49
- Source field:
cost_full_rubric_whole_frame_v31_usd - 1,261
- Source field:
cost_second_screener_v31_usd - 18
- Source field:
total_no_prefilter_v31_usd - 1,279
- Source field:
award_pays_for_compute - no
- Source field:
award_purpose_in_the_calls_words - travel, accommodation, and related participation costs
- Source field:
prefilter_needed - no
- Source field:
prefilter_design_total_usd - 1,004
- Source field:
prefilter_saving_usd - -106
- Source field:
prefilter_deleted_because - It is a RETRIEVAL STEP, in a project whose central finding is that retrieval destroys these maps (finding 12: the topic route finds 12% of the field). It existed only because the rubric was miscosted at once per WORK rather than once per CALL, overstating the alternative 5.1x. With the arithmetic right it saves almost nothing and costs the thesis. Deleted.
- Source field:
pilot_works_screened - 6,202
- Source field:
frame_to_pilot_ratio - 693
- Source field:
method_that_scales - batch inference, not agent fan-out
- Source field:
supersedes - an earlier version charged the rubric once per work, reported $24,379 for the full-frame two-screener design, called it 'eight times the grant', and used that to justify a cheap prefilter. The rubric is a system prompt sent once per CALL, and the pilot's own chunks batch 155 works per call. DEVIATIONS.md D7.
- Source field:
caveat - Token counts use a 4-chars-per-token approximation and list prices as of 2026-07; both will move, and the conclusion is robust to +/-25% in either. Screening the frame with a cheaper model than the pilot used makes the CHOICE OF MODEL more consequential, not less: finding 10 shows two screeners already imply base rates a factor of two apart. That is precisely why the second screener and the screener-swap range are reported, and why the human audit is the study.
screening_cost · computed 2026-07-13T12:23:27Z
22
OpenAlex is metered
The OpenAlex API is metered (1,000 credits per ~11h; $0.10 free tier). Enumerating the 4,299,418-work Canadian frame needs 21,498 cursor-paged calls: 10 days per pass on the free tier. An API-based pipeline at this scale is neither free nor reproducible; the pinned snapshot is both.
openalex_is_metered · computed 2026-07-13T11:54:45Z
23
active_learning
THE BATCH-OF-100 LOOP WORKS, AND THE 20 RANDOM ANCHORS IN IT DO NOT. Against a random-batch control at identical budget, active acquisition (50 max-teacher-disagreement + 30 max-uncertainty + 20 random) takes held-out AP from 0.0143 to 0.1391 while the control reaches 0.0502: 2.8x, winning 18 of 20 rounds, with held-out positive-set churn falling to 0.1111. At a ~1% base rate a random batch of 100 holds one positive, and the control curve is what that buys. BUT the anchors, which were supposed to keep an unbiased evaluation stream alive, hold 3 positives after 400 draws: prevalence 0.0057 with a 95% interval 0.01594 wide, WIDER THAN THE QUANTITY ITSELF. They are demoted to drift sentinels. THE LOOP LEARNS; THE PREREGISTERED HUMAN AUDIT MEASURES. And AP here is agreement with the majority TEACHER, so the whole curve is imitation, not accuracy. The first version of this script said the loop DEGRADED, because it measured AP on the shrinking pool the algorithm was itself editing; the churn sat at exactly 1.000 for twenty rounds, which is a constant, not a measurement (D33).
active_learning · computed 2026-07-15T22:52:19Z
24
adjudication
AN INDEPENDENT BLINDED JUDGE SAYS THE RUBRIC IS SILENT ON 89% OF THE WORKS THE MODELS FOUGHT OVER. Fable 5, which is not one of the three screened arms, adjudicated all 179 contested works seeing three screener opinions as A/B/C in random order with no model names, so it was asked which reading of the RUBRIC is right rather than which model to trust. It shows no arm preference (spread 1.16x, chi-square p = 0.493), so it is not a fourth vote for one of the screeners, and it overruled ALL THREE on 6 works. Its verdict on the instrument: the rubric is SILENT on 159 of 179 contested works, and only 17 splits are a screener misapplying a rule that actually exists. THE FIELD'S BOUNDARY IS BEING SET BY THE SCREENER, NOT BY THE INSTRUMENT. The seams that cost the most agreement are methods_dev_vs_study (39 works), lis_sts_asymmetry (26), and history_of_science (19). This is what converts adjudication from a tiebreak into a measurement: not 179 answers, but a frequency-weighted census of which sentences the rubric is missing and exactly how many works each one costs. The judge's TIERS are used for nothing and produce no estimate: a model cannot certify a model, and that argument applies to the judge too.
adjudication · computed 2026-07-13T11:55:28Z
25
canadian_linkage_misnames_itself
BOTH CLAUSES OF 'CANADIAN' MEASURE SOMETHING OTHER THAN WHAT THEY ARE NAMED. This project has spent all its effort on what retrieval MISSES; this is the first hard evidence about what it wrongly ADMITS. (1) CA-AFF, the PRIMARY estimand's main clause: 17,466 works enter the frame on a Canadian 'institution' that does not exist. OpenAlex hands 'Discovery Air (Canada)' to a Brazilian linguistics paper, 'Musee de la Civilisation' to a Brazilian food-science paper, and 'Impact', which is not an institution at all but a parse failure with an institution id, to a Spanish COVID essay. The CA-AFF route also carries thousands of works in Latvian and Indonesian. That count is a LOWER BOUND: the artifact strings tested are only the ones screening agents noticed BY EYE while doing something else, and nobody has swept the institution vocabulary, so the true precision is UNKNOWN rather than merely unmeasured. (2) ABOUT-CA, the SECONDARY estimand: the retrieval route means 'Canada appears in the text' and the rubric field means 'the Canadian research SYSTEM is a substantive object of study', and in the strata built entirely on the former, the latter fires on 1.8% to 3.5% of works. Worse, `about_ca` cannot be true unless the work is ALREADY about research, so it is nearly collinear with tier: only 39 works in 5,600 are both in-scope AND about the Canadian research system. THE SECONDARY ESTIMAND CANNOT BE ESTIMATED FROM THE STRATUM DESIGNED FOR IT. Four screening agents found this independently and all four proposed the same repair: split the field into `about_ca_system` and `about_ca_topic`, because one boolean cannot separate 'Canadian data, universal claim' from 'a claim about Canada'.
canadian_linkage_misnames_itself · computed 2026-07-13T11:55:27Z
26
classifier
THE CLASSIFIER MAY NOT LABEL, AND THE NUMBERS SAY WHY. Design-weighted, out-of-fold, venue-grouped: the teacher heads reach AP 0.17 to 0.135 at AUC ~0.8, with a 4.68x top-decile lift that is real and is lift toward MACHINE labels. THE ESTIMAND MOVES THE SCORE 2.7x (unanimous 0.0783, majority 0.1364, any 0.2093), so there is no 'the' target and choosing one silently settles the question the audit exists to answer. AND THE UNANIMOUS CORE IS THE HARDEST TO LEARN, not the easiest: if the boundary were a shared line with noise around it, the agreed core would be the separable part, and it is not. The between-head spread predicts where the teachers split at 3.74x the base rate: real, weak, and reported as an efficiency gain rather than a map. Finally, the abstract is worth MORE THAN ANY MODELING CHOICE (+0.1145 AP, 84% relative), and the frame cannot serve it: works.csv stores has_abstract, a boolean, and no text. The DEPLOYABLE model is the WORSE one, the build now FAILS if a model is trained on a field inference cannot supply, and the gap is a priced budget decision rather than an assumption.
classifier · computed 2026-07-15T22:32:09Z
27
distillation_ceiling
A CLASSIFIER TRAINED ON LLM LABELS CANNOT EVEN REPRODUCE ITS OWN TEACHER, LET ALONE ADJUDICATE THREE. Students distilled from each model match their own teacher at Jaccard 0.17, while the three teachers match EACH OTHER at 0.5 to 0.56: the classifier is a worse approximation of Opus than Grok is. Distilling from Opus also moves you AWAY from Grok (-0.35), so distillation does not bridge the models' disagreement, it degrades away from all of them. THE SELF-TRAINING LOOP THEN REFUTED MY OWN PREDICTION. I argued at length that it would shrink the positive class toward the confident anglophone core; run on 20,000 works the models never saw, it did not: positive rate drift +0.00 points, French share drift +0.00 points, because class weighting prevents the collapse, exactly as the adversarial review warned before the run ('a serious risk, NOT A THEOREM'). What the loop DID do is reach a fixed point immediately and recycle its own labels (224 -> 424 training positives, then nothing): it CONVERGED WITHOUT LEARNING, and a fixed point looks identical from the inside whether its labels are right or wrong. What the classifier IS worth is the stratifier: 48.2% of the in-scope works sit in its top decile against a blind baseline of 10%, a 4.8x lift in in-scope works found per unit of human coding effort, and design-weighted estimates stay unbiased however noisy it is. THE MACHINE DECIDES WHERE THE HUMANS LOOK. IT NEVER SUPPLIES A LABEL.
distillation_ceiling · computed 2026-07-15T22:58:15Z
28
gemma_gate
THE FREE MODEL FAILED THE GATE, AND THE GATE WAS AIMED AT THE WRONG BAR. Gemma-4-31B screened 700 works under rubric v3.1 (100% schema-clean) and scored Jaccard 0.44 on metaresearch and 0.515 on any-category against the frontier union, under the prespecified 0.60 bar, SO IT DOES NOT LABEL THE FRAME: the rule stands because rules that move after the numbers exist are not rules. But the frontier models agree with each other at only 0.384 and 0.36 on the same works, because every loop batch is selected FOR disagreement: the gate demanded of a free model what the frontier models do not achieve among themselves there, and Gemma beat the frontier's own mutual agreement on both targets. Test 2 is prespecified on the random_baseline stratum with the rationale as the bar: J(gemma, union) >= J(opus, gpt).
gemma_gate · computed 2026-07-14T00:08:32Z
29
instrument_contradicts_itself
THE LOCKED INSTRUMENT CONTRADICTS ITSELF, AND NOTHING CAUGHT IT. The rubric and the output schema name TWO DIFFERENT controlled vocabularies for the same field, `genre`, overlapping on 2 values (empirical and other); 4 terms exist only in the rubric and 7 only in the schema. Every screener was handed both and told to obey both, and each invented its own reconciliation: GPT-5.6 and Grok followed the rubric and are therefore 26.9% and 20.9% ILLEGAL against the schema, while Opus drew from both lists at once and emitted 13 distinct values. The validator checked `tier` and never checked `genre`, so 16,800 labels passed every check that ran. It was found by an agent mentioning it in one clause of a report about something else. Nothing crashed; the variable simply meant a different thing in each arm, and the variance would have been attributed to the models. The labels are NOT repaired, because harmonizing the arms after seeing them would destroy the only evidence that they diverged; the genre field is reported as unusable and the fix belongs in the instrument, at a version boundary.
instrument_contradicts_itself · computed 2026-07-15T22:32:11Z
30
the_frame
The frame is BUILT, not estimated. All 482 partitions of a pinned OpenAlex snapshot were streamed and filtered, and the Canadian frame holds 4,299,418 works, each exactly once. For most of this project's life 'the frame' was 3,507,205, an EXTRAPOLATION from a single partition, hardcoded in six scripts; the estimate was 18% LOW, and every cost, audit-power and field-size figure computed against it has moved. The route that matters: 1,565,226 works (36.4% of the frame) are INVISIBLE TO AFFILIATION ALONE. A frame built on Canadian affiliation would hold 2,734,192 works and would never see them. That is finding 2 at frame scale, and it is why the frame is a union of four routes and why every record carries the provenance of the route that admitted it.
the_frame · computed 2026-07-13T11:55:21Z
31
v1_to_v2
WRITING THE MISSING SENTENCES RESOLVED 54% OF THE DISAGREEMENTS THEY WERE WRITTEN AGAINST. The 179 works three frontier models split on under rubric v1 were re-screened, by the same three models, under a v2 whose rules were written against a published, frequency-weighted census of WHICH SENTENCES THE RUBRIC WAS MISSING. 96 of 179 are now unanimous, and the pairwise Jaccard of the in-scope sets moves from 0.19 to 0.36. This is a measurement almost nobody makes, and it is available ONLY because the ambiguity was enumerated, adjudicated and PUBLISHED BEFORE it was resolved: the census went out first, so v2 could not be quietly tuned until the numbers looked good. SCOPE, stated rather than blurred: these works were SELECTED FOR DISAGREEMENT, so this says the rules resolve the cases they were written for; it does NOT say frame-wide agreement rose, because regression to the mean alone would move a set selected this way. The unbiased test is a full re-screen of all 5,600 under v2, prespecified, and it is what the funded work runs.
v1_to_v2 · computed 2026-07-13T11:55:28Z
32
zero_probability_region
THE SAMPLING DESIGN COULD NOT REACH 12.9% OF THE FRAME, AND NO WEIGHT CAN FIX THAT. A stratified design rests on one identity, sum(N_h) = N. It was never checked. 549,370 works (12.9% of the sampling frame) had an inclusion probability of EXACTLY ZERO: 366,856 that no predicate claimed, because aff_core excluded everything ABOUT Canada while about_only excluded everything AFFILIATED with Canada, so a work that was both fell between them; and 182,514 more whose predicate evaluated to SQL NULL, which `WHERE p` and `WHERE NOT p` BOTH decline to select, so they were in no stratum and were not even orphans. Nothing threw. Every stratum returned exactly the n it asked for, because a stratum cannot know about the works it was never asked about, and the design drew a clean textbook probability sample OF 87% OF THE FRAME while every number computed from it said 'the frame'. A zero-probability work is not underweighted, it is UNREACHABLE, and design weights are the guarantee this project leans on hardest. THE CELL IT DELETED WAS THE SECONDARY ESTIMAND'S: 328,912 Canadian-affiliated works about Canada, in a study whose secondary estimand is 'metaresearch about the Canadian research system'. Repaired by adding the exact NULL-safe complement as two strata; the 7 strata now partition the frame by construction (4,255,410 = 4,255,410), asserted on every build. It was found by an adversarial model adding up five numbers I handed it. The defect made the sample TIDIER and the variance SMALLER, which is why nothing about it felt wrong: the errors that survive are the ones that flatter you.
zero_probability_region · computed 2026-07-13T11:55:25Z