MétaCan
Menu
Findings

What the pilot measured.

All 32 findings, rendered directly from pilot/results/findings.json: the file the pilot scripts write. No number on this page was typed by a human, which is the only way to guarantee the site and the analysis cannot drift apart.

the same file over the API →

01

The affiliation gap

508,744 of 793,883 works (64%) in the metaresearch topic space have no raw affiliation strings in OpenAlex.

Source field:topic_space_total
793,883
Source field:with_raw_affiliation
285,139
Source field:without_raw_affiliation
508,744
Source field:pct_without
64.1
affiliation_gap · computed 2026-07-13T11:54:42Z
02

The topic route

Of 4,516 OpenAlex topics, 0 name metaresearch as a field. The 11 topics that do carry metaresearch content are scattered across 7 different OpenAlex fields.

Source field:n_topics_in_taxonomy
4,516
Source field:n_topics_naming_field
0
Source field:topics_naming_field
    Source field:n_candidate_topics
    11
    Source field:n_fields_spanned
    7
    Source field:fields_spanned
    • Arts and Humanities
    • Computer Science
    • Decision Sciences
    • Mathematics
    • Medicine
    • Psychology
    • Social Sciences
    Source field:candidate_topic_ids
    • T10102
    • T13607
    • T13516
    • T11937
    • T10206
    • T10582
    • T10267
    • T10778
    • T13558
    • T13284
    • T11875
    topics · computed 2026-07-13T11:54:42Z
    03

    Polysemy defeats the lexicon

    The single term reproducibility retrieves 43,392 Canadian works, of which only 0.8% fall in the metaresearch topic space. Keyword retrieval cannot separate the metaresearch sense from the everyday one.

    Source field:hits_alone
    reproducibility:
    43,392
    "peer review":
    21,234
    "open access":
    6,609
    "open science":
    1,925
    Source field:hits_alone_and_on_topic
    reproducibility:
    362
    "peer review":
    795
    "open access":
    634
    "open science":
    325
    Source field:topic_space_precision_pct
    reproducibility:
    0.8
    "peer review":
    3.7
    "open access":
    9.6
    "open science":
    16.9
    Source field:disciplined_lexicon_hits
    8,026
    Source field:worst_term
    reproducibility
    polysemy · computed 2026-07-13T11:54:44Z
    04

    The language gap

    French is 2.7% (395/14,873) of Canadian metaresearch in OpenAlex. A dedicated French lexicon finds 168 Canadian works, against 8,026 in English.

    Source field:canadian_topic_works
    14,873
    Source field:n_english
    14,028
    Source field:n_french
    395
    Source field:pct_french
    2.7
    Source field:en_lexicon_canadian_hits
    8,026
    Source field:fr_lexicon_canadian_hits
    168
    Source field:fr_lexicon_world_hits
    4,682
    Source field:canada_share_of_world_french
    3.6
    language_gap · computed 2026-07-13T11:54:43Z
    05

    Érudit is invisible to OpenAlex

    Erudit matches 0 sources in OpenAlex, but its OAI-PMH endpoint is live and exposes 379 harvestable sets.

    Source field:openalex_sources_matching_erudit
    0
    Source field:oai_endpoint
    https://oai.erudit.org/oai/request
    Source field:oai_repository_name
    Erudit
    Source field:oai_earliest_datestamp
    2011-06-03
    Source field:oai_harvestable_sets
    379
    erudit · computed 2026-07-13T11:54:43Z
    06

    Capture-recapture is void here

    Naive two-route capture-recapture estimates 467,541 Canadian metaresearch works, implying Canada produces 59% of the world's metaresearch against an observed 1.9%. The estimator is void here; we cut it rather than dress it up as a lower bound.

    Source field:route1_topic
    14,873
    Source field:route2_naive_lexical
    77,583
    Source field:overlap
    2,468
    Source field:observed_union
    89,988
    Source field:lincoln_petersen_estimate
    467,541
    Source field:entire_topic_space_all_countries
    793,883
    Source field:canada_observed_share_pct
    1.9
    Source field:canada_implied_share_pct
    58.9
    Source field:estimator_void
    yes
    capture_recapture_fails · computed 2026-07-13T11:54:44Z
    07

    The three-model screen

    Three frontier models (Opus 4.8, GPT-5.6 high, Grok 4.5) screened the same 5,600 works, drawn from the real 4.3M frame under a design whose seven strata PARTITION it (an earlier five-stratum design could not reach 12.9% of the frame at all; D22). Design-weighted base rates span 2.54% to 3.81% (1.5x). But the RATE is not the finding, the SETS are: of the 274 works ANY model called metaresearch, only 104 (38%) were called metaresearch by ALL THREE, and 117 (43%) rest on a SINGLE model's opinion; pairwise Jaccard on the in-scope sets is about 50%. THE FIELD'S BOUNDARY IS NOT A LINE THE MODELS SHARE; IT IS A REGION THEY EACH CUT DIFFERENTLY, and that result is STABLE across n = 1,000, 2,000 and 5,600 (unanimity 37%, 37%, 38%). A SECOND, PRETTIER CLAIM DID NOT SURVIVE: at n = 2,000 the models agreed markedly more on 'is this about research at all' (1.43x here) than on 'is it in scope' (1.51x here), and this project said so in capitals; at n = 5,600 the two spreads are within noise (ratio 1.06) and the claim is WITHDRAWN (D23). The largest tier confusion is OUT-vs-T2, every time, at every sample size: the adjacent traditions the inclusiveness criterion exists to protect. The deliverable is not a base rate. It is the disagreement dossier, the 274 works that mark the empirical boundary, each carrying all three models' stated reasons, and the criteria that have to be written against them.

    Source field:frame
    the real 4.3M-work Canadian frame (all 482 OpenAlex partitions)
    Source field:payload
    the rubric's FULL eight fields, including venue (repairs D1)
    Source field:sample
    5,600 works across seven exhaustive strata, with known selection probabilities and French oversampled
    Source field:models
    Claude Opus 4.8; GPT-5.6 (high effort); Grok 4.5 (medium effort)
    Source field:harness
    chunks randomized and manifest-logged before any model ran; the harness writes label files, never the model (repairs D11); every arm reconciled against the manifest (repairs D2)
    Source field:n_labelled_by_all_three
    5,600
    Source field:tranche_homogeneity_p
    0.63
    Source field:tranches_pool
    yes
    Source field:base_rate_weighted_pct
    opus:
    3.81
    gpt:
    2.92
    grok:
    2.54
    Source field:between_model_spread_x
    1.5
    Source field:jaccard_opus_gpt
    50
    Source field:jaccard_opus_grok
    50
    Source field:jaccard_gpt_grok
    56
    Source field:n_about_research_at_all
    opus:
    391
    gpt:
    325
    grok:
    274
    Source field:n_in_scope
    opus:
    224
    gpt:
    163
    grok:
    148
    Source field:spread_about_research_x
    1.43
    Source field:spread_in_scope_x
    1.51
    Source field:variance_is_in_the_rubric_not_the_models
    yes
    Source field:called_in_scope_by_any
    274
    Source field:unanimous_in_scope
    104
    Source field:pct_unanimous_of_any
    38
    Source field:in_scope_by_one_model_only
    117
    Source field:pct_single_model_of_any
    43
    Source field:contested_by_stratum
    about_only:
    67
    venue_new:
    65
    residual:
    64
    aff_core:
    62
    fund_new:
    61
    aff_about:
    59
    french:
    54
    Source field:tier_disagreement_patterns
    OUT/T2:
    77
    T1:
    65
    OUT/T1:
    57
    T2:
    30
    T1/T3:
    11
    T2/T3:
    10
    T1/T2:
    9
    OUT/T1/T2:
    8
    OUT/T1/T3:
    3
    OUT/T2/T3:
    3
    Source field:gpt_schema_violations_first_pass
    18
    Source field:gpt_violation_note
    GPT-5.6 (high) wrote GENRE values ('empirical', 'conceptual') into the TIER field on 18 of 1,000 records in its first pass, in 3 of 20 chunks. The validator caught it because the harness reconciles files against a manifest rather than trusting the model's report. Those chunks were RE-RUN, not repaired: coercing a model's output to the schema is fitting the instrument to the data.
    Source field:deliverable
    pilot/screening/frame1k/disagreement_dossier.json: every work any model called in-scope, with all three labels. This, not the base rate, is what the criteria must be written against.
    Source field:caveat
    These are MACHINE labels and none of them is truth (finding 15). The unanimity rate is not accuracy: three models sharing training data can be wrong together, and they are most correlated exactly on the boundary cases the field's definition turns on. What this measures is where the RUBRIC is underspecified, which is a property of the instrument and is exactly what a criteria document needs. Base rates are design-weighted from a stratified sample, so they estimate the frame; the Jaccard and unanimity figures are unweighted set quantities over the sample and are NOT frame estimates.
    three_model_screen · computed 2026-07-15T22:59:18Z
    08

    Swap the screener, move the answer

    Swap which model is called 'the screener' and the base rate moves from 1.06% to 2.37%: a 2.2x spread, from 45,397 to 101,776 works in the frame. The two screeners agree on in/out for 98.4% of the frame (design-weighted), but that figure is dominated by the settled rejects: agreement falls to 95% inside the contested boundary. THE SCREENER-SWAP RANGE, NOT THE BINOMIAL CI ON EITHER MODEL ALONE, IS THE HONEST UNCERTAINTY ON THE FIELD'S SIZE.

    Source field:n_double_screened
    1,290
    Source field:base_rate_screener_a_pct
    1.06
    Source field:base_rate_screener_b_pct
    2.37
    Source field:swap_ratio_x
    2.24
    Source field:field_size_screener_a
    45,397
    Source field:field_size_screener_b
    101,776
    Source field:published_binomial_ci_contains_b
    no
    Source field:screener_a
    claude-sonnet-4-6 (40 agents, medium effort)
    Source field:screener_b
    gpt-5.6-sol (codex)
    Source field:sampling
    stratified on screener A's label; positives/paratext/boundary taken whole, settled-OUT sampled
    Source field:raw_agreement_inout_pct
    96.6
    Source field:weighted_agreement_pct
    98.4
    Source field:cohens_kappa_inout
    0.681
    Source field:n_disagree_inout
    44
    Source field:gpt_in_claude_out
    37
    Source field:claude_in_gpt_out
    7
    Source field:agreement_by_stratum
    stratum:
    settled_out,boundary,positive,paratext
    n:
    600,599,58,33
    sel_prob:
    0.124921923797626,1,1,1
    agree_pct:
    99,95,87.9,97
    a_says_in:
    0,0,58,0
    b_says_in:
    6,30,51,1
    Source field:caveat
    PROCESS METRIC, NOT ACCURACY. Two LLMs share training data and failure modes; their errors are correlated and most correlated on the boundary. This is not 'duplicate screening': that term's warrant comes from independent human judgment. Accuracy rests on the human-coded probability sample (PROTOCOL s6.2). Agreement is reported per stratum because a pooled kappa on a 1.3% base rate is dominated by the cell where agreement is free (the kappa paradox). And note what the high agreement figure CONCEALS: it is dominated by the settled-OUT mass, while the two screeners imply base rates a factor of two apart. Quoting agreement without the swap would be presenting the reassuring statistic.
    agreement · computed 2026-07-15T22:52:08Z
    09

    The base rate

    Screening 5,737 unfiltered Canadian works against the rubric puts metaresearch at 1.31% of Canadian research, implying ~56,206 works in the 4,299,418-work frame, sizing the field without a search strategy at all. The 95% CI on this screener's labels is 1.03-1.64%, but that is sampling error, NOT the uncertainty: swap the screener and the estimate lands outside it (finding 10). The machine-screener range in finding 10 is the honest uncertainty on field size, and only the human audit can narrow it.

    Source field:n_screened
    5,737
    Source field:screener
    claude-sonnet-4-6, 40 agents, medium effort, locked rubric
    Source field:sampling_frame
    unfiltered Canadian works from the pinned 2026-06-24 snapshot partition
    Source field:tier_counts
    OUT:
    5,621
    T1:
    40
    T2:
    35
    T3:
    41
    Source field:n_in_scope_t1_t2
    75
    Source field:base_rate_pct
    1.31
    Source field:base_rate_ci_lo_pct
    1.03
    Source field:base_rate_ci_hi_pct
    1.64
    Source field:canadian_frame_size
    4,299,418
    Source field:estimated_field_size
    56,206
    Source field:estimated_field_lo
    44,268
    Source field:estimated_field_hi
    70,338
    Source field:topic_route_retrieved
    14,873
    Source field:binomial_ci_is_not_the_uncertainty
    yes
    Source field:caveat
    Machine labels, not a human gold standard, and the binomial CI above is sampling error on ONE screener; the honest uncertainty is the screener-swap range in finding 10 (1.06% to 2.37%). The partition is also not a uniform draw: it under-represents works with abstracts, where the screen finds 2x more metaresearch (finding 11). This is a hypothesis with a denominator; the human-coded probability sample tests it. Do NOT divide topic_route_retrieved by estimated_field_size; see finding 12.
    base_rate · computed 2026-07-15T22:33:57Z
    10

    Base-rate robustness

    The base rate's real bias is not recency but MISSING ABSTRACTS: 31.5% of the partition has none, and the screen finds 0.78% metaresearch there against 1.55% where an abstract exists (chi-square p = 0.023, robust to adjustment for year and language). A third of the frame is screened on its title alone. What this does NOT show, and an earlier draft wrongly claimed, is that the blindness is DIFFERENTIAL by tradition: the T2 no-abstract cell holds 4 works and the interaction is not significant (p = 0.141). That claim is withdrawn, as is 'Erudit's exact profile' (the stratum is 99% English and the works are NEWER, not older). The main effect is the finding.

    Source field:partition
    updated_date=2026-06-24
    Source field:records_sent_to_screener
    6,202
    Source field:records_silently_lost
    465
    Source field:records_lost_pct
    7.5
    Source field:pct_no_abstract_among_lost
    40.2
    Source field:pct_no_abstract_among_labelled
    31.5
    Source field:chisq_p_loss_bias
    0.00013
    Source field:losses_are_biased
    yes
    Source field:works_2000_09
    2,704
    Source field:works_2020_25
    1,195
    Source field:old_to_new_ratio
    2.3
    Source field:partition_skews_old
    yes
    Source field:base_rate_by_era_pct
    2000-09:
    1.24
    2010-19:
    1.35
    2020-25:
    1.4
    Source field:chisq_p_era
    0.906
    Source field:era_events
    75
    Source field:era_check_is_underpowered
    yes
    Source field:no_abstract_share_pct
    31.5
    Source field:base_rate_no_abstract_pct
    0.78
    Source field:base_rate_has_abstract_pct
    1.55
    Source field:chisq_p_abstract
    0.023
    Source field:abstract_effect_x
    2
    Source field:t1_penalty_x
    1.4
    Source field:t2_penalty_x
    3.6
    Source field:t2_no_abstract_cell_count
    4
    Source field:interaction_p
    0.141
    Source field:differential_is_supported
    no
    Source field:differential_claim_withdrawn
    yes
    Source field:no_abstract_mean_year
    2013.7
    Source field:has_abstract_mean_year
    2010.3
    Source field:no_abstract_stratum_pct_english
    99
    Source field:p_no_abstract_given_english_pct
    32
    Source field:p_no_abstract_given_non_english_pct
    12.7
    Source field:erudit_profile_claim_holds
    no
    Source field:residual_bias_direction
    anti-conservative for coverage claims: the partition over-represents abstract-less works (31.5%), where the screen finds 2x less metaresearch, so 1.31% likely UNDER-states the frame's base rate, and the field is larger than the headline implies
    Source field:supersedes
    THREE retractions live here. (1) An earlier version tested ERA only, called the base rate robust, and published 3100/2900/1500 as percentages (a dplyr summarise() column-masking bug). (2) It then claimed the blindness is DIFFERENTIAL, T2 losing 3.6x against T1's 1.4x, and made that the proposal's whole answer to the inclusiveness criterion. The T2 no-abstract cell holds FOUR works and the interaction is not significant (p = 0.141). WITHDRAWN. (3) It claimed missing abstracts track 'older, non-English' records, 'Erudit's exact profile'. Backwards: they are NEWER, and the stratum is 99% ENGLISH. WITHDRAWN. See DEVIATIONS.md D4, D5, D6.
    Source field:caveat
    What survives is the MAIN EFFECT and only the main effect: a third of the frame is screened on its title alone and the screen finds half as much metaresearch there (p = 0.023, robust to adjustment for year and language). That is a real coverage problem and a reason to stratify the audit on abstract availability. It is NOT evidence of differential blindness by tradition, and this finding no longer says it is. Separately, the era check is underpowered (75 events, 3 strata) and cannot refute an era effect; it merely fails to detect one. And note D2: the harness silently dropped 465 records, non-randomly, on this very covariate.
    base_rate_robustness · computed 2026-07-13T11:54:47Z
    11

    What the labels cannot tell us

    Two limits on the machine labels, found by attacking the fixes. (A) The topic route's recall is 12% against screener A and 7% against screener B: finding 14's instrument (ii) scores filters against MACHINE labels, so it measures agreement with a machine, not accuracy, and finding 10 already showed that swings by a factor of two. The conclusion strengthens (the second screener thinks the route is WORSE) but 12% is not truth. (B) The pilot holds exactly 1 French in-scope work, so seeing 20 French positives needs ~1,620 coded French records against an audit budget of 1,000. FRENCH SENSITIVITY IS NOT ESTIMABLE IN AN OPENALEX-ONLY FRAME. That is an argument for the Erudit harvest, not against the French claim, but the pilot did not run that harvest, so the power is stated as a condition rather than a promise.

    Source field:route_recall_vs_screener_a_pct
    12
    Source field:route_recall_vs_screener_b_pct
    7
    Source field:route_recall_a_ci
    • 5.6
    • 21.6
    Source field:route_recall_b_ci
    • 2.5
    • 14.3
    Source field:positives_screener_a
    75
    Source field:positives_screener_b
    88
    Source field:instrument_ii_is_model_dependent
    yes
    Source field:instrument_ii_caveat
    Finding 14's instrument (ii) scores a filter against the 5,737 MACHINE labels. That measures agreement with a machine, not accuracy. Swap the machine and the topic route's recall moves from 12% to 7%. The conclusion (the route finds a small fraction) survives and strengthens; the NUMBER is not a measurement against truth.
    Source field:french_records_in_pilot
    81
    Source field:french_in_scope_in_pilot
    1
    Source field:french_records_needed_for_20_positives
    1,620
    Source field:audit_budget_records
    1,000
    Source field:french_stratum_is_powered
    no
    Source field:french_power_depends_on
    the Erudit harvest, which the pilot did NOT run (it verified the endpoint: 379 live sets)
    Source field:caveat
    (A) is a limit on every recall number this project quotes against machine labels, including its own headline. (B) is a limit on the inclusiveness promise: French sensitivity cannot be estimated in an OpenAlex-only frame, because Erudit matches zero OpenAlex sources and the francophone literature is therefore largely absent from the frame rather than merely sparse in it. Both are stated in the proposal rather than left for a reviewer.
    label_limits · computed 2026-07-13T11:54:49Z
    12

    Agent variance

    Haiku, the model finding 13 budgets the entire screen on, lands near Sonnet's base rate (1.27% vs 1.06%) and agrees with it on 98.1% of the frame, but their in-scope SETS overlap 16% unweighted and 10% design-weighted (Sonnet-GPT: 54%/37%); of Sonnet's 58 positives Haiku agrees on 12. RATE AGREEMENT IS NOT SET AGREEMENT. Worse: agents of the SAME model on the SAME prompt disagree beyond chance in BOTH arms after conditioning on stratum (CMH p = 0.0056 and 0.015), with raw spreads 3.1x and 5.2x against 2.2x between models, and the agents' ordering FLIPS between arms. The noise inside one model is at least the size of the difference between models, and the pilot's own 40-agent screen is too underpowered to rule the same instability out (5 of 37 chunks found zero metaresearch; p = 0.113).

    Source field:n_works
    1,290
    Source field:base_rate_sonnet_pct
    1.06
    Source field:base_rate_gpt_pct
    2.37
    Source field:base_rate_haiku_pct
    1.27
    Source field:agreement_haiku_sonnet_pct
    98.1
    Source field:jaccard_sonnet_gpt_pct
    54
    Source field:jaccard_sonnet_haiku_pct
    16
    Source field:jaccard_gpt_haiku_pct
    12
    Source field:wjaccard_sonnet_gpt_pct
    37
    Source field:wjaccard_sonnet_haiku_pct
    10
    Source field:wjaccard_gpt_haiku_pct
    6
    Source field:sonnet_positives
    58
    Source field:haiku_agrees_on
    12
    Source field:agent_rates_raw_pct
    agent-1:
    1.6
    agent-2:
    1.25
    agent-3:
    3.85
    Source field:stratum_mix_differs_by_agent_p
    0.0000000000000000134
    Source field:arm1_cmh_p
    0.00564
    Source field:arm1_permutation_p
    0.0275
    Source field:arm1_raw_spread_x
    3.1
    Source field:arm2_cmh_p
    0.0152
    Source field:arm2_permutation_p
    0.0075
    Source field:arm2_raw_spread_x
    5.2
    Source field:agent_order_replicates
    no
    Source field:spread_weighted_x_leverage_sensitive
    13.2
    Source field:between_model_spread_x
    2.2
    Source field:within_at_least_matches_between
    yes
    Source field:pilot_chunks
    37
    Source field:pilot_agents_finding_zero
    5
    Source field:pilot_between_agent_p
    0.113
    Source field:pilot_underpowered_not_homogeneous
    yes
    Source field:caveat
    The first draft's between-agent test was CONFOUNDED: the stratum mix differs by agent (p = 1.3e-17), and the draft asserted a verification that did not exist (DEVIATIONS.md D13). The tests above condition on stratum, and the heterogeneity survives in both arms. The 13.2x design-weighted spread the draft led with rests on five high-weight events and is demoted to a recorded, leverage-sensitive descriptive. THE INFERENCE IS SCOPED: there are three agents per arm, assigned consecutive chunk blocks without randomization or a run-time manifest, so these p-values license 'these runs are not exchangeable', not a population claim about agents in general; that is exactly enough to break a budget that assumed exchangeability, and the full study assigns agents randomized, manifest-logged, fixed-size chunks with a duplicate-agent reliability arm. The pilot's own agents give p = 0.113 on ~2 expected events per chunk: underpowered, so the pilot is uninformative on agent homogeneity, not exonerated. Haiku was tested on the six-field payload (the pilot's own deviation, D1) and on a guided eight-field arm, so the defensible conclusion is 'not shown to be an interchangeable measurer, and unstable in the arms tested', not 'cannot screen'. An eight-field neutral-prompt arm is not used at all: one agent claimed six label files it never wrote (D11), so no payload effect is reported from any arm. The guided arm is used ONLY for the between-agent contrast, which its shared prompt leaves internally valid.
    agent_variance · computed 2026-07-15T22:52:18Z
    13

    The abstract gap is structural

    The screen's largest measured bias is the abstract gap: 23.3% of the frame (1,003,117 works) has NO ABSTRACT, and finding 11 showed the screen finds HALF as much metaresearch there. Cascading PubMed, Europe PMC and Crossref recovers 37.8% of a 500-work sample, cutting title-only exposure to ~14.5% of the frame. But I BUILT THE CASCADE AROUND CROSSREF as the discipline-agnostic rescue, and it recovered 2 abstracts against PubMed's 180: publishers do not deposit them, so THAT RESCUE DOES NOT EXIST (D15). The gap is therefore not a metadata failure a better index fixes; it is STRUCTURAL. Recovery is 91.2% for reviews against 6.2% for book chapters, 38.8% English against 15.4% French. So the tempting shortcut, 'just screen the works that have abstracts', is a SELECTION ON A COVARIATE THAT PREDICTS THE OUTCOME which would delete 61.6% of book chapters against 22.1% of articles, AND the works it deletes are exactly the works no cascade can rescue. Defensible only as a DECLARED exclusion with a measured cost, and the audit keeps a sampling floor in it.

    Source field:frame_works
    4,299,418
    Source field:frame_works_no_abstract
    1,003,117
    Source field:pct_frame_no_abstract
    23.3
    Source field:pct_dropped_by_type
    book-chapter:
    61.6
    letter:
    52.1
    editorial:
    43.3
    review:
    29.5
    article:
    22.1
    book:
    21.6
    other:
    21.1
    report:
    18.6
    preprint:
    14
    dataset:
    8.2
    dissertation:
    4.3
    Source field:pct_dropped_by_language
    fr:
    21.6
    en:
    23.7
    Source field:abstracts_only_is_a_selection_on_the_outcome
    yes
    Source field:sampled
    500
    Source field:sources
    PubMed (Entrez) -> Europe PMC (REST) -> Crossref (REST)
    Source field:recovered
    189
    Source field:pct_gap_recovered
    37.8
    Source field:recovered_by_source
    pubmed:
    180
    europepmc:
    7
    crossref:
    2
    Source field:hypothesis_crossref_would_be_load_bearing
    no
    Source field:crossref_recovered
    2
    Source field:europepmc_recovered
    7
    Source field:pubmed_recovered
    180
    Source field:no_discipline_agnostic_rescue_exists
    yes
    Source field:the_gap_is_structural_not_a_metadata_failure
    yes
    Source field:recovery_pct_by_type
    review:
    91.2
    article:
    40.5
    preprint:
    38.5
    book-chapter:
    6.2
    letter:
    0
    Source field:recovery_pct_english
    38.8
    Source field:recovery_pct_french
    15.4
    Source field:residual_pct_frame_title_only
    14.5
    Source field:caveat
    Run on a 500-work hash-ordered sample of the no-abstract stratum, not the frame: the cascade is rate-limited and a million lookups is days. The recovery rate is an estimate with sampling error, and it is an estimate of a CEILING (an abstract that EXISTS is recoverable; it does not follow the screen then classifies the work correctly). Only works with a DOI can be looked up, so the DOI-less part of the stratum is untouched and its size bounds what any cascade can do. PubMed and Europe PMC are biomedical; Crossref is not, and it is in the chain for exactly that reason: a cascade of biomedical indexes would close the gap unevenly and make the residual bias MORE discipline-shaped while appearing to improve coverage. That reasoning was right and the remedy is not available: Crossref recovered 2 of 189 and Europe PMC 7, because publishers largely do not deposit abstracts to Crossref, so no discipline-agnostic rescue exists (D15). Restricting screening to abstract-bearing works remains a DECLARED EXCLUSION with a measured cost, not a scoping convenience, and the audit keeps a sampling floor in the excluded stratum so that cost stays estimable.
    abstract_cascade · computed 2026-07-13T11:55:17Z
    14

    A boolean over a four-state space

    Joined to the Canadian frame by DOI, Retraction Watch records 143 works that OpenAlex does NOT flag as retracted, including 49 outright retractions. But the undercount is the smaller problem. 52 of these carry an EXPRESSION OF CONCERN, and OpenAlex HAS NO FIELD FOR ONE: `is_retracted` is a boolean over a state space with at least four values (retraction, expression of concern, correction, reinstatement), so it can express one and silently reports the rest as FALSE, which reads as 'fine'. Nor can a boolean carry WHY. This is finding 1's disease in a second schema: the canonical database cannot express the distinction the field turns on.

    Source field:source
    Retraction Watch (Crossref-licensed), joined by bare lowercased DOI
    Source field:frame_works_with_doi
    3,690,953
    Source field:openalex_is_retracted_flags
    1,584
    Source field:matched_in_retraction_watch
    1,052
    Source field:openalex_misses
    143
    Source field:outright_retractions_missed
    49
    Source field:expressions_of_concern
    52
    Source field:missed_by_nature
    Expression of concern:
    52
    Retraction:
    49
    Correction:
    32
    Reinstatement:
    10
    Source field:top_reasons
    Investigation by Journal/Publisher:
    335
    Unreliable Results and/or Conclusions:
    259
    Concerns/Issues about Data:
    233
    Investigation by Third Party:
    169
    Concerns/Issues about Referencing/Attributions:
    151
    Concerns/Issues about Results and/or Conclusions:
    122
    Concerns/Issues about Peer Review:
    109
    Investigation by Company/Institution:
    101
    Source field:openalex_has_eoc_field
    no
    Source field:is_a_frame_route
    no
    Source field:caveat
    This is an ATTRIBUTE, not a frame route. A retracted cardiology paper is retracted cardiology, not metaresearch, and admitting works on the strength of a retraction would let an interesting signal masquerade as the estimand. The DOI join can only see works that HAVE a DOI, and OpenAlex flags some works Retraction Watch does not match, which may be DOI drift rather than disagreement; the asymmetry reported here is one-directional on purpose (what RW adds), because that is the direction the join can support.
    retraction_record · computed 2026-07-13T11:55:00Z
    15

    Topic-route recall

    Scored against the rubric, the topic route retrieves 12% of Canadian metaresearch (95% CI 5.6-21.6%) at 60% precision: it misses 66 of 75. It fails because OpenAlex files a work by what it is about, and metaresearch about cardiology reads as cardiology: the field is invisible to topic retrieval precisely because it is about other fields.

    Source field:n_screened
    5,737
    Source field:n_metaresearch
    75
    Source field:n_retrieved_by_route
    15
    Source field:true_positives
    9
    Source field:false_positives
    6
    Source field:false_negatives
    66
    Source field:recall_pct
    12
    Source field:recall_ci_lo_pct
    5.6
    Source field:recall_ci_hi_pct
    21.6
    Source field:precision_pct
    60
    Source field:precision_ci_lo_pct
    32.3
    Source field:precision_ci_hi_pct
    83.7
    Source field:missed_works_by_field
    Social Sciences:
    17
    Medicine:
    16
    Health Professions:
    7
    Business, Management and Accounting:
    4
    Economics, Econometrics and Finance:
    4
    Computer Science:
    3
    Source field:scored_by
    primary_topic.id (the key R/frame.R defines the route with), not display name
    Source field:route_size_implied_by_sample
    11,241
    Source field:route_size_from_api
    14,873
    Source field:reconciliation_gap_x
    1.32
    Source field:supersedes
    the earlier 32.4% figure (14,873/45,850), which divided a retrieved set by a true field size
    Source field:caveat
    Small n on the positive class ( 75 metaresearch works, of which 15 were on the route), so the intervals are wide. Separately, the sample-implied route size does not reconcile with the API's 14,873 and I cannot say why at n = 15 on-route works; the likeliest cause is that one updated_date partition is not a uniform draw (finding 11). Recall is unaffected: it is a within-sample ratio, not an extrapolation.
    topic_route_recall · computed 2026-07-13T11:54:47Z
    16

    Funder-route recall

    The PRIMARY ESTIMAND leans on CA-FUND to rescue works whose affiliation is missing (finding 2: 64% have no affiliation string), and nothing had tested it. Tested against CIHR's own database of 44,190 funded projects, the result REFUTED THE HYPOTHESIS I WROTE BEFORE RUNNING IT: OpenAlex tags 178,133 frame works with CIHR, or 4.03 per grant, a plausible rate showing no under-tagging (DEVIATIONS.md D14). What the data DOES support needs no hypothesis of mine: 71.2% OF THE FRAME CARRIES NO FUNDER METADATA AT ALL, which is CA-FUND's hard ceiling, and 65.9% of Canadian-AFFILIATED works carry none either. Both clauses of the estimand rest on metadata that is mostly absent, which is why the frame is a union of four routes and why the audit must sample the works no route reached.

    Source field:external_criterion
    CIHR's own project database (44,190 projects), which owes nothing to OpenAlex
    Source field:cihr_projects
    44,190
    Source field:cihr_distinct_pis
    27,536
    Source field:cihr_funder_id
    https://openalex.org/F4320334506
    Source field:frame_works
    4,299,418
    Source field:frame_works_tagged_cihr
    178,133
    Source field:tagged_publications_per_funded_project
    4.03
    Source field:papers_per_grant_is_a_plausible_rate_not_a_defect
    yes
    Source field:expectation_i_wrote_before_running_and_that_was_false
    that OpenAlex under-tags CIHR so badly it falls below one paper per grant. It is 4.03 per grant, a plausible rate. See DEVIATIONS.md D14.
    Source field:works_ca_fund_rescues_alone
    166,743
    Source field:frame_works_with_any_funder
    1,239,950
    Source field:pct_frame_with_any_funder
    28.8
    Source field:pct_frame_with_no_funder
    71.2
    Source field:ca_aff_works
    2,734,192
    Source field:ca_aff_works_with_no_funder
    1,802,605
    Source field:pct_ca_aff_with_no_funder
    65.9
    Source field:top_recorded_funders
    Natural Sciences and Engineering Research Council of Canada:
    294,401
    Canadian Institutes of Health Research:
    178,133
    National Institutes of Health:
    81,262
    National Natural Science Foundation of China:
    67,358
    National Science Foundation:
    62,497
    Canada Research Chairs:
    35,113
    European Commission:
    34,039
    Social Sciences and Humanities Research Council of Canada:
    33,473
    Source field:is_an_aggregate_not_a_record_level_join
    yes
    Source field:caveat
    CIHR's CSV carries no DOIs and no publication links, so this is an AGGREGATE reconciliation, not a record-level known-item join, and NO RECALL POINT ESTIMATE is claimed. It establishes a CEILING on CA-FUND (the route cannot see a funder OpenAlex never recorded), which is a bound, not a measurement. The papers-per-grant ratio is reported because I ran it, and it REFUTES the hypothesis I wrote before running it: at 4.03 per grant it is a plausible publication rate and shows no CIHR under-tagging at all. It is also the wrong instrument, for the same reason the retracted 32.4% coverage figure was (D3): its numerator and denominator are not linked record to record, so the quotient has no estimand behind it. Record-level linkage is finding 21.
    funder_route_recall · computed 2026-07-13T11:55:02Z
    17

    Canadian linkage

    Affiliation finds 14,873 works; a further 3,964 are about Canada with no Canadian affiliation. NSERC has 9.2x SSHRC's linked works, so funder-based rules under-count the social sciences.

    Source field:by_affiliation
    14,873
    Source field:by_funder
    1,331
    Source field:about_canada
    5,704
    Source field:about_canada_no_affiliation
    3,964
    Source field:funder_works
    Canadian Institutes of Health Research:
    194,681
    Natural Sciences and Engineering Research Council of Canada:
    433,090
    Social Sciences and Humanities Research Council of Canada:
    46,913
    Canada Foundation for Innovation:
    16,753
    Source field:sshrc_works
    46,913
    Source field:nserc_works
    433,090
    Source field:nserc_to_sshrc_ratio
    9.2
    Source field:affiliation_noise
    University of London:
    494
    Impact:
    475
    canadian_linkage · computed 2026-07-13T11:54:44Z
    18

    Trial linkage

    The registry is the only REFERENCE STANDARD in this project not made of machine labels: ClinicalTrials.gov knows a Canadian trial happened independently of any pipeline, so it cannot be wrong in the pipeline's favour. Of 304 publications SPONSORS THEMSELVES reported as results of completed Canadian-located trials, the frame holds 160: a naive recall of 52.6%, which fell so close to the 44.5% THIS PROPOSAL OPENS WITH that it read as a replication. IT IS AN ARTIFACT, and running the disambiguation is the only thing that caught it: 137 of the 145 'misses' have NO CANADIAN AUTHOR (multi-site international trials with a Canadian SITE), and a frame of Canadian RESEARCH is CORRECT to exclude them. A trial with a Canadian site is not a publication with a Canadian author. Against the population the frame actually claims, recall is 95.2% (95% CI 90.8-97.9), and the real defect is 8 works OpenAlex holds WITH a Canadian author that the frame's own routes still missed. The frame is GOOD at this, the dramatic parallel was a coincidence between two different populations, and I had every incentive not to check. DEVIATIONS.md D16.

    Source field:reference_standard
    ClinicalTrials.gov: completed trials with a Canadian location, and the RESULT publications the sponsors themselves reported
    Source field:why_not_a_frame_route
    A trial registration is not a publication and not metaresearch. Registrations contribute NO records to the frame; the registry is a reference standard, not a source.
    Source field:why_it_matters
    Every other recall number in this project is scored against MACHINE labels (finding 15). A registry knows a trial happened independently of any pipeline, so it cannot be wrong in the pipeline's favour. This is the only instrument here that measures the frame against a world that exists without it.
    Source field:trials_retrieved
    1,000
    Source field:trials_with_result_publication
    123
    Source field:pct_trials_with_result_publication
    12.3
    Source field:known_result_pmids
    596
    Source field:resolvable_to_doi
    304
    Source field:present_in_frame
    160
    Source field:naive_frame_recall_pct
    52.6
    Source field:naive_recall_is_an_artifact
    yes
    Source field:naive_recall_ci
    • 46.9
    • 58.4
    Source field:missing_total
    145
    Source field:missing_no_canadian_author
    137
    Source field:missing_route_gap_canadian_author
    8
    Source field:missing_not_in_openalex
    0
    Source field:claimable_population
    168
    Source field:adjusted_recall_pct
    95.2
    Source field:adjusted_recall_ci
    • 90.8
    • 97.9
    Source field:is_metaresearch_recall
    no
    Source field:caveat
    This measures FRAME recall (does the Canadian frame hold the publication at all?), NOT metaresearch recall: trial reports are primary research and the rubric screens them OUT. It is the precondition for screening, not the screen. The reference standard is the sponsor's OWN reported result publications, so it is incomplete in a known direction: sponsors under-report, which means the true set of trial publications is LARGER than the standard and this recall figure is measured only on the ones we can see. Trials are matched by Canadian LOCATION, which is not the same as Canadian authorship, so some result publications may have no Canadian author and legitimately fall outside the frame; that direction is not controlled here and it bounds the interpretation. Only PMIDs resolvable to a DOI can be looked up.
    trial_linkage · computed 2026-07-13T11:55:20Z
    19

    Preprint coverage

    The frame carries 156,086 preprints, and the tempting move was to assert that OpenAlex covers the preprint servers and skip the ingest. That is a COVERAGE CLAIM, and this project does not get to make one from inside the pipeline being claimed for: it is the topic route certifying its own recall (finding 12). Measured instead against the servers' OWN API, OpenAlex indexes 99.6% of the 705 bioRxiv and medRxiv preprints enumerated (3 missing). The claim survives measurement, so no separate preprint ingest is built.

    Source field:external_criterion
    the bioRxiv/medRxiv details API, which enumerates the servers' own corpus and owes nothing to OpenAlex
    Source field:window
    2023-01-01 to 2025-12-31
    Source field:preprints_enumerated
    705
    Source field:indexed_by_openalex
    702
    Source field:missing_from_openalex
    3
    Source field:pct_indexed
    99.6
    Source field:sampled_with_canadian_author
    45
    Source field:of_those_present_in_frame
    45
    Source field:frame_preprints_total
    156,086
    Source field:separate_ingest_needed
    no
    Source field:caveat
    The bioRxiv API exposes no author country, so the sample is mostly non-Canadian and the clean quantity is INDEX coverage (does OpenAlex hold the preprint at all?), not Canadian recall: a preprint the index lacks cannot enter any frame by any route, so index coverage is the binding upper bound. The Canadian sub-count is small and is reported as a check, not as an estimate. arXiv is NOT tested here (its OAI endpoint pages differently); the frame carries 11,647 arXiv works and that claim remains untested against arXiv itself.
    preprint_coverage · computed 2026-07-13T11:55:18Z
    20

    The power of a human audit

    The audit as first specified could not measure what it exists to measure. A simple random sample of the 4,243,096-record screened-out mass yields an expected 0.4 hits from the 600 records budgeted; seeing 20 would take 2,009 coder-hours against the 65 planned. Score-stratified oversampling rescues it only partly (the pilot shows 30 of 37 disputed works sit at the contested boundary), and it stays BLIND to the confident rejects. Which works those are is a MECHANISM, not a measured differential: the rubric says judge on the title alone with no abstract, and says T2 work may use none of the field's vocabulary, so a work with neither is rejected confidently and never sampled. (An earlier version cited a 3.6x figure here as if it were measured; it rested on four works and is withdrawn.) So recall is measured two other ways: against the 5,737 works that already carry rubric labels, and by KNOWN-ITEM recall on a venue reference set, because venue is an external criterion immune to the aboutness that defeats everything else.

    Source field:problem
    A simple random sample of the screened-out stratum cannot measure screening sensitivity: the works the screen wrongly rejected are a vanishing fraction of a 3.2M-record rejected mass.
    Source field:screened_out_pool
    4,243,096
    Source field:screened_out_is_the_screens_rejects
    yes
    Source field:audit_screened_out_budgeted
    600
    Source field:expected_hits_at_95_recall
    0.4
    Source field:codings_needed_for_20_hits_at_95_recall
    30,134
    Source field:coder_hours_needed
    2,009
    Source field:coder_hours_budgeted
    65
    Source field:naive_audit_is_powered
    no
    Source field:disputed_in_boundary
    30
    Source field:disputed_in_settled_rejects
    6
    Source field:misses_concentrate_near_threshold
    yes
    Source field:blind_spot
    Score-stratified oversampling finds the works the screener ALMOST caught; it is blind to the ones it rejected CONFIDENTLY. WHICH works those are is a MECHANISM, not a measurement: the rubric says judge on the title alone when the abstract is missing, and the rubric also says T2 work may use none of the field's vocabulary, so a work with neither is rejected confidently and sits deep in the settled rejects. An earlier version cited a 3.6x differential from finding 11 as if this were measured. It is not, and that figure is withdrawn (DEVIATIONS.md D6). The venue instrument is how the prediction gets tested rather than asserted.
    Source field:blind_spot_is_a_mechanism_not_a_measurement
    yes
    Source field:instrument_1
    Measure any filter's recall against the 5737 works that already carry full-rubric labels, exactly as finding 12 scored the topic route. Free, and it needs no needle-hunting in the discarded mass.
    Source field:instrument_2
    Known-item recall on an external criterion: VENUE. A Canadian-authored paper in Social Studies of Science is T2 by where it was published, whatever its abstract is about. Venue is immune to aboutness, which is what defeats topic retrieval (finding 12) and title-only screening (finding 11), so it is the only instrument that can see into the blind spot.
    Source field:reference_venues
    • Social Studies of Science
    • Scientometrics
    • Quantitative Science Studies
    • Research Integrity and Peer Review
    • Journal of the Association for Information Science and Technology
    • Research Evaluation
    • Accountability in Research
    • PLOS ONE (metaresearch collection)
    • Recherches qualitatives
    • Documentation et bibliotheques
    Source field:n_labelled_works
    5,737
    Source field:n_in_scope
    75
    Source field:caveat
    The recall grid assumes the screen's misses are uniform in the rejected mass, which the pilot shows they are not (they concentrate at the boundary). That makes the naive design LESS hopeless than the grid implies but does not save it, and it does nothing at all about the confident-reject blind spot. Known-item recall on a venue reference set is a non-probability estimate: it bounds and diagnoses recall on the hard cases, it does not replace the design-weighted population estimate.
    audit_power · computed 2026-07-13T11:54:49Z
    21

    What screening costs

    The full rubric over the 4,299,418-work frame with two screeners costs $6,567, not the $30,633 an earlier version of this script reported: the rubric is a system prompt sent once per CALL, and the pilot's own chunks batch 155 works per call, so it was overcharged 4.7x. That error was not cosmetic. It made me propose a cheap TRIAGE in front of the screen, which is a RETRIEVAL STEP in a project whose central finding is that retrieval destroys these maps. With the arithmetic right the triage is unnecessary: the full rubric over EVERY work in the frame, plus a second screener on a 20,000-record sample, costs $1,110 at the v1 instrument and $1,279 at the locked v3.1 instrument that will actually run. THE PREFILTER IS DELETED. And the compute is SELF-FUNDED: the call's CAD $4,000 covers, in its own words, 'travel, accommodation, and related participation costs', not inference and not coder wages (D30).

    Source field:measured_from
    pilot/screening/chunks/*.json (6-field), pilot/screening/haiku/p8_guided/chunk_*.json (8-field), docs/protocol/rubric.md; not asserted
    Source field:chars_per_token_assumed
    4
    Source field:tokens_per_work
    312
    Source field:tokens_per_work_six_field_deviation
    256
    Source field:payload_inflation_8_over_6
    1.22
    Source field:costed_at_the_rubric_payload_not_the_deviation
    yes
    Source field:tokens_rubric
    1,878
    Source field:tokens_per_label
    37
    Source field:batch_works_per_call
    155
    Source field:frame_size
    4,299,418
    Source field:grant_usd_approx
    2,900
    Source field:naive_cost_charging_rubric_per_work_usd
    30,633
    Source field:corrected_cost_two_screeners_usd
    6,567
    Source field:cost_overstatement_x
    4.7
    Source field:cost_full_rubric_whole_frame_cheap_usd
    1,094
    Source field:cost_full_rubric_whole_frame_sonnet_usd
    3,283
    Source field:cost_second_screener_on_sample_usd
    15
    Source field:second_screener_n
    20,000
    Source field:total_no_prefilter_usd
    1,110
    Source field:tokens_rubric_v31
    4,604
    Source field:tokens_per_label_v31
    49
    Source field:cost_full_rubric_whole_frame_v31_usd
    1,261
    Source field:cost_second_screener_v31_usd
    18
    Source field:total_no_prefilter_v31_usd
    1,279
    Source field:award_pays_for_compute
    no
    Source field:award_purpose_in_the_calls_words
    travel, accommodation, and related participation costs
    Source field:prefilter_needed
    no
    Source field:prefilter_design_total_usd
    1,004
    Source field:prefilter_saving_usd
    -106
    Source field:prefilter_deleted_because
    It is a RETRIEVAL STEP, in a project whose central finding is that retrieval destroys these maps (finding 12: the topic route finds 12% of the field). It existed only because the rubric was miscosted at once per WORK rather than once per CALL, overstating the alternative 5.1x. With the arithmetic right it saves almost nothing and costs the thesis. Deleted.
    Source field:pilot_works_screened
    6,202
    Source field:frame_to_pilot_ratio
    693
    Source field:method_that_scales
    batch inference, not agent fan-out
    Source field:supersedes
    an earlier version charged the rubric once per work, reported $24,379 for the full-frame two-screener design, called it 'eight times the grant', and used that to justify a cheap prefilter. The rubric is a system prompt sent once per CALL, and the pilot's own chunks batch 155 works per call. DEVIATIONS.md D7.
    Source field:caveat
    Token counts use a 4-chars-per-token approximation and list prices as of 2026-07; both will move, and the conclusion is robust to +/-25% in either. Screening the frame with a cheaper model than the pilot used makes the CHOICE OF MODEL more consequential, not less: finding 10 shows two screeners already imply base rates a factor of two apart. That is precisely why the second screener and the screener-swap range are reported, and why the human audit is the study.
    screening_cost · computed 2026-07-13T12:23:27Z
    22

    OpenAlex is metered

    The OpenAlex API is metered (1,000 credits per ~11h; $0.10 free tier). Enumerating the 4,299,418-work Canadian frame needs 21,498 cursor-paged calls: 10 days per pass on the free tier. An API-based pipeline at this scale is neither free nor reproducible; the pinned snapshot is both.

    Source field:observed_retry_after_s
    40,268
    Source field:observed_ratelimit_limit
    1,000
    Source field:observed_free_tier_usd
    0.1
    Source field:observed_cost_per_call_usd
    0.0001
    Source field:frame_size_works
    4,299,418
    Source field:per_page_max
    200
    Source field:calls_for_one_pass
    21,498
    Source field:days_on_free_tier_per_pass
    10
    Source field:prepaid_cost_per_pass_usd
    2.15
    openalex_is_metered · computed 2026-07-13T11:54:45Z
    23

    active_learning

    THE BATCH-OF-100 LOOP WORKS, AND THE 20 RANDOM ANCHORS IN IT DO NOT. Against a random-batch control at identical budget, active acquisition (50 max-teacher-disagreement + 30 max-uncertainty + 20 random) takes held-out AP from 0.0143 to 0.1391 while the control reaches 0.0502: 2.8x, winning 18 of 20 rounds, with held-out positive-set churn falling to 0.1111. At a ~1% base rate a random batch of 100 holds one positive, and the control curve is what that buys. BUT the anchors, which were supposed to keep an unbiased evaluation stream alive, hold 3 positives after 400 draws: prevalence 0.0057 with a 95% interval 0.01594 wide, WIDER THAN THE QUANTITY ITSELF. They are demoted to drift sentinels. THE LOOP LEARNS; THE PREREGISTERED HUMAN AUDIT MEASURES. And AP here is agreement with the majority TEACHER, so the whole curve is imitation, not accuracy. The first version of this script said the loop DEGRADED, because it measured AP on the shrinking pool the algorithm was itself editing; the churn sat at exactly 1.000 for twenty rounds, which is a constant, not a measurement (D33).

    Source field:batch
    batch_size:
    100
    max_disagreement:
    50
    max_uncertainty:
    30
    random_anchor:
    20
    Source field:rounds
    20
    Source field:ap_active_start
    0.0143
    Source field:ap_active_end
    0.1391
    Source field:ap_random_start
    0.0238
    Source field:ap_random_end
    0.0502
    Source field:active_over_random_x
    2.8
    Source field:active_beats_random_in_rounds
    18
    Source field:holdout_churn_final
    0.1111
    Source field:pi_directive_vindicated
    yes
    Source field:what_ap_measures_here
    Agreement with the MAJORITY TEACHER, not accuracy. A rising curve is better imitation of three LLMs that overlap each other at Jaccard ~0.5. No amount of imitation crosses that gap, and the audit is the only instrument that can.
    Source field:anchors_n
    400
    Source field:anchors_positives
    3
    Source field:anchors_prevalence
    0.0057
    Source field:anchors_ci_width
    0.01594
    Source field:anchors_are_a_usable_evaluation_stream
    no
    Source field:anchors_demoted_to
    drift sentinels, rescored under the frozen final model, never used to pick a batch
    Source field:simulation_caveat
    Every query reveals a label that already exists: 16,800 labels over 5,600 works. This is in-sample recycling. It demonstrates the loop's MECHANICS and cannot show that the loop matures on 4.3M unlabeled works, nor validate the v3.1 instrument, under which no screening has run.
    Source field:superseded_result
    The first version of this experiment measured AP and churn on the POOL of unrevealed works. Active learning REMOVES contested works from the pool by construction, so the pool gets easier every round and the evaluation target moves under the metric. It reported AP FALLING from 0.019 to 0.011 with churn pinned at exactly 1.000 for twenty straight rounds, and the conclusion was going to be 'the loop degrades'. The constant churn is what gave it away. A metric computed on a set the algorithm is actively editing is not a metric. See DEVIATIONS.md D33.
    active_learning · computed 2026-07-15T22:52:19Z
    24

    adjudication

    AN INDEPENDENT BLINDED JUDGE SAYS THE RUBRIC IS SILENT ON 89% OF THE WORKS THE MODELS FOUGHT OVER. Fable 5, which is not one of the three screened arms, adjudicated all 179 contested works seeing three screener opinions as A/B/C in random order with no model names, so it was asked which reading of the RUBRIC is right rather than which model to trust. It shows no arm preference (spread 1.16x, chi-square p = 0.493), so it is not a fourth vote for one of the screeners, and it overruled ALL THREE on 6 works. Its verdict on the instrument: the rubric is SILENT on 159 of 179 contested works, and only 17 splits are a screener misapplying a rule that actually exists. THE FIELD'S BOUNDARY IS BEING SET BY THE SCREENER, NOT BY THE INSTRUMENT. The seams that cost the most agreement are methods_dev_vs_study (39 works), lis_sts_asymmetry (26), and history_of_science (19). This is what converts adjudication from a tiebreak into a measurement: not 179 answers, but a frequency-weighted census of which sentences the rubric is missing and exactly how many works each one costs. The judge's TIERS are used for nothing and produce no estimate: a model cannot certify a model, and that argument applies to the judge too.

    Source field:judge
    Fable 5 (xhigh), which is NOT one of the three screened arms
    Source field:blinded
    yes
    Source field:blinding
    three screener opinions presented as A/B/C, order randomized per work, no model names in the prompt
    Source field:contested_works
    179
    Source field:judge_agrees_with
    • 94
    • 95
    • 109
    Source field:judge_agrees_with_none
    6
    Source field:arm_preference_spread
    1.16
    Source field:arm_preference_chisq_p
    0.493
    Source field:judge_is_captured
    no
    Source field:rubric_is_silent
    159
    Source field:pct_rubric_is_silent
    89
    Source field:splits_from_screener_error
    17
    Source field:seam_census
    • 39
    • 26
    • 19
    • 19
    • 17
    • 12
    • 10
    • 8
    • 8
    • 7
    • 6
    • 5
    • 3
    Source field:adjudicated_tiers
    • 90
    • 23
    • 36
    • 30
    Source field:adjudicated_in_scope
    59
    Source field:caveat
    The judge is a MODEL, not a human, and this is NOT a reference standard. It is an independent, blinded, documented fourth reading whose value is the SEAM CENSUS, not the tiers. Its tiers are used for nothing downstream and produce no estimate. The arm-preference test shows the judge is not CAPTURED; it says nothing about whether it is CORRECT, and correctness is exactly what the human audit exists to establish. A model cannot certify a model, which is the whole argument of this project and it applies to the judge too.
    adjudication · computed 2026-07-13T11:55:28Z
    25

    canadian_linkage_misnames_itself

    BOTH CLAUSES OF 'CANADIAN' MEASURE SOMETHING OTHER THAN WHAT THEY ARE NAMED. This project has spent all its effort on what retrieval MISSES; this is the first hard evidence about what it wrongly ADMITS. (1) CA-AFF, the PRIMARY estimand's main clause: 17,466 works enter the frame on a Canadian 'institution' that does not exist. OpenAlex hands 'Discovery Air (Canada)' to a Brazilian linguistics paper, 'Musee de la Civilisation' to a Brazilian food-science paper, and 'Impact', which is not an institution at all but a parse failure with an institution id, to a Spanish COVID essay. The CA-AFF route also carries thousands of works in Latvian and Indonesian. That count is a LOWER BOUND: the artifact strings tested are only the ones screening agents noticed BY EYE while doing something else, and nobody has swept the institution vocabulary, so the true precision is UNKNOWN rather than merely unmeasured. (2) ABOUT-CA, the SECONDARY estimand: the retrieval route means 'Canada appears in the text' and the rubric field means 'the Canadian research SYSTEM is a substantive object of study', and in the strata built entirely on the former, the latter fires on 1.8% to 3.5% of works. Worse, `about_ca` cannot be true unless the work is ALREADY about research, so it is nearly collinear with tier: only 39 works in 5,600 are both in-scope AND about the Canadian research system. THE SECONDARY ESTIMAND CANNOT BE ESTIMATED FROM THE STRATUM DESIGNED FOR IT. Four screening agents found this independently and all four proposed the same repair: split the field into `about_ca_system` and `about_ca_topic`, because one boolean cannot separate 'Canadian data, universal claim' from 'a claim about Canada'.

    Source field:ca_aff_works
    2,714,734
    Source field:ca_aff_only_institution_is_artifact
    17,466
    Source field:artifact_strings_tested
    • Impact
    • Discovery Air (Canada)
    • Musée de la Civilisation
    • Encana (Canada)
    • Kellogg's (Canada)
    • The Alberta Paraplegic Foundation
    Source field:artifact_count_is_a_lower_bound
    yes
    Source field:ca_aff_non_official_languages
    • 8,734
    • 8,463
    • 5,289
    • 5,160
    • 4,708
    • 1,200
    Source field:about_ca_pct_in_about_strata_opus
    • 1.8
    • 3.5
    Source field:about_ca_pct_in_about_strata_gpt
    • 1
    • 2
    Source field:about_ca_pct_in_about_strata_grok
    • 0.8
    • 1.2
    Source field:n_screened
    5,600
    Source field:about_ca_and_in_scope
    39
    Source field:about_ca_and_out_of_scope
    26
    Source field:about_ca_nearly_collinear_with_tier
    yes
    Source field:found_by
    Screening agents, reporting records they were given for an unrelated reason. Four of them independently reported that the ABOUT-CA route and the rubric's about_ca field are different constructs, and all four proposed the same repair without having seen each other's reports.
    Source field:caveat
    The artifact count is a LOWER BOUND, and deliberately reported as one: the strings tested are only those agents noticed by eye. No sweep of the OpenAlex institution vocabulary has been done, so the true CA-AFF precision is unknown, not merely unmeasured. That is the honest state and it is why the human audit samples the retrieved stratum as well as the non-retrieved one: precision is measured, not assumed.
    canadian_linkage_misnames_itself · computed 2026-07-13T11:55:27Z
    26

    classifier

    THE CLASSIFIER MAY NOT LABEL, AND THE NUMBERS SAY WHY. Design-weighted, out-of-fold, venue-grouped: the teacher heads reach AP 0.17 to 0.135 at AUC ~0.8, with a 4.68x top-decile lift that is real and is lift toward MACHINE labels. THE ESTIMAND MOVES THE SCORE 2.7x (unanimous 0.0783, majority 0.1364, any 0.2093), so there is no 'the' target and choosing one silently settles the question the audit exists to answer. AND THE UNANIMOUS CORE IS THE HARDEST TO LEARN, not the easiest: if the boundary were a shared line with noise around it, the agreed core would be the separable part, and it is not. The between-head spread predicts where the teachers split at 3.74x the base rate: real, weak, and reported as an efficiency gain rather than a map. Finally, the abstract is worth MORE THAN ANY MODELING CHOICE (+0.1145 AP, 84% relative), and the frame cannot serve it: works.csv stores has_abstract, a boolean, and no text. The DEPLOYABLE model is the WORSE one, the build now FAILS if a model is trained on a field inference cannot supply, and the gap is a priced budget decision rather than an assumption.

    Source field:evaluation
    out-of-fold, venue-grouped, vectorizer fitted on train folds only, DESIGN-WEIGHTED
    Source field:why_design_weighted
    The 5,600 works are a stratified sample: French carries weight 311, aff_core 1,119. An unweighted AP is an AP over a population that does not exist. Both are reported; the gap is a fact about the design.
    Source field:ap_design_weighted_by_teacher
    opus:
    0.17
    gpt:
    0.135
    grok:
    0.1417
    Source field:ap_design_weighted_by_consensus_rule
    unanimous:
    0.0783
    majority:
    0.1364
    any:
    0.2093
    Source field:estimand_moves_performance_x
    2.7
    Source field:unanimous_core_is_hardest_to_learn
    yes
    Source field:top_decile_lift
    opus:
    4.68
    gpt:
    3.93
    grok:
    5.28
    Source field:ap_by_language_opus
    en:
    0.1757
    fr:
    0.181
    Source field:ap_by_language_gpt
    en:
    0.1425
    fr:
    0.089
    Source field:french_is_model_dependent
    GPT's boundary is markedly less learnable in French (AP 0.089 vs 0.1425 in English) while Opus's is not ( 0.181 vs 0.1757 ). 'Character n-grams do the French work' was an assertion; measured per language, it is true for one teacher and false for another.
    Source field:predicting_teacher_disagreement
    n_contested:
    89
    base_rate:
    0.0159
    average_precision:
    0.0594
    lift_over_base:
    3.74
    verdict:
    Real but weak. It buys an efficiency gain for the audit's disagreement stratum. It is NOT a trustworthy standalone map of the contested region, and shared teacher bias is invisible to it: if all three are wrong the same way the spread is zero and the work looks settled.
    Source field:payload_skew
    ap_frame_parity:
    0.1364
    ap_with_abstract:
    0.2509
    abstract_is_worth_ap:
    0.1145
    why_the_deployable_model_is_the_worse_one:
    data/db/works.csv stores has_abstract, a boolean, and no abstract text. A model trained with abstracts scores a payload the 4.3M works do not have. Reporting the abstract model's number as the classifier's performance would be reporting a metric for a model that cannot run. The difference is what re-extracting abstracts from the snapshot would buy, and it is a budget decision with a number attached rather than an assumption.
    Source field:ships_a_category_label
    no
    Source field:why_no_label
    Distilled from our own metaresearch labels, a student reproduced its own teacher's positive set at Jaccard 0.17 (finding 25). Two independent adversarial reviews reached the same verdict from different directions: per-teacher heads are useful for ALLOCATING HUMAN EFFORT and cosmetic as a solution to the contested boundary. Scores ship. Labels do not.
    Source field:what_the_classifier_is_actually_for
    Not cost. Finding 13 says the LLM can read every work in the frame for $1,261, so there was never anything to avoid, and a proposal that budgeted the full screen while justifying a cheap substitute for it was describing two studies (D32). It is for: a calibrated score the audit stratifies on; rubric revisions testable in minutes instead of one $1,261 pass each; a frame-wide estimate of where the three teachers would split, which three full passes would cost 3x and ~25 days; and learnability as evidence about the boundary.
    classifier · computed 2026-07-15T22:32:09Z
    27

    distillation_ceiling

    A CLASSIFIER TRAINED ON LLM LABELS CANNOT EVEN REPRODUCE ITS OWN TEACHER, LET ALONE ADJUDICATE THREE. Students distilled from each model match their own teacher at Jaccard 0.17, while the three teachers match EACH OTHER at 0.5 to 0.56: the classifier is a worse approximation of Opus than Grok is. Distilling from Opus also moves you AWAY from Grok (-0.35), so distillation does not bridge the models' disagreement, it degrades away from all of them. THE SELF-TRAINING LOOP THEN REFUTED MY OWN PREDICTION. I argued at length that it would shrink the positive class toward the confident anglophone core; run on 20,000 works the models never saw, it did not: positive rate drift +0.00 points, French share drift +0.00 points, because class weighting prevents the collapse, exactly as the adversarial review warned before the run ('a serious risk, NOT A THEOREM'). What the loop DID do is reach a fixed point immediately and recycle its own labels (224 -> 424 training positives, then nothing): it CONVERGED WITHOUT LEARNING, and a fixed point looks identical from the inside whether its labels are right or wrong. What the classifier IS worth is the stratifier: 48.2% of the in-scope works sit in its top decile against a blind baseline of 10%, a 4.8x lift in in-scope works found per unit of human coding effort, and design-weighted estimates stay unbiased however noisy it is. THE MACHINE DECIDES WHERE THE HUMANS LOOK. IT NEVER SUPPLIES A LABEL.

    Source field:n_three_way_labelled
    5,600
    Source field:student_vs_own_teacher_jaccard
    • 0.176
    • 0.173
    • 0.17
    Source field:teacher_vs_teacher_jaccard
    • 0.5
    • 0.5
    • 0.563
    Source field:bridging_gain
    • -0.348297213622291
    • -0.348214285714286
    • -0.3375
    Source field:selftrain_rounds
    • 0
    • 1
    • 2
    • 3
    Source field:selftrain_train_positives
    • 224
    • 424
    • 424
    • 424
    Source field:selftrain_pool_positive_rate_pct
    • 1
    • 1
    • 1
    • 1
    Source field:selftrain_french_share_pct
    • 6.5
    • 6.5
    • 6.5
    • 6.5
    Source field:selftrain_drift_positive_rate_pts
    0
    Source field:selftrain_drift_french_share_pts
    0
    Source field:my_shrinkage_prediction_was_refuted
    yes
    Source field:stratifier_top_decile_recall_pct
    48.2
    Source field:stratifier_lift_vs_blind
    4.8
    Source field:used_to_produce_any_estimate
    no
    Source field:caveat
    The student is TF-IDF word+char n-grams with a cross-validated logistic head: the cheap classifier the question was actually about. A fine-tuned multilingual encoder would raise the student-vs-teacher number and change none of the argument, because the ceiling is set by the LABELS, not by the model class. The self-training loop's non-collapse is CONDITIONAL on class weighting and should not be read as a general safety result: it says the collapse is avoidable, not that the loop is informative. And nothing here produces a prevalence estimate; the classifier is a stratifier, and the human audit is the instrument.
    distillation_ceiling · computed 2026-07-15T22:58:15Z
    28

    gemma_gate

    THE FREE MODEL FAILED THE GATE, AND THE GATE WAS AIMED AT THE WRONG BAR. Gemma-4-31B screened 700 works under rubric v3.1 (100% schema-clean) and scored Jaccard 0.44 on metaresearch and 0.515 on any-category against the frontier union, under the prespecified 0.60 bar, SO IT DOES NOT LABEL THE FRAME: the rule stands because rules that move after the numbers exist are not rules. But the frontier models agree with each other at only 0.384 and 0.36 on the same works, because every loop batch is selected FOR disagreement: the gate demanded of a free model what the frontier models do not achieve among themselves there, and Gemma beat the frontier's own mutual agreement on both targets. Test 2 is prespecified on the random_baseline stratum with the rationale as the bar: J(gemma, union) >= J(opus, gpt).

    Source field:model
    google/gemma-4-31b-it (free)
    Source field:n_works
    700
    Source field:schema_clean_return_rate
    1
    Source field:jaccard_metaresearch_gemma_vs_union
    0.44
    Source field:jaccard_metaresearch_opus_vs_gpt
    0.384
    Source field:jaccard_anycat_gemma_vs_union
    0.515
    Source field:jaccard_anycat_opus_vs_gpt
    0.36
    Source field:study_design_exact
    gemma_vs_opus:
    0.643
    gemma_vs_gpt:
    0.614
    opus_vs_gpt_for_reference:
    0.73
    Source field:prespecified_rule
    Jaccard >= 0.60 vs the frontier union, on metaresearch AND any-category
    Source field:passed
    no
    Source field:rule_honored
    yes
    Source field:gate_was_miscalibrated_because
    the 0.60 bar was set from round-001's 0.44-0.75 pairwise range without registering that loop batches are selected FOR disagreement, which depresses every agreement statistic computed on them. On the same 700 works the frontier models agree with each other at 0.384/0.360: below the bar Gemma was held to, and Gemma beat both numbers.
    Source field:test2_prespecified
    On the enriched sample's random_baseline stratum (1,500 works, not selected for disagreement), Gemma passes if J(gemma, frontier_union) >= J(opus, gpt) on both metaresearch and any-category. Written before any of those labels exist.
    gemma_gate · computed 2026-07-14T00:08:32Z
    29

    instrument_contradicts_itself

    THE LOCKED INSTRUMENT CONTRADICTS ITSELF, AND NOTHING CAUGHT IT. The rubric and the output schema name TWO DIFFERENT controlled vocabularies for the same field, `genre`, overlapping on 2 values (empirical and other); 4 terms exist only in the rubric and 7 only in the schema. Every screener was handed both and told to obey both, and each invented its own reconciliation: GPT-5.6 and Grok followed the rubric and are therefore 26.9% and 20.9% ILLEGAL against the schema, while Opus drew from both lists at once and emitted 13 distinct values. The validator checked `tier` and never checked `genre`, so 16,800 labels passed every check that ran. It was found by an agent mentioning it in one clause of a report about something else. Nothing crashed; the variable simply meant a different thing in each arm, and the variance would have been attributed to the models. The labels are NOT repaired, because harmonizing the arms after seeing them would destroy the only evidence that they diverged; the genre field is reported as unusable and the fix belongs in the instrument, at a version boundary.

    Source field:field
    genre
    Source field:rubric_vocabulary
    • empirical
    • conceptual
    • editorial/commentary
    • policy
    • infrastructure/announcement
    • other
    Source field:schema_vocabulary
    • empirical
    • review
    • methods
    • commentary
    • editorial
    • protocol
    • dataset
    • software
    • other
    Source field:shared_values
    • empirical
    • other
    Source field:n_shared
    2
    Source field:only_in_rubric
    • conceptual
    • editorial/commentary
    • policy
    • infrastructure/announcement
    Source field:only_in_schema
    • review
    • methods
    • commentary
    • editorial
    • protocol
    • dataset
    • software
    Source field:arms
    • opus
    • gpt
    • grok
    Source field:n_labels
    • 5,600
    • 5,600
    • 5,600
    Source field:n_distinct_values_by_arm
    • 13
    • 6
    • 12
    Source field:pct_legal_against_schema
    • 90.8
    • 73.1
    • 79.1
    Source field:pct_legal_against_rubric
    • 86.1
    • 100
    • 99.8
    Source field:validator_checked_tier
    yes
    Source field:validator_checked_genre
    no
    Source field:labels_repaired
    no
    Source field:found_by
    An Opus screening agent mentioned it in one clause of a report about something else, while working chunks it had been given for an unrelated reason. It was not looking for this, no check was watching for it, and it had already survived 6,000 labels across three models.
    Source field:caveat
    The genre variable from this screen is reported as UNUSABLE and is used for nothing. It is not remapped to a common vocabulary: the three arms resolved the contradiction differently, and harmonizing them after the fact would destroy the only evidence that they did.
    instrument_contradicts_itself · computed 2026-07-15T22:32:11Z
    30

    the_frame

    The frame is BUILT, not estimated. All 482 partitions of a pinned OpenAlex snapshot were streamed and filtered, and the Canadian frame holds 4,299,418 works, each exactly once. For most of this project's life 'the frame' was 3,507,205, an EXTRAPOLATION from a single partition, hardcoded in six scripts; the estimate was 18% LOW, and every cost, audit-power and field-size figure computed against it has moved. The route that matters: 1,565,226 works (36.4% of the frame) are INVISIBLE TO AFFILIATION ALONE. A frame built on Canadian affiliation would hold 2,734,192 works and would never see them. That is finding 2 at frame scale, and it is why the frame is a union of four routes and why every record carries the provenance of the route that admitted it.

    Source field:built_from
    all 482 partitions of a pinned OpenAlex snapshot, streamed and filtered; each work appears exactly once
    Source field:partitions
    482
    Source field:works
    4,299,418
    Source field:prior_estimate
    3,507,205
    Source field:estimate_was_low_by_pct
    18
    Source field:estimate_was_an_extrapolation_from_one_partition
    yes
    Source field:route_ca_aff
    2,734,192
    Source field:route_about_ca
    1,316,408
    Source field:route_ca_fund
    788,103
    Source field:route_ca_venue
    638,161
    Source field:invisible_to_affiliation
    1,565,226
    Source field:pct_invisible_to_affiliation
    36.4
    Source field:with_venue
    3,902,048
    Source field:pct_with_venue
    90.8
    Source field:with_abstract
    3,296,301
    Source field:pct_with_abstract
    76.7
    Source field:pct_no_abstract
    23.3
    Source field:french
    237,207
    Source field:pct_french
    5.5
    Source field:built_utc
    2026-07-12T17:59:59Z
    Source field:caveat
    The frame is bounded by OpenAlex. A work whose Canadian link is invisible to the metadata (no affiliation, no funder, no textual mention, no Canadian venue) cannot enter ANY frame by any method, and scholarship indexed by neither OpenAlex nor Erudit is not estimated here. That is a hard boundary of the data, not of this design, and it is stated rather than hidden. Erudit matches ZERO OpenAlex sources (finding 3), so the francophone share reported here (5.5%) measures the pipeline, NOT Canadian scholarship.
    the_frame · computed 2026-07-13T11:55:21Z
    31

    v1_to_v2

    WRITING THE MISSING SENTENCES RESOLVED 54% OF THE DISAGREEMENTS THEY WERE WRITTEN AGAINST. The 179 works three frontier models split on under rubric v1 were re-screened, by the same three models, under a v2 whose rules were written against a published, frequency-weighted census of WHICH SENTENCES THE RUBRIC WAS MISSING. 96 of 179 are now unanimous, and the pairwise Jaccard of the in-scope sets moves from 0.19 to 0.36. This is a measurement almost nobody makes, and it is available ONLY because the ambiguity was enumerated, adjudicated and PUBLISHED BEFORE it was resolved: the census went out first, so v2 could not be quietly tuned until the numbers looked good. SCOPE, stated rather than blurred: these works were SELECTED FOR DISAGREEMENT, so this says the rules resolve the cases they were written for; it does NOT say frame-wide agreement rose, because regression to the mean alone would move a set selected this way. The unbiased test is a full re-screen of all 5,600 under v2, prespecified, and it is what the funded work runs.

    Source field:contested_under_v1
    179
    Source field:unanimous_under_v2
    96
    Source field:pct_resolved
    54
    Source field:still_contested
    83
    Source field:in_scope_by_model_v1
    • 129
    • 68
    • 53
    Source field:in_scope_by_model_v2
    • 73
    • 32
    • 44
    Source field:jaccard_v1_mean
    0.187
    Source field:jaccard_v2_mean
    0.362
    Source field:sequence
    lock v1; screen 5,600 works with three models; have an independent blinded judge adjudicate every disagreement AND name the seam that caused it; PUBLISH the seam census; write v2 against the census and lock it BEFORE re-screening; re-screen; report the difference, whatever it is. Publishing the census before resolving it is what stops v2 being tuned until the numbers improve.
    Source field:caveat
    SELECTED FOR DISAGREEMENT. These 179 works are exactly the ones the three models split on under v1, so this measures whether v2's rules resolve the cases they were written for. It CANNOT support a claim about agreement across the frame: regression to the mean alone would move a set selected this way. The unbiased test is a full re-screen of all 5,600 under v2, prespecified, and it is what the funded work runs.
    v1_to_v2 · computed 2026-07-13T11:55:28Z
    32

    zero_probability_region

    THE SAMPLING DESIGN COULD NOT REACH 12.9% OF THE FRAME, AND NO WEIGHT CAN FIX THAT. A stratified design rests on one identity, sum(N_h) = N. It was never checked. 549,370 works (12.9% of the sampling frame) had an inclusion probability of EXACTLY ZERO: 366,856 that no predicate claimed, because aff_core excluded everything ABOUT Canada while about_only excluded everything AFFILIATED with Canada, so a work that was both fell between them; and 182,514 more whose predicate evaluated to SQL NULL, which `WHERE p` and `WHERE NOT p` BOTH decline to select, so they were in no stratum and were not even orphans. Nothing threw. Every stratum returned exactly the n it asked for, because a stratum cannot know about the works it was never asked about, and the design drew a clean textbook probability sample OF 87% OF THE FRAME while every number computed from it said 'the frame'. A zero-probability work is not underweighted, it is UNREACHABLE, and design weights are the guarantee this project leans on hardest. THE CELL IT DELETED WAS THE SECONDARY ESTIMAND'S: 328,912 Canadian-affiliated works about Canada, in a study whose secondary estimand is 'metaresearch about the Canadian research system'. Repaired by adding the exact NULL-safe complement as two strata; the 7 strata now partition the frame by construction (4,255,410 = 4,255,410), asserted on every build. It was found by an adversarial model adding up five numbers I handed it. The defect made the sample TIDIER and the variance SMALLER, which is why nothing about it felt wrong: the errors that survive are the ones that flatter you.

    Source field:full_frame
    4,299,418
    Source field:unscreenable_excluded
    44,008
    Source field:sampling_frame
    4,255,410
    Source field:reachable_under_shipped_design
    3,706,040
    Source field:orphaned_predicate_false
    366,856
    Source field:invisible_predicate_null
    182,514
    Source field:zero_probability_works
    549,370
    Source field:pct_of_sampling_frame
    12.9
    Source field:aff_and_about_cell
    328,912
    Source field:strata_after_repair
    7
    Source field:strata_sum_after_repair
    4,255,410
    Source field:partition_holds
    yes
    Source field:found_by
    An adversarial model asked to attack the classifier design. Its first move was to add up the five design weights in a summary table I had handed it: 2000x1119 + 1000x664.2 + 750x536.8 + 750x310.9 + 500x335.8 = 3,705,875, against a frame of 4,299,418. I had never added them up.
    Source field:caveat
    The repair does not retro-fix numbers produced under the broken design; those estimated a 3.7M subpopulation and were reported under the frame's name, and saying so IS the finding. Every design-weighted figure is now re-derived against the seven-stratum design. The five original predicates are kept byte-for-byte because 5,000 works had already been drawn from them by hash order, and widening a stratum silently re-draws it.
    zero_probability_region · computed 2026-07-13T11:55:25Z