MétaCan
Menu
Constats

Ce que le pilote a mesuré.

Les 32 constats, rendus directement depuis pilot/results/findings.json, le fichier qu'écrivent les scripts du pilote. Aucun nombre de cette page n'a été saisi par un humain : c'est la seule façon de garantir que le site et l'analyse ne peuvent pas diverger.

le même fichier par l'API →

01

L'écart d'affiliation

Dans l'espace thématique de la métarecherche, 508 744 travaux sur 793 883 (64 %) ne portent aucune chaîne d'affiliation brute dans OpenAlex.

Énoncé original (findings.json) : 508,744 of 793,883 works (64%) in the metaresearch topic space have no raw affiliation strings in OpenAlex.

Champ source:topic_space_total
793 883
Champ source:with_raw_affiliation
285 139
Champ source:without_raw_affiliation
508 744
Champ source:pct_without
64,1
affiliation_gap · calculé le 2026-07-13T11:54:42Z
02

La voie thématique

Sur 4 516 thématiques OpenAlex, 0 ne nomme la métarecherche comme domaine. Les 11 thématiques qui portent bel et bien du contenu métarecherche sont dispersées dans 7 domaines OpenAlex différents.

Énoncé original (findings.json) : Of 4,516 OpenAlex topics, 0 name metaresearch as a field. The 11 topics that do carry metaresearch content are scattered across 7 different OpenAlex fields.

Champ source:n_topics_in_taxonomy
4 516
Champ source:n_topics_naming_field
0
Champ source:topics_naming_field
    Champ source:n_candidate_topics
    11
    Champ source:n_fields_spanned
    7
    Champ source:fields_spanned
    • Arts and Humanities
    • Computer Science
    • Decision Sciences
    • Mathematics
    • Medicine
    • Psychology
    • Social Sciences
    Champ source:candidate_topic_ids
    • T10102
    • T13607
    • T13516
    • T11937
    • T10206
    • T10582
    • T10267
    • T10778
    • T13558
    • T13284
    • T11875
    topics · calculé le 2026-07-13T11:54:42Z
    03

    La polysémie met le lexique en échec

    Le seul terme reproducibility repère 43 392 travaux canadiens, dont seulement 0,8 % relèvent de l'espace thématique de la métarecherche. Le repérage par mots-clés ne peut pas séparer le sens métascientifique du sens courant.

    Énoncé original (findings.json) : The single term reproducibility retrieves 43,392 Canadian works, of which only 0.8% fall in the metaresearch topic space. Keyword retrieval cannot separate the metaresearch sense from the everyday one.

    Champ source:hits_alone
    reproducibility:
    43 392
    "peer review":
    21 234
    "open access":
    6 609
    "open science":
    1 925
    Champ source:hits_alone_and_on_topic
    reproducibility:
    362
    "peer review":
    795
    "open access":
    634
    "open science":
    325
    Champ source:topic_space_precision_pct
    reproducibility:
    0,8
    "peer review":
    3,7
    "open access":
    9,6
    "open science":
    16,9
    Champ source:disciplined_lexicon_hits
    8 026
    Champ source:worst_term
    reproducibility
    polysemy · calculé le 2026-07-13T11:54:44Z
    04

    L'écart linguistique

    Le français représente 2,7 % (395/14 873) de la métarecherche canadienne dans OpenAlex. Un lexique français dédié trouve 168 travaux canadiens, contre 8 026 en anglais.

    Énoncé original (findings.json) : French is 2.7% (395/14,873) of Canadian metaresearch in OpenAlex. A dedicated French lexicon finds 168 Canadian works, against 8,026 in English.

    Champ source:canadian_topic_works
    14 873
    Champ source:n_english
    14 028
    Champ source:n_french
    395
    Champ source:pct_french
    2,7
    Champ source:en_lexicon_canadian_hits
    8 026
    Champ source:fr_lexicon_canadian_hits
    168
    Champ source:fr_lexicon_world_hits
    4 682
    Champ source:canada_share_of_world_french
    3,6
    language_gap · calculé le 2026-07-13T11:54:43Z
    05

    Érudit est invisible pour OpenAlex

    Érudit ne correspond à aucune source dans OpenAlex (0), mais son point d'accès OAI-PMH est actif et expose 379 ensembles moissonnables.

    Énoncé original (findings.json) : Erudit matches 0 sources in OpenAlex, but its OAI-PMH endpoint is live and exposes 379 harvestable sets.

    Champ source:openalex_sources_matching_erudit
    0
    Champ source:oai_endpoint
    https://oai.erudit.org/oai/request
    Champ source:oai_repository_name
    Erudit
    Champ source:oai_earliest_datestamp
    2011-06-03
    Champ source:oai_harvestable_sets
    379
    erudit · calculé le 2026-07-13T11:54:43Z
    06

    La capture-recapture est ici sans valeur

    La capture-recapture naïve à deux voies estime 467 541 travaux canadiens de métarecherche, ce qui impliquerait que le Canada produit 59 % de la métarecherche mondiale contre 1,9 % observé. L'estimateur est ici sans valeur ; nous l'avons supprimé plutôt que de le maquiller en borne inférieure.

    Énoncé original (findings.json) : Naive two-route capture-recapture estimates 467,541 Canadian metaresearch works, implying Canada produces 59% of the world's metaresearch against an observed 1.9%. The estimator is void here; we cut it rather than dress it up as a lower bound.

    Champ source:route1_topic
    14 873
    Champ source:route2_naive_lexical
    77 583
    Champ source:overlap
    2 468
    Champ source:observed_union
    89 988
    Champ source:lincoln_petersen_estimate
    467 541
    Champ source:entire_topic_space_all_countries
    793 883
    Champ source:canada_observed_share_pct
    1,9
    Champ source:canada_implied_share_pct
    58,9
    Champ source:estimator_void
    oui
    capture_recapture_fails · calculé le 2026-07-13T11:54:44Z
    07

    Le tri à trois modèles

    Trois modèles de pointe, Opus 4.8, GPT-5.6 high et Grok 4.5, ont trié les mêmes 5 600 travaux tirés de la vraie base de 4,3 millions selon un plan dont les sept strates la PARTITIONNENT ; un plan antérieur à cinq strates ne pouvait atteindre 12,9 % de la base (D22). Les taux de base pondérés par le plan s'étendent de 2,54 % à 3,81 %, soit un rapport de 1,5. Mais le TAUX n'est pas le constat ; ce sont les ENSEMBLES. Parmi les 274 travaux qu'AU MOINS UN modèle a qualifiés de métarecherche, seulement 104, soit 38 %, l'ont été par LES TROIS, et 117, soit 43 %, reposent sur l'avis d'un seul modèle ; l'indice de Jaccard par paire des ensembles dans le champ avoisine 50 %. LA FRONTIÈRE DU DOMAINE N'EST PAS UNE LIGNE QUE LES MODÈLES PARTAGENT ; C'EST UNE RÉGION QUE CHACUN DÉCOUPE À SA FAÇON. Ce résultat reste STABLE à n = 1 000, 2 000 et 5 600, avec 37 %, 37 % et 38 % d'unanimité. UNE SECONDE AFFIRMATION, PLUS SÉDUISANTE, N'A PAS SURVÉCU : à n = 2 000, les modèles semblaient s'accorder nettement davantage sur « ce travail porte-t-il sur la recherche ? » que sur « est-il dans le champ ? », avec des écarts de 1,43 et 1,51 ; à n = 5 600, les deux écarts ne se distinguent plus du bruit, avec un rapport de 1,06, et l'affirmation est RETIRÉE (D23). La plus grande confusion entre niveaux reste OUT contre T2, à chaque taille d'échantillon : ce sont les traditions adjacentes que le critère d'inclusivité doit protéger. Le livrable n'est pas un taux de base ; c'est le dossier des désaccords, soit les 274 travaux qui marquent la frontière empirique, chacun accompagné des raisons données par les trois modèles, et les critères qui doivent être rédigés à partir d'eux.

    Énoncé original (findings.json) : Three frontier models (Opus 4.8, GPT-5.6 high, Grok 4.5) screened the same 5,600 works, drawn from the real 4.3M frame under a design whose seven strata PARTITION it (an earlier five-stratum design could not reach 12.9% of the frame at all; D22). Design-weighted base rates span 2.54% to 3.81% (1.5x). But the RATE is not the finding, the SETS are: of the 274 works ANY model called metaresearch, only 104 (38%) were called metaresearch by ALL THREE, and 117 (43%) rest on a SINGLE model's opinion; pairwise Jaccard on the in-scope sets is about 50%. THE FIELD'S BOUNDARY IS NOT A LINE THE MODELS SHARE; IT IS A REGION THEY EACH CUT DIFFERENTLY, and that result is STABLE across n = 1,000, 2,000 and 5,600 (unanimity 37%, 37%, 38%). A SECOND, PRETTIER CLAIM DID NOT SURVIVE: at n = 2,000 the models agreed markedly more on 'is this about research at all' (1.43x here) than on 'is it in scope' (1.51x here), and this project said so in capitals; at n = 5,600 the two spreads are within noise (ratio 1.06) and the claim is WITHDRAWN (D23). The largest tier confusion is OUT-vs-T2, every time, at every sample size: the adjacent traditions the inclusiveness criterion exists to protect. The deliverable is not a base rate. It is the disagreement dossier, the 274 works that mark the empirical boundary, each carrying all three models' stated reasons, and the criteria that have to be written against them.

    Champ source:frame
    the real 4.3M-work Canadian frame (all 482 OpenAlex partitions)
    Champ source:payload
    the rubric's FULL eight fields, including venue (repairs D1)
    Champ source:sample
    5,600 works across seven exhaustive strata, with known selection probabilities and French oversampled
    Champ source:models
    Claude Opus 4.8; GPT-5.6 (high effort); Grok 4.5 (medium effort)
    Champ source:harness
    chunks randomized and manifest-logged before any model ran; the harness writes label files, never the model (repairs D11); every arm reconciled against the manifest (repairs D2)
    Champ source:n_labelled_by_all_three
    5 600
    Champ source:tranche_homogeneity_p
    0,63
    Champ source:tranches_pool
    oui
    Champ source:base_rate_weighted_pct
    opus:
    3,81
    gpt:
    2,92
    grok:
    2,54
    Champ source:between_model_spread_x
    1,5
    Champ source:jaccard_opus_gpt
    50
    Champ source:jaccard_opus_grok
    50
    Champ source:jaccard_gpt_grok
    56
    Champ source:n_about_research_at_all
    opus:
    391
    gpt:
    325
    grok:
    274
    Champ source:n_in_scope
    opus:
    224
    gpt:
    163
    grok:
    148
    Champ source:spread_about_research_x
    1,43
    Champ source:spread_in_scope_x
    1,51
    Champ source:variance_is_in_the_rubric_not_the_models
    oui
    Champ source:called_in_scope_by_any
    274
    Champ source:unanimous_in_scope
    104
    Champ source:pct_unanimous_of_any
    38
    Champ source:in_scope_by_one_model_only
    117
    Champ source:pct_single_model_of_any
    43
    Champ source:contested_by_stratum
    about_only:
    67
    venue_new:
    65
    residual:
    64
    aff_core:
    62
    fund_new:
    61
    aff_about:
    59
    french:
    54
    Champ source:tier_disagreement_patterns
    OUT/T2:
    77
    T1:
    65
    OUT/T1:
    57
    T2:
    30
    T1/T3:
    11
    T2/T3:
    10
    T1/T2:
    9
    OUT/T1/T2:
    8
    OUT/T1/T3:
    3
    OUT/T2/T3:
    3
    Champ source:gpt_schema_violations_first_pass
    18
    Champ source:gpt_violation_note
    GPT-5.6 (high) wrote GENRE values ('empirical', 'conceptual') into the TIER field on 18 of 1,000 records in its first pass, in 3 of 20 chunks. The validator caught it because the harness reconciles files against a manifest rather than trusting the model's report. Those chunks were RE-RUN, not repaired: coercing a model's output to the schema is fitting the instrument to the data.
    Champ source:deliverable
    pilot/screening/frame1k/disagreement_dossier.json: every work any model called in-scope, with all three labels. This, not the base rate, is what the criteria must be written against.
    Champ source:caveat
    These are MACHINE labels and none of them is truth (finding 15). The unanimity rate is not accuracy: three models sharing training data can be wrong together, and they are most correlated exactly on the boundary cases the field's definition turns on. What this measures is where the RUBRIC is underspecified, which is a property of the instrument and is exactly what a criteria document needs. Base rates are design-weighted from a stratified sample, so they estimate the frame; the Jaccard and unanimity figures are unweighted set quantities over the sample and are NOT frame estimates.
    three_model_screen · calculé le 2026-07-15T22:59:18Z
    08

    Changez de trieur, la réponse bouge

    Changez le modèle qu'on appelle « le trieur » et le taux de base passe de 1,06 % à 2,37 % : un écart de 2,2 fois, soit de 45 397 à 101 776 travaux dans la base. Les deux trieurs s'accordent sur inclus ou exclu pour 98,4 % de la base après pondération du plan, mais ce chiffre est dominé par les rejets établis : l'accord tombe à 95 % à l'intérieur de la frontière contestée. L'ÉTENDUE OBTENUE EN CHANGEANT DE TRIEUR, ET NON L'IC BINOMIAL D'UN SEUL MODÈLE, EST L'INCERTITUDE HONNÊTE SUR LA TAILLE DU DOMAINE.

    Énoncé original (findings.json) : Swap which model is called 'the screener' and the base rate moves from 1.06% to 2.37%: a 2.2x spread, from 45,397 to 101,776 works in the frame. The two screeners agree on in/out for 98.4% of the frame (design-weighted), but that figure is dominated by the settled rejects: agreement falls to 95% inside the contested boundary. THE SCREENER-SWAP RANGE, NOT THE BINOMIAL CI ON EITHER MODEL ALONE, IS THE HONEST UNCERTAINTY ON THE FIELD'S SIZE.

    Champ source:n_double_screened
    1 290
    Champ source:base_rate_screener_a_pct
    1,06
    Champ source:base_rate_screener_b_pct
    2,37
    Champ source:swap_ratio_x
    2,24
    Champ source:field_size_screener_a
    45 397
    Champ source:field_size_screener_b
    101 776
    Champ source:published_binomial_ci_contains_b
    non
    Champ source:screener_a
    claude-sonnet-4-6 (40 agents, medium effort)
    Champ source:screener_b
    gpt-5.6-sol (codex)
    Champ source:sampling
    stratified on screener A's label; positives/paratext/boundary taken whole, settled-OUT sampled
    Champ source:raw_agreement_inout_pct
    96,6
    Champ source:weighted_agreement_pct
    98,4
    Champ source:cohens_kappa_inout
    0,681
    Champ source:n_disagree_inout
    44
    Champ source:gpt_in_claude_out
    37
    Champ source:claude_in_gpt_out
    7
    Champ source:agreement_by_stratum
    stratum:
    settled_out,boundary,positive,paratext
    n:
    600,599,58,33
    sel_prob:
    0.124921923797626,1,1,1
    agree_pct:
    99,95,87.9,97
    a_says_in:
    0,0,58,0
    b_says_in:
    6,30,51,1
    Champ source:caveat
    PROCESS METRIC, NOT ACCURACY. Two LLMs share training data and failure modes; their errors are correlated and most correlated on the boundary. This is not 'duplicate screening': that term's warrant comes from independent human judgment. Accuracy rests on the human-coded probability sample (PROTOCOL s6.2). Agreement is reported per stratum because a pooled kappa on a 1.3% base rate is dominated by the cell where agreement is free (the kappa paradox). And note what the high agreement figure CONCEALS: it is dominated by the settled-OUT mass, while the two screeners imply base rates a factor of two apart. Quoting agreement without the swap would be presenting the reassuring statistic.
    agreement · calculé le 2026-07-15T22:52:08Z
    09

    Le taux de base

    Le tri de 5 737 travaux canadiens non filtrés selon la grille situe la métarecherche à 1,31 % de la recherche canadienne, soit environ 56 206 travaux dans la base de 4 299 418, ce qui dimensionne le domaine sans aucune stratégie de recherche. L'IC à 95 % sur les étiquettes de ce trieur va de 1,03 à 1,64 %, mais c'est de l'erreur d'échantillonnage, PAS l'incertitude : changez de trieur et l'estimation tombe hors de l'intervalle (constat 10). L'étendue produite par le changement de trieur au constat 10 représente l'incertitude honnête sur la taille du domaine, et seul l'audit humain peut la resserrer.

    Énoncé original (findings.json) : Screening 5,737 unfiltered Canadian works against the rubric puts metaresearch at 1.31% of Canadian research, implying ~56,206 works in the 4,299,418-work frame, sizing the field without a search strategy at all. The 95% CI on this screener's labels is 1.03-1.64%, but that is sampling error, NOT the uncertainty: swap the screener and the estimate lands outside it (finding 10). The machine-screener range in finding 10 is the honest uncertainty on field size, and only the human audit can narrow it.

    Champ source:n_screened
    5 737
    Champ source:screener
    claude-sonnet-4-6, 40 agents, medium effort, locked rubric
    Champ source:sampling_frame
    unfiltered Canadian works from the pinned 2026-06-24 snapshot partition
    Champ source:tier_counts
    OUT:
    5 621
    T1:
    40
    T2:
    35
    T3:
    41
    Champ source:n_in_scope_t1_t2
    75
    Champ source:base_rate_pct
    1,31
    Champ source:base_rate_ci_lo_pct
    1,03
    Champ source:base_rate_ci_hi_pct
    1,64
    Champ source:canadian_frame_size
    4 299 418
    Champ source:estimated_field_size
    56 206
    Champ source:estimated_field_lo
    44 268
    Champ source:estimated_field_hi
    70 338
    Champ source:topic_route_retrieved
    14 873
    Champ source:binomial_ci_is_not_the_uncertainty
    oui
    Champ source:caveat
    Machine labels, not a human gold standard, and the binomial CI above is sampling error on ONE screener; the honest uncertainty is the screener-swap range in finding 10 (1.06% to 2.37%). The partition is also not a uniform draw: it under-represents works with abstracts, where the screen finds 2x more metaresearch (finding 11). This is a hypothesis with a denominator; the human-coded probability sample tests it. Do NOT divide topic_route_retrieved by estimated_field_size; see finding 12.
    base_rate · calculé le 2026-07-15T22:33:57Z
    10

    La robustesse du taux de base

    Le vrai biais du taux de base n'est pas la récence mais les RÉSUMÉS MANQUANTS : 31,5 % de la partition n'en a aucun, et le tri y repère 0,78 % de métarecherche contre 1,55 % là où un résumé existe (khi carré p = 0,023, robuste à l'ajustement pour l'année et la langue). Le tiers de la base est trié sur son seul titre. Ce que cela ne montre PAS, et qu'une version antérieure affirmait à tort, c'est que l'aveuglement serait DIFFÉRENTIEL selon la tradition : la case T2 sans résumé contient 4 travaux et l'interaction n'est pas significative (p = 0,141). Cette affirmation est retirée, comme l'est « le portrait exact d'Érudit » (la strate est anglophone à 99 % et les travaux sont plus RÉCENTS, non plus anciens). L'effet principal est le constat.

    Énoncé original (findings.json) : The base rate's real bias is not recency but MISSING ABSTRACTS: 31.5% of the partition has none, and the screen finds 0.78% metaresearch there against 1.55% where an abstract exists (chi-square p = 0.023, robust to adjustment for year and language). A third of the frame is screened on its title alone. What this does NOT show, and an earlier draft wrongly claimed, is that the blindness is DIFFERENTIAL by tradition: the T2 no-abstract cell holds 4 works and the interaction is not significant (p = 0.141). That claim is withdrawn, as is 'Erudit's exact profile' (the stratum is 99% English and the works are NEWER, not older). The main effect is the finding.

    Champ source:partition
    updated_date=2026-06-24
    Champ source:records_sent_to_screener
    6 202
    Champ source:records_silently_lost
    465
    Champ source:records_lost_pct
    7,5
    Champ source:pct_no_abstract_among_lost
    40,2
    Champ source:pct_no_abstract_among_labelled
    31,5
    Champ source:chisq_p_loss_bias
    0,00013
    Champ source:losses_are_biased
    oui
    Champ source:works_2000_09
    2 704
    Champ source:works_2020_25
    1 195
    Champ source:old_to_new_ratio
    2,3
    Champ source:partition_skews_old
    oui
    Champ source:base_rate_by_era_pct
    2000-09:
    1,24
    2010-19:
    1,35
    2020-25:
    1,4
    Champ source:chisq_p_era
    0,906
    Champ source:era_events
    75
    Champ source:era_check_is_underpowered
    oui
    Champ source:no_abstract_share_pct
    31,5
    Champ source:base_rate_no_abstract_pct
    0,78
    Champ source:base_rate_has_abstract_pct
    1,55
    Champ source:chisq_p_abstract
    0,023
    Champ source:abstract_effect_x
    2
    Champ source:t1_penalty_x
    1,4
    Champ source:t2_penalty_x
    3,6
    Champ source:t2_no_abstract_cell_count
    4
    Champ source:interaction_p
    0,141
    Champ source:differential_is_supported
    non
    Champ source:differential_claim_withdrawn
    oui
    Champ source:no_abstract_mean_year
    2013,7
    Champ source:has_abstract_mean_year
    2010,3
    Champ source:no_abstract_stratum_pct_english
    99
    Champ source:p_no_abstract_given_english_pct
    32
    Champ source:p_no_abstract_given_non_english_pct
    12,7
    Champ source:erudit_profile_claim_holds
    non
    Champ source:residual_bias_direction
    anti-conservative for coverage claims: the partition over-represents abstract-less works (31.5%), where the screen finds 2x less metaresearch, so 1.31% likely UNDER-states the frame's base rate, and the field is larger than the headline implies
    Champ source:supersedes
    THREE retractions live here. (1) An earlier version tested ERA only, called the base rate robust, and published 3100/2900/1500 as percentages (a dplyr summarise() column-masking bug). (2) It then claimed the blindness is DIFFERENTIAL, T2 losing 3.6x against T1's 1.4x, and made that the proposal's whole answer to the inclusiveness criterion. The T2 no-abstract cell holds FOUR works and the interaction is not significant (p = 0.141). WITHDRAWN. (3) It claimed missing abstracts track 'older, non-English' records, 'Erudit's exact profile'. Backwards: they are NEWER, and the stratum is 99% ENGLISH. WITHDRAWN. See DEVIATIONS.md D4, D5, D6.
    Champ source:caveat
    What survives is the MAIN EFFECT and only the main effect: a third of the frame is screened on its title alone and the screen finds half as much metaresearch there (p = 0.023, robust to adjustment for year and language). That is a real coverage problem and a reason to stratify the audit on abstract availability. It is NOT evidence of differential blindness by tradition, and this finding no longer says it is. Separately, the era check is underpowered (75 events, 3 strata) and cannot refute an era effect; it merely fails to detect one. And note D2: the harness silently dropped 465 records, non-randomly, on this very covariate.
    base_rate_robustness · calculé le 2026-07-13T11:54:47Z
    11

    Ce que les étiquettes ne peuvent pas dire

    Deux limites des étiquettes machines, trouvées en attaquant les correctifs. (A) Le rappel de la voie thématique est de 12 % selon le trieur A et de 7 % selon le trieur B : l'instrument (ii) du constat 14 évalue les filtres sur des étiquettes de MACHINE, il mesure donc l'accord avec une machine et non l'exactitude, et le constat 10 avait déjà montré que cela varie du simple au double. La conclusion se renforce (le second trieur juge la voie encore PIRE), mais 12 % n'est pas la vérité. (B) Le pilote contient exactement 1 travail francophone dans le champ ; en voir 20 exigerait environ 1 620 notices françaises codées contre un budget d'audit de 1 000. LA SENSIBILITÉ POUR LE FRANÇAIS N'EST PAS ESTIMABLE DANS UNE BASE FONDÉE SUR LE SEUL OPENALEX. C'est un argument pour le moissonnage d'Érudit, non contre la revendication du français, mais le pilote n'a pas exécuté ce moissonnage : la puissance est donc énoncée comme une condition plutôt que comme une promesse.

    Énoncé original (findings.json) : Two limits on the machine labels, found by attacking the fixes. (A) The topic route's recall is 12% against screener A and 7% against screener B: finding 14's instrument (ii) scores filters against MACHINE labels, so it measures agreement with a machine, not accuracy, and finding 10 already showed that swings by a factor of two. The conclusion strengthens (the second screener thinks the route is WORSE) but 12% is not truth. (B) The pilot holds exactly 1 French in-scope work, so seeing 20 French positives needs ~1,620 coded French records against an audit budget of 1,000. FRENCH SENSITIVITY IS NOT ESTIMABLE IN AN OPENALEX-ONLY FRAME. That is an argument for the Erudit harvest, not against the French claim, but the pilot did not run that harvest, so the power is stated as a condition rather than a promise.

    Champ source:route_recall_vs_screener_a_pct
    12
    Champ source:route_recall_vs_screener_b_pct
    7
    Champ source:route_recall_a_ci
    • 5,6
    • 21,6
    Champ source:route_recall_b_ci
    • 2,5
    • 14,3
    Champ source:positives_screener_a
    75
    Champ source:positives_screener_b
    88
    Champ source:instrument_ii_is_model_dependent
    oui
    Champ source:instrument_ii_caveat
    Finding 14's instrument (ii) scores a filter against the 5,737 MACHINE labels. That measures agreement with a machine, not accuracy. Swap the machine and the topic route's recall moves from 12% to 7%. The conclusion (the route finds a small fraction) survives and strengthens; the NUMBER is not a measurement against truth.
    Champ source:french_records_in_pilot
    81
    Champ source:french_in_scope_in_pilot
    1
    Champ source:french_records_needed_for_20_positives
    1 620
    Champ source:audit_budget_records
    1 000
    Champ source:french_stratum_is_powered
    non
    Champ source:french_power_depends_on
    the Erudit harvest, which the pilot did NOT run (it verified the endpoint: 379 live sets)
    Champ source:caveat
    (A) is a limit on every recall number this project quotes against machine labels, including its own headline. (B) is a limit on the inclusiveness promise: French sensitivity cannot be estimated in an OpenAlex-only frame, because Erudit matches zero OpenAlex sources and the francophone literature is therefore largely absent from the frame rather than merely sparse in it. Both are stated in the proposal rather than left for a reviewer.
    label_limits · calculé le 2026-07-13T11:54:49Z
    12

    La variance des agents

    Haiku, le modèle sur lequel le constat 13 budgète tout le tri, aboutit près du taux de base de Sonnet (1,27 % contre 1,06 %) et s'accorde avec lui sur 98,1 % de la base, mais leurs ENSEMBLES dans le champ ne se recoupent qu'à 16 % sans pondération et à 10 % avec la pondération du plan de sondage (Sonnet-GPT : 54 %/37 %) ; des 58 positifs de Sonnet, Haiku n'en confirme que 12. L'ACCORD SUR LE TAUX N'EST PAS L'ACCORD SUR L'ENSEMBLE. Pire : des agents du MÊME modèle sur la MÊME consigne divergent au-delà du hasard dans les DEUX volets après conditionnement sur la strate (CMH p = 0,0056 et 0,015), avec des écarts bruts de 3,1 et 5,2 fois contre 2,2 fois entre modèles, et l'ordre des agents S'INVERSE d'un volet à l'autre. Le bruit à l'intérieur d'un même modèle est au moins de la taille de la différence entre modèles, et le tri à 40 agents du pilote manque lui-même de puissance pour exclure la même instabilité (5 des 37 lots n'ont trouvé aucune métarecherche ; p = 0,113).

    Énoncé original (findings.json) : Haiku, the model finding 13 budgets the entire screen on, lands near Sonnet's base rate (1.27% vs 1.06%) and agrees with it on 98.1% of the frame, but their in-scope SETS overlap 16% unweighted and 10% design-weighted (Sonnet-GPT: 54%/37%); of Sonnet's 58 positives Haiku agrees on 12. RATE AGREEMENT IS NOT SET AGREEMENT. Worse: agents of the SAME model on the SAME prompt disagree beyond chance in BOTH arms after conditioning on stratum (CMH p = 0.0056 and 0.015), with raw spreads 3.1x and 5.2x against 2.2x between models, and the agents' ordering FLIPS between arms. The noise inside one model is at least the size of the difference between models, and the pilot's own 40-agent screen is too underpowered to rule the same instability out (5 of 37 chunks found zero metaresearch; p = 0.113).

    Champ source:n_works
    1 290
    Champ source:base_rate_sonnet_pct
    1,06
    Champ source:base_rate_gpt_pct
    2,37
    Champ source:base_rate_haiku_pct
    1,27
    Champ source:agreement_haiku_sonnet_pct
    98,1
    Champ source:jaccard_sonnet_gpt_pct
    54
    Champ source:jaccard_sonnet_haiku_pct
    16
    Champ source:jaccard_gpt_haiku_pct
    12
    Champ source:wjaccard_sonnet_gpt_pct
    37
    Champ source:wjaccard_sonnet_haiku_pct
    10
    Champ source:wjaccard_gpt_haiku_pct
    6
    Champ source:sonnet_positives
    58
    Champ source:haiku_agrees_on
    12
    Champ source:agent_rates_raw_pct
    agent-1:
    1,6
    agent-2:
    1,25
    agent-3:
    3,85
    Champ source:stratum_mix_differs_by_agent_p
    0,0000000000000000134
    Champ source:arm1_cmh_p
    0,00564
    Champ source:arm1_permutation_p
    0,0275
    Champ source:arm1_raw_spread_x
    3,1
    Champ source:arm2_cmh_p
    0,0152
    Champ source:arm2_permutation_p
    0,0075
    Champ source:arm2_raw_spread_x
    5,2
    Champ source:agent_order_replicates
    non
    Champ source:spread_weighted_x_leverage_sensitive
    13,2
    Champ source:between_model_spread_x
    2,2
    Champ source:within_at_least_matches_between
    oui
    Champ source:pilot_chunks
    37
    Champ source:pilot_agents_finding_zero
    5
    Champ source:pilot_between_agent_p
    0,113
    Champ source:pilot_underpowered_not_homogeneous
    oui
    Champ source:caveat
    The first draft's between-agent test was CONFOUNDED: the stratum mix differs by agent (p = 1.3e-17), and the draft asserted a verification that did not exist (DEVIATIONS.md D13). The tests above condition on stratum, and the heterogeneity survives in both arms. The 13.2x design-weighted spread the draft led with rests on five high-weight events and is demoted to a recorded, leverage-sensitive descriptive. THE INFERENCE IS SCOPED: there are three agents per arm, assigned consecutive chunk blocks without randomization or a run-time manifest, so these p-values license 'these runs are not exchangeable', not a population claim about agents in general; that is exactly enough to break a budget that assumed exchangeability, and the full study assigns agents randomized, manifest-logged, fixed-size chunks with a duplicate-agent reliability arm. The pilot's own agents give p = 0.113 on ~2 expected events per chunk: underpowered, so the pilot is uninformative on agent homogeneity, not exonerated. Haiku was tested on the six-field payload (the pilot's own deviation, D1) and on a guided eight-field arm, so the defensible conclusion is 'not shown to be an interchangeable measurer, and unstable in the arms tested', not 'cannot screen'. An eight-field neutral-prompt arm is not used at all: one agent claimed six label files it never wrote (D11), so no payload effect is reported from any arm. The guided arm is used ONLY for the between-agent contrast, which its shared prompt leaves internally valid.
    agent_variance · calculé le 2026-07-15T22:52:18Z
    13

    L'écart des résumés est structurel

    Le plus grand biais mesuré du tri est l'écart des résumés : 23,3 % de la base (1 003 117 travaux) n'a AUCUN RÉSUMÉ, et le constat 11 a montré que le tri y repère MOITIÉ moins de métarecherche. La cascade PubMed, Europe PMC puis Crossref récupère 37,8 % d'un échantillon de 500 travaux, ramenant l'exposition au tri sur seul titre à environ 14,5 % de la base. Mais J'AVAIS BÂTI LA CASCADE AUTOUR DE CROSSREF comme voie de secours indépendante des disciplines, et il a récupéré 2 résumés contre 180 pour PubMed : les éditeurs ne les déposent pas, donc CETTE VOIE DE SECOURS N'EXISTE PAS (D15). L'écart n'est donc pas une défaillance de métadonnées qu'un meilleur index corrigerait ; il est STRUCTUREL. La récupération atteint 91,2 % pour les articles de synthèse contre 6,2 % pour les chapitres de livre, 38,8 % pour l'anglais contre 15,4 % pour le français. Le raccourci tentant, « ne trier que les travaux qui ont un résumé », est donc une SÉLECTION SUR UNE COVARIABLE QUI PRÉDIT LE RÉSULTAT : il supprimerait 61,6 % des chapitres de livre contre 22,1 % des articles, ET les travaux qu'il supprime sont exactement ceux qu'aucune cascade ne peut récupérer. Défendable seulement comme exclusion DÉCLARÉE au coût mesuré, et l'audit y conserve un plancher d'échantillonnage.

    Énoncé original (findings.json) : The screen's largest measured bias is the abstract gap: 23.3% of the frame (1,003,117 works) has NO ABSTRACT, and finding 11 showed the screen finds HALF as much metaresearch there. Cascading PubMed, Europe PMC and Crossref recovers 37.8% of a 500-work sample, cutting title-only exposure to ~14.5% of the frame. But I BUILT THE CASCADE AROUND CROSSREF as the discipline-agnostic rescue, and it recovered 2 abstracts against PubMed's 180: publishers do not deposit them, so THAT RESCUE DOES NOT EXIST (D15). The gap is therefore not a metadata failure a better index fixes; it is STRUCTURAL. Recovery is 91.2% for reviews against 6.2% for book chapters, 38.8% English against 15.4% French. So the tempting shortcut, 'just screen the works that have abstracts', is a SELECTION ON A COVARIATE THAT PREDICTS THE OUTCOME which would delete 61.6% of book chapters against 22.1% of articles, AND the works it deletes are exactly the works no cascade can rescue. Defensible only as a DECLARED exclusion with a measured cost, and the audit keeps a sampling floor in it.

    Champ source:frame_works
    4 299 418
    Champ source:frame_works_no_abstract
    1 003 117
    Champ source:pct_frame_no_abstract
    23,3
    Champ source:pct_dropped_by_type
    book-chapter:
    61,6
    letter:
    52,1
    editorial:
    43,3
    review:
    29,5
    article:
    22,1
    book:
    21,6
    other:
    21,1
    report:
    18,6
    preprint:
    14
    dataset:
    8,2
    dissertation:
    4,3
    Champ source:pct_dropped_by_language
    fr:
    21,6
    en:
    23,7
    Champ source:abstracts_only_is_a_selection_on_the_outcome
    oui
    Champ source:sampled
    500
    Champ source:sources
    PubMed (Entrez) -> Europe PMC (REST) -> Crossref (REST)
    Champ source:recovered
    189
    Champ source:pct_gap_recovered
    37,8
    Champ source:recovered_by_source
    pubmed:
    180
    europepmc:
    7
    crossref:
    2
    Champ source:hypothesis_crossref_would_be_load_bearing
    non
    Champ source:crossref_recovered
    2
    Champ source:europepmc_recovered
    7
    Champ source:pubmed_recovered
    180
    Champ source:no_discipline_agnostic_rescue_exists
    oui
    Champ source:the_gap_is_structural_not_a_metadata_failure
    oui
    Champ source:recovery_pct_by_type
    review:
    91,2
    article:
    40,5
    preprint:
    38,5
    book-chapter:
    6,2
    letter:
    0
    Champ source:recovery_pct_english
    38,8
    Champ source:recovery_pct_french
    15,4
    Champ source:residual_pct_frame_title_only
    14,5
    Champ source:caveat
    Run on a 500-work hash-ordered sample of the no-abstract stratum, not the frame: the cascade is rate-limited and a million lookups is days. The recovery rate is an estimate with sampling error, and it is an estimate of a CEILING (an abstract that EXISTS is recoverable; it does not follow the screen then classifies the work correctly). Only works with a DOI can be looked up, so the DOI-less part of the stratum is untouched and its size bounds what any cascade can do. PubMed and Europe PMC are biomedical; Crossref is not, and it is in the chain for exactly that reason: a cascade of biomedical indexes would close the gap unevenly and make the residual bias MORE discipline-shaped while appearing to improve coverage. That reasoning was right and the remedy is not available: Crossref recovered 2 of 189 and Europe PMC 7, because publishers largely do not deposit abstracts to Crossref, so no discipline-agnostic rescue exists (D15). Restricting screening to abstract-bearing works remains a DECLARED EXCLUSION with a measured cost, not a scoping convenience, and the audit keeps a sampling floor in the excluded stratum so that cost stays estimable.
    abstract_cascade · calculé le 2026-07-13T11:55:17Z
    14

    Un booléen sur un espace à quatre états

    Jointe à la base canadienne par DOI, Retraction Watch consigne 143 travaux qu'OpenAlex ne signale PAS comme rétractés, dont 49 rétractations pures et simples. Mais le sous-compte est le moindre des problèmes. 52 de ces travaux portent une EXPRESSION DE PRÉOCCUPATION, et OpenAlex N'A AUCUN CHAMP POUR CELA : is_retracted est un booléen sur un espace d'états à au moins quatre valeurs (rétractation, expression de préoccupation, correction, rétablissement) ; il peut en exprimer une et rapporte silencieusement les autres comme FALSE, ce qui se lit comme « rien à signaler ». Un booléen ne peut pas non plus porter le POURQUOI. C'est la maladie du constat 1 dans un second schéma : la base de données canonique ne peut pas exprimer la distinction sur laquelle le domaine repose.

    Énoncé original (findings.json) : Joined to the Canadian frame by DOI, Retraction Watch records 143 works that OpenAlex does NOT flag as retracted, including 49 outright retractions. But the undercount is the smaller problem. 52 of these carry an EXPRESSION OF CONCERN, and OpenAlex HAS NO FIELD FOR ONE: `is_retracted` is a boolean over a state space with at least four values (retraction, expression of concern, correction, reinstatement), so it can express one and silently reports the rest as FALSE, which reads as 'fine'. Nor can a boolean carry WHY. This is finding 1's disease in a second schema: the canonical database cannot express the distinction the field turns on.

    Champ source:source
    Retraction Watch (Crossref-licensed), joined by bare lowercased DOI
    Champ source:frame_works_with_doi
    3 690 953
    Champ source:openalex_is_retracted_flags
    1 584
    Champ source:matched_in_retraction_watch
    1 052
    Champ source:openalex_misses
    143
    Champ source:outright_retractions_missed
    49
    Champ source:expressions_of_concern
    52
    Champ source:missed_by_nature
    Expression of concern:
    52
    Retraction:
    49
    Correction:
    32
    Reinstatement:
    10
    Champ source:top_reasons
    Investigation by Journal/Publisher:
    335
    Unreliable Results and/or Conclusions:
    259
    Concerns/Issues about Data:
    233
    Investigation by Third Party:
    169
    Concerns/Issues about Referencing/Attributions:
    151
    Concerns/Issues about Results and/or Conclusions:
    122
    Concerns/Issues about Peer Review:
    109
    Investigation by Company/Institution:
    101
    Champ source:openalex_has_eoc_field
    non
    Champ source:is_a_frame_route
    non
    Champ source:caveat
    This is an ATTRIBUTE, not a frame route. A retracted cardiology paper is retracted cardiology, not metaresearch, and admitting works on the strength of a retraction would let an interesting signal masquerade as the estimand. The DOI join can only see works that HAVE a DOI, and OpenAlex flags some works Retraction Watch does not match, which may be DOI drift rather than disagreement; the asymmetry reported here is one-directional on purpose (what RW adds), because that is the direction the join can support.
    retraction_record · calculé le 2026-07-13T11:55:00Z
    15

    Le rappel de la voie thématique

    Évaluée selon la grille, la voie thématique repère 12 % de la métarecherche canadienne (IC à 95 % de 5,6 à 21,6 %) avec une précision de 60 % : elle manque 66 travaux sur 75. Elle échoue parce qu'OpenAlex classe un travail selon ce dont il traite, et la métarecherche sur la cardiologie se lit comme de la cardiologie : le domaine est invisible au repérage thématique précisément parce qu'il porte sur d'autres domaines.

    Énoncé original (findings.json) : Scored against the rubric, the topic route retrieves 12% of Canadian metaresearch (95% CI 5.6-21.6%) at 60% precision: it misses 66 of 75. It fails because OpenAlex files a work by what it is about, and metaresearch about cardiology reads as cardiology: the field is invisible to topic retrieval precisely because it is about other fields.

    Champ source:n_screened
    5 737
    Champ source:n_metaresearch
    75
    Champ source:n_retrieved_by_route
    15
    Champ source:true_positives
    9
    Champ source:false_positives
    6
    Champ source:false_negatives
    66
    Champ source:recall_pct
    12
    Champ source:recall_ci_lo_pct
    5,6
    Champ source:recall_ci_hi_pct
    21,6
    Champ source:precision_pct
    60
    Champ source:precision_ci_lo_pct
    32,3
    Champ source:precision_ci_hi_pct
    83,7
    Champ source:missed_works_by_field
    Social Sciences:
    17
    Medicine:
    16
    Health Professions:
    7
    Business, Management and Accounting:
    4
    Economics, Econometrics and Finance:
    4
    Computer Science:
    3
    Champ source:scored_by
    primary_topic.id (the key R/frame.R defines the route with), not display name
    Champ source:route_size_implied_by_sample
    11 241
    Champ source:route_size_from_api
    14 873
    Champ source:reconciliation_gap_x
    1,32
    Champ source:supersedes
    the earlier 32.4% figure (14,873/45,850), which divided a retrieved set by a true field size
    Champ source:caveat
    Small n on the positive class ( 75 metaresearch works, of which 15 were on the route), so the intervals are wide. Separately, the sample-implied route size does not reconcile with the API's 14,873 and I cannot say why at n = 15 on-route works; the likeliest cause is that one updated_date partition is not a uniform draw (finding 11). Recall is unaffected: it is a within-sample ratio, not an extrapolation.
    topic_route_recall · calculé le 2026-07-13T11:54:47Z
    16

    Le rappel de la voie du financement

    L'ESTIMANDE PRINCIPALE s'appuie sur CA-FUND pour récupérer les travaux dont l'affiliation manque (constat 2 : 64 % n'ont aucune chaîne d'affiliation), et rien ne l'avait mise à l'épreuve. Confrontée à la base de données des IRSC eux-mêmes, soit 44 190 projets financés, l'épreuve A RÉFUTÉ L'HYPOTHÈSE QUE J'AVAIS ÉCRITE AVANT DE L'EXÉCUTER : OpenAlex étiquette 178 133 travaux de la base avec les IRSC, soit 4,03 par subvention, un taux plausible sans sous-étiquetage (DEVIATIONS.md D14). Ce que les données soutiennent, elles, n'exige aucune hypothèse de ma part : 71,2 % DE LA BASE NE PORTE AUCUNE MÉTADONNÉE DE FINANCEMENT, ce qui est le plafond absolu de CA-FUND, et 65,9 % des travaux à affiliation canadienne n'en portent pas non plus. Les deux clauses de l'estimande reposent sur des métadonnées le plus souvent absentes : c'est pourquoi la base est l'union de quatre voies, et pourquoi l'audit doit échantillonner les travaux qu'aucune voie n'a atteints.

    Énoncé original (findings.json) : The PRIMARY ESTIMAND leans on CA-FUND to rescue works whose affiliation is missing (finding 2: 64% have no affiliation string), and nothing had tested it. Tested against CIHR's own database of 44,190 funded projects, the result REFUTED THE HYPOTHESIS I WROTE BEFORE RUNNING IT: OpenAlex tags 178,133 frame works with CIHR, or 4.03 per grant, a plausible rate showing no under-tagging (DEVIATIONS.md D14). What the data DOES support needs no hypothesis of mine: 71.2% OF THE FRAME CARRIES NO FUNDER METADATA AT ALL, which is CA-FUND's hard ceiling, and 65.9% of Canadian-AFFILIATED works carry none either. Both clauses of the estimand rest on metadata that is mostly absent, which is why the frame is a union of four routes and why the audit must sample the works no route reached.

    Champ source:external_criterion
    CIHR's own project database (44,190 projects), which owes nothing to OpenAlex
    Champ source:cihr_projects
    44 190
    Champ source:cihr_distinct_pis
    27 536
    Champ source:cihr_funder_id
    https://openalex.org/F4320334506
    Champ source:frame_works
    4 299 418
    Champ source:frame_works_tagged_cihr
    178 133
    Champ source:tagged_publications_per_funded_project
    4,03
    Champ source:papers_per_grant_is_a_plausible_rate_not_a_defect
    oui
    Champ source:expectation_i_wrote_before_running_and_that_was_false
    that OpenAlex under-tags CIHR so badly it falls below one paper per grant. It is 4.03 per grant, a plausible rate. See DEVIATIONS.md D14.
    Champ source:works_ca_fund_rescues_alone
    166 743
    Champ source:frame_works_with_any_funder
    1 239 950
    Champ source:pct_frame_with_any_funder
    28,8
    Champ source:pct_frame_with_no_funder
    71,2
    Champ source:ca_aff_works
    2 734 192
    Champ source:ca_aff_works_with_no_funder
    1 802 605
    Champ source:pct_ca_aff_with_no_funder
    65,9
    Champ source:top_recorded_funders
    Natural Sciences and Engineering Research Council of Canada:
    294 401
    Canadian Institutes of Health Research:
    178 133
    National Institutes of Health:
    81 262
    National Natural Science Foundation of China:
    67 358
    National Science Foundation:
    62 497
    Canada Research Chairs:
    35 113
    European Commission:
    34 039
    Social Sciences and Humanities Research Council of Canada:
    33 473
    Champ source:is_an_aggregate_not_a_record_level_join
    oui
    Champ source:caveat
    CIHR's CSV carries no DOIs and no publication links, so this is an AGGREGATE reconciliation, not a record-level known-item join, and NO RECALL POINT ESTIMATE is claimed. It establishes a CEILING on CA-FUND (the route cannot see a funder OpenAlex never recorded), which is a bound, not a measurement. The papers-per-grant ratio is reported because I ran it, and it REFUTES the hypothesis I wrote before running it: at 4.03 per grant it is a plausible publication rate and shows no CIHR under-tagging at all. It is also the wrong instrument, for the same reason the retracted 32.4% coverage figure was (D3): its numerator and denominator are not linked record to record, so the quotient has no estimand behind it. Record-level linkage is finding 21.
    funder_route_recall · calculé le 2026-07-13T11:55:02Z
    17

    Le lien canadien

    L'affiliation trouve 14 873 travaux ; 3 964 autres portent sur le Canada sans affiliation canadienne. Le CRSNG a 9,2 fois plus de travaux liés que le CRSH : les règles fondées sur le financement sous-comptent donc les sciences sociales.

    Énoncé original (findings.json) : Affiliation finds 14,873 works; a further 3,964 are about Canada with no Canadian affiliation. NSERC has 9.2x SSHRC's linked works, so funder-based rules under-count the social sciences.

    Champ source:by_affiliation
    14 873
    Champ source:by_funder
    1 331
    Champ source:about_canada
    5 704
    Champ source:about_canada_no_affiliation
    3 964
    Champ source:funder_works
    Canadian Institutes of Health Research:
    194 681
    Natural Sciences and Engineering Research Council of Canada:
    433 090
    Social Sciences and Humanities Research Council of Canada:
    46 913
    Canada Foundation for Innovation:
    16 753
    Champ source:sshrc_works
    46 913
    Champ source:nserc_works
    433 090
    Champ source:nserc_to_sshrc_ratio
    9,2
    Champ source:affiliation_noise
    University of London:
    494
    Impact:
    475
    canadian_linkage · calculé le 2026-07-13T11:54:44Z
    18

    Le lien aux essais cliniques

    Le registre est le seul ÉTALON DE RÉFÉRENCE de ce projet qui ne soit pas fait d'étiquettes de machine : ClinicalTrials.gov sait qu'un essai canadien a eu lieu indépendamment de toute chaîne de traitement, il ne peut donc pas se tromper en faveur de la chaîne. Des 304 publications que les PROMOTEURS EUX-MÊMES ont déclarées comme résultats d'essais achevés menés au Canada, la base en détient 160 : un rappel naïf de 52,6 %, tombé si près du 44,5 % PAR LEQUEL CETTE PROPOSITION S'OUVRE qu'il se lisait comme une réplication. C'EST UN ARTEFACT, et seule l'exécution de la désambiguïsation l'a détecté : 137 des 145 « manqués » n'ont AUCUN AUTEUR CANADIEN (essais internationaux multicentriques avec un SITE canadien), et une base de la RECHERCHE canadienne a RAISON de les exclure. Un essai avec un site canadien n'est pas une publication avec un auteur canadien. Contre la population que la base revendique réellement, le rappel est de 95,2 % (IC à 95 % de 90,8 à 97,9), et le vrai défaut tient aux 8 travaux qu'OpenAlex détient AVEC un auteur canadien et que les voies de la base ont tout de même manqués. La base est BONNE à cet exercice, le parallèle spectaculaire était une coïncidence entre deux populations différentes, et j'avais toutes les raisons de ne pas vérifier. DEVIATIONS.md D16.

    Énoncé original (findings.json) : The registry is the only REFERENCE STANDARD in this project not made of machine labels: ClinicalTrials.gov knows a Canadian trial happened independently of any pipeline, so it cannot be wrong in the pipeline's favour. Of 304 publications SPONSORS THEMSELVES reported as results of completed Canadian-located trials, the frame holds 160: a naive recall of 52.6%, which fell so close to the 44.5% THIS PROPOSAL OPENS WITH that it read as a replication. IT IS AN ARTIFACT, and running the disambiguation is the only thing that caught it: 137 of the 145 'misses' have NO CANADIAN AUTHOR (multi-site international trials with a Canadian SITE), and a frame of Canadian RESEARCH is CORRECT to exclude them. A trial with a Canadian site is not a publication with a Canadian author. Against the population the frame actually claims, recall is 95.2% (95% CI 90.8-97.9), and the real defect is 8 works OpenAlex holds WITH a Canadian author that the frame's own routes still missed. The frame is GOOD at this, the dramatic parallel was a coincidence between two different populations, and I had every incentive not to check. DEVIATIONS.md D16.

    Champ source:reference_standard
    ClinicalTrials.gov: completed trials with a Canadian location, and the RESULT publications the sponsors themselves reported
    Champ source:why_not_a_frame_route
    A trial registration is not a publication and not metaresearch. Registrations contribute NO records to the frame; the registry is a reference standard, not a source.
    Champ source:why_it_matters
    Every other recall number in this project is scored against MACHINE labels (finding 15). A registry knows a trial happened independently of any pipeline, so it cannot be wrong in the pipeline's favour. This is the only instrument here that measures the frame against a world that exists without it.
    Champ source:trials_retrieved
    1 000
    Champ source:trials_with_result_publication
    123
    Champ source:pct_trials_with_result_publication
    12,3
    Champ source:known_result_pmids
    596
    Champ source:resolvable_to_doi
    304
    Champ source:present_in_frame
    160
    Champ source:naive_frame_recall_pct
    52,6
    Champ source:naive_recall_is_an_artifact
    oui
    Champ source:naive_recall_ci
    • 46,9
    • 58,4
    Champ source:missing_total
    145
    Champ source:missing_no_canadian_author
    137
    Champ source:missing_route_gap_canadian_author
    8
    Champ source:missing_not_in_openalex
    0
    Champ source:claimable_population
    168
    Champ source:adjusted_recall_pct
    95,2
    Champ source:adjusted_recall_ci
    • 90,8
    • 97,9
    Champ source:is_metaresearch_recall
    non
    Champ source:caveat
    This measures FRAME recall (does the Canadian frame hold the publication at all?), NOT metaresearch recall: trial reports are primary research and the rubric screens them OUT. It is the precondition for screening, not the screen. The reference standard is the sponsor's OWN reported result publications, so it is incomplete in a known direction: sponsors under-report, which means the true set of trial publications is LARGER than the standard and this recall figure is measured only on the ones we can see. Trials are matched by Canadian LOCATION, which is not the same as Canadian authorship, so some result publications may have no Canadian author and legitimately fall outside the frame; that direction is not controlled here and it bounds the interpretation. Only PMIDs resolvable to a DOI can be looked up.
    trial_linkage · calculé le 2026-07-13T11:55:20Z
    19

    La couverture des prépublications

    La base porte 156 086 prépublications, et la tentation était d'affirmer qu'OpenAlex couvre les serveurs de prépublications et de sauter l'ingestion. C'est une AFFIRMATION DE COUVERTURE, et ce projet n'a pas le droit d'en faire une depuis l'intérieur de la chaîne de traitement même qu'elle concerne : ce serait la voie thématique certifiant son propre rappel (constat 12). Mesurée plutôt contre la PROPRE API des serveurs, OpenAlex indexe 99,6 % des 705 prépublications bioRxiv et medRxiv énumérées (3 manquantes). L'affirmation survit à la mesure ; aucune ingestion séparée de prépublications n'est donc construite.

    Énoncé original (findings.json) : The frame carries 156,086 preprints, and the tempting move was to assert that OpenAlex covers the preprint servers and skip the ingest. That is a COVERAGE CLAIM, and this project does not get to make one from inside the pipeline being claimed for: it is the topic route certifying its own recall (finding 12). Measured instead against the servers' OWN API, OpenAlex indexes 99.6% of the 705 bioRxiv and medRxiv preprints enumerated (3 missing). The claim survives measurement, so no separate preprint ingest is built.

    Champ source:external_criterion
    the bioRxiv/medRxiv details API, which enumerates the servers' own corpus and owes nothing to OpenAlex
    Champ source:window
    2023-01-01 to 2025-12-31
    Champ source:preprints_enumerated
    705
    Champ source:indexed_by_openalex
    702
    Champ source:missing_from_openalex
    3
    Champ source:pct_indexed
    99,6
    Champ source:sampled_with_canadian_author
    45
    Champ source:of_those_present_in_frame
    45
    Champ source:frame_preprints_total
    156 086
    Champ source:separate_ingest_needed
    non
    Champ source:caveat
    The bioRxiv API exposes no author country, so the sample is mostly non-Canadian and the clean quantity is INDEX coverage (does OpenAlex hold the preprint at all?), not Canadian recall: a preprint the index lacks cannot enter any frame by any route, so index coverage is the binding upper bound. The Canadian sub-count is small and is reported as a check, not as an estimate. arXiv is NOT tested here (its OAI endpoint pages differently); the frame carries 11,647 arXiv works and that claim remains untested against arXiv itself.
    preprint_coverage · calculé le 2026-07-13T11:55:18Z
    20

    La puissance d'un audit humain

    L'audit tel que d'abord spécifié ne pouvait pas mesurer ce qu'il existe pour mesurer. Un échantillon aléatoire simple de la masse rejetée de 4 243 096 notices donne 0,4 occurrence attendue sur les 600 notices budgétées ; en voir 20 exigerait 2 009 heures de codage contre les 65 prévues. Le suréchantillonnage stratifié par score ne le sauve qu'en partie, car le pilote montre que 30 des 37 travaux disputés se trouvent à la frontière contestée, et il reste AVEUGLE aux rejets assurés. Savoir quels travaux ce sont relève d'un MÉCANISME, non d'un écart mesuré : la grille dit de juger sur le seul titre en l'absence de résumé, et dit qu'un travail T2 peut n'employer aucun mot du vocabulaire du domaine ; un travail privé des deux est donc rejeté avec assurance et jamais échantillonné. Une version antérieure citait ici un chiffre de 3,6 fois comme s'il était mesuré ; il reposait sur quatre travaux et il est retiré. Le rappel se mesure donc de deux autres façons : sur les 5 737 travaux qui portent déjà des étiquettes de la grille, et par un rappel sur cibles connues à partir d'un ensemble de revues de référence, parce que la revue est un critère externe insensible à cette « aboutness » qui met tout le reste en échec.

    Énoncé original (findings.json) : The audit as first specified could not measure what it exists to measure. A simple random sample of the 4,243,096-record screened-out mass yields an expected 0.4 hits from the 600 records budgeted; seeing 20 would take 2,009 coder-hours against the 65 planned. Score-stratified oversampling rescues it only partly (the pilot shows 30 of 37 disputed works sit at the contested boundary), and it stays BLIND to the confident rejects. Which works those are is a MECHANISM, not a measured differential: the rubric says judge on the title alone with no abstract, and says T2 work may use none of the field's vocabulary, so a work with neither is rejected confidently and never sampled. (An earlier version cited a 3.6x figure here as if it were measured; it rested on four works and is withdrawn.) So recall is measured two other ways: against the 5,737 works that already carry rubric labels, and by KNOWN-ITEM recall on a venue reference set, because venue is an external criterion immune to the aboutness that defeats everything else.

    Champ source:problem
    A simple random sample of the screened-out stratum cannot measure screening sensitivity: the works the screen wrongly rejected are a vanishing fraction of a 3.2M-record rejected mass.
    Champ source:screened_out_pool
    4 243 096
    Champ source:screened_out_is_the_screens_rejects
    oui
    Champ source:audit_screened_out_budgeted
    600
    Champ source:expected_hits_at_95_recall
    0,4
    Champ source:codings_needed_for_20_hits_at_95_recall
    30 134
    Champ source:coder_hours_needed
    2 009
    Champ source:coder_hours_budgeted
    65
    Champ source:naive_audit_is_powered
    non
    Champ source:disputed_in_boundary
    30
    Champ source:disputed_in_settled_rejects
    6
    Champ source:misses_concentrate_near_threshold
    oui
    Champ source:blind_spot
    Score-stratified oversampling finds the works the screener ALMOST caught; it is blind to the ones it rejected CONFIDENTLY. WHICH works those are is a MECHANISM, not a measurement: the rubric says judge on the title alone when the abstract is missing, and the rubric also says T2 work may use none of the field's vocabulary, so a work with neither is rejected confidently and sits deep in the settled rejects. An earlier version cited a 3.6x differential from finding 11 as if this were measured. It is not, and that figure is withdrawn (DEVIATIONS.md D6). The venue instrument is how the prediction gets tested rather than asserted.
    Champ source:blind_spot_is_a_mechanism_not_a_measurement
    oui
    Champ source:instrument_1
    Measure any filter's recall against the 5737 works that already carry full-rubric labels, exactly as finding 12 scored the topic route. Free, and it needs no needle-hunting in the discarded mass.
    Champ source:instrument_2
    Known-item recall on an external criterion: VENUE. A Canadian-authored paper in Social Studies of Science is T2 by where it was published, whatever its abstract is about. Venue is immune to aboutness, which is what defeats topic retrieval (finding 12) and title-only screening (finding 11), so it is the only instrument that can see into the blind spot.
    Champ source:reference_venues
    • Social Studies of Science
    • Scientometrics
    • Quantitative Science Studies
    • Research Integrity and Peer Review
    • Journal of the Association for Information Science and Technology
    • Research Evaluation
    • Accountability in Research
    • PLOS ONE (metaresearch collection)
    • Recherches qualitatives
    • Documentation et bibliotheques
    Champ source:n_labelled_works
    5 737
    Champ source:n_in_scope
    75
    Champ source:caveat
    The recall grid assumes the screen's misses are uniform in the rejected mass, which the pilot shows they are not (they concentrate at the boundary). That makes the naive design LESS hopeless than the grid implies but does not save it, and it does nothing at all about the confident-reject blind spot. Known-item recall on a venue reference set is a non-probability estimate: it bounds and diagnoses recall on the hard cases, it does not replace the design-weighted population estimate.
    audit_power · calculé le 2026-07-13T11:54:49Z
    21

    Ce que coûte le tri

    La grille complète appliquée par deux trieurs à la base de 4 299 418 travaux coûte 6 567 $, et non les 30 633 $ qu'une version antérieure du script annonçait : la grille est une consigne système envoyée une fois par APPEL, et les lots du pilote regroupent eux-mêmes 155 travaux par appel ; la facture était donc gonflée de 4,7 fois. Cette erreur n'était pas cosmétique. Elle m'a fait proposer un TRIAGE bon marché devant le tri, soit une ÉTAPE DE REPÉRAGE dans un projet dont le constat central est que le repérage détruit ces cartes. L'arithmétique corrigée, le triage devient inutile : la grille complète sur CHAQUE travail de la base, plus un second trieur sur un échantillon de 20 000 notices, coûte 1 110 $ avec l'instrument v1 et 1 279 $ avec l'instrument v3.1 verrouillé qui sera réellement exécuté. LE PRÉFILTRE EST SUPPRIMÉ. Le calcul est AUTOFINANCÉ : les 4 000 $ CA de l'appel couvrent, selon son propre texte, les déplacements, l'hébergement et les frais de participation connexes, non l'inférence ni la rémunération du codage (D30).

    Énoncé original (findings.json) : The full rubric over the 4,299,418-work frame with two screeners costs $6,567, not the $30,633 an earlier version of this script reported: the rubric is a system prompt sent once per CALL, and the pilot's own chunks batch 155 works per call, so it was overcharged 4.7x. That error was not cosmetic. It made me propose a cheap TRIAGE in front of the screen, which is a RETRIEVAL STEP in a project whose central finding is that retrieval destroys these maps. With the arithmetic right the triage is unnecessary: the full rubric over EVERY work in the frame, plus a second screener on a 20,000-record sample, costs $1,110 at the v1 instrument and $1,279 at the locked v3.1 instrument that will actually run. THE PREFILTER IS DELETED. And the compute is SELF-FUNDED: the call's CAD $4,000 covers, in its own words, 'travel, accommodation, and related participation costs', not inference and not coder wages (D30).

    Champ source:measured_from
    pilot/screening/chunks/*.json (6-field), pilot/screening/haiku/p8_guided/chunk_*.json (8-field), docs/protocol/rubric.md; not asserted
    Champ source:chars_per_token_assumed
    4
    Champ source:tokens_per_work
    312
    Champ source:tokens_per_work_six_field_deviation
    256
    Champ source:payload_inflation_8_over_6
    1,22
    Champ source:costed_at_the_rubric_payload_not_the_deviation
    oui
    Champ source:tokens_rubric
    1 878
    Champ source:tokens_per_label
    37
    Champ source:batch_works_per_call
    155
    Champ source:frame_size
    4 299 418
    Champ source:grant_usd_approx
    2 900
    Champ source:naive_cost_charging_rubric_per_work_usd
    30 633
    Champ source:corrected_cost_two_screeners_usd
    6 567
    Champ source:cost_overstatement_x
    4,7
    Champ source:cost_full_rubric_whole_frame_cheap_usd
    1 094
    Champ source:cost_full_rubric_whole_frame_sonnet_usd
    3 283
    Champ source:cost_second_screener_on_sample_usd
    15
    Champ source:second_screener_n
    20 000
    Champ source:total_no_prefilter_usd
    1 110
    Champ source:tokens_rubric_v31
    4 604
    Champ source:tokens_per_label_v31
    49
    Champ source:cost_full_rubric_whole_frame_v31_usd
    1 261
    Champ source:cost_second_screener_v31_usd
    18
    Champ source:total_no_prefilter_v31_usd
    1 279
    Champ source:award_pays_for_compute
    non
    Champ source:award_purpose_in_the_calls_words
    travel, accommodation, and related participation costs
    Champ source:prefilter_needed
    non
    Champ source:prefilter_design_total_usd
    1 004
    Champ source:prefilter_saving_usd
    -106
    Champ source:prefilter_deleted_because
    It is a RETRIEVAL STEP, in a project whose central finding is that retrieval destroys these maps (finding 12: the topic route finds 12% of the field). It existed only because the rubric was miscosted at once per WORK rather than once per CALL, overstating the alternative 5.1x. With the arithmetic right it saves almost nothing and costs the thesis. Deleted.
    Champ source:pilot_works_screened
    6 202
    Champ source:frame_to_pilot_ratio
    693
    Champ source:method_that_scales
    batch inference, not agent fan-out
    Champ source:supersedes
    an earlier version charged the rubric once per work, reported $24,379 for the full-frame two-screener design, called it 'eight times the grant', and used that to justify a cheap prefilter. The rubric is a system prompt sent once per CALL, and the pilot's own chunks batch 155 works per call. DEVIATIONS.md D7.
    Champ source:caveat
    Token counts use a 4-chars-per-token approximation and list prices as of 2026-07; both will move, and the conclusion is robust to +/-25% in either. Screening the frame with a cheaper model than the pilot used makes the CHOICE OF MODEL more consequential, not less: finding 10 shows two screeners already imply base rates a factor of two apart. That is precisely why the second screener and the screener-swap range are reported, and why the human audit is the study.
    screening_cost · calculé le 2026-07-13T12:23:27Z
    22

    OpenAlex est facturé à l'usage

    L'API OpenAlex est facturée à l'usage : 1 000 crédits par environ 11 heures, avec un palier gratuit de 0,10 $. Énumérer la base canadienne de 4 299 418 travaux exige 21 498 appels paginés par curseur, soit 10 jours par passage au palier gratuit. Une chaîne de traitement fondée sur l'API à cette échelle n'est ni gratuite ni reproductible ; l'instantané épinglé est les deux.

    Énoncé original (findings.json) : The OpenAlex API is metered (1,000 credits per ~11h; $0.10 free tier). Enumerating the 4,299,418-work Canadian frame needs 21,498 cursor-paged calls: 10 days per pass on the free tier. An API-based pipeline at this scale is neither free nor reproducible; the pinned snapshot is both.

    Champ source:observed_retry_after_s
    40 268
    Champ source:observed_ratelimit_limit
    1 000
    Champ source:observed_free_tier_usd
    0,1
    Champ source:observed_cost_per_call_usd
    0,0001
    Champ source:frame_size_works
    4 299 418
    Champ source:per_page_max
    200
    Champ source:calls_for_one_pass
    21 498
    Champ source:days_on_free_tier_per_pass
    10
    Champ source:prepaid_cost_per_pass_usd
    2,15
    openalex_is_metered · calculé le 2026-07-13T11:54:45Z
    23

    active_learning

    LA BOUCLE PAR LOTS DE 100 FONCTIONNE, MAIS SES 20 ANCRES ALÉATOIRES NE FONCTIONNENT PAS. À budget identique, l'acquisition active (50 cas au désaccord maximal entre enseignants, 30 à l'incertitude maximale et 20 aléatoires) fait passer l'AP hors échantillon de 0,0143 à 0,1391, alors que le témoin par lots aléatoires atteint 0,0502 : un rapport de 2,8, avec une victoire dans 18 des 20 rondes et une rotation de l'ensemble positif hors échantillon qui tombe à 0,1111. À un taux de base d'environ 1 %, un lot aléatoire de 100 contient un positif, et la courbe témoin montre ce que cela permet. MAIS les ancres, censées maintenir un flux d'évaluation non biaisé, ne contiennent que 3 positifs après 400 tirages : une prévalence de 0,0057 avec un intervalle à 95 % d'une largeur de 0,01594, PLUS LARGE QUE LA QUANTITÉ ELLE-MÊME. Elles sont rétrogradées au rôle de sentinelles de dérive. LA BOUCLE APPREND ; L'AUDIT HUMAIN PRÉENREGISTRÉ MESURE. L'AP mesure ici l'accord avec l'ENSEIGNANT majoritaire ; toute la courbe mesure donc l'imitation et non l'exactitude. La première version du script concluait que la boucle SE DÉGRADAIT, parce qu'elle mesurait l'AP sur le bassin décroissant que l'algorithme modifiait lui-même ; la rotation restait exactement à 1,000 pendant vingt rondes, soit une constante et non une mesure (D33).

    Énoncé original (findings.json) : THE BATCH-OF-100 LOOP WORKS, AND THE 20 RANDOM ANCHORS IN IT DO NOT. Against a random-batch control at identical budget, active acquisition (50 max-teacher-disagreement + 30 max-uncertainty + 20 random) takes held-out AP from 0.0143 to 0.1391 while the control reaches 0.0502: 2.8x, winning 18 of 20 rounds, with held-out positive-set churn falling to 0.1111. At a ~1% base rate a random batch of 100 holds one positive, and the control curve is what that buys. BUT the anchors, which were supposed to keep an unbiased evaluation stream alive, hold 3 positives after 400 draws: prevalence 0.0057 with a 95% interval 0.01594 wide, WIDER THAN THE QUANTITY ITSELF. They are demoted to drift sentinels. THE LOOP LEARNS; THE PREREGISTERED HUMAN AUDIT MEASURES. And AP here is agreement with the majority TEACHER, so the whole curve is imitation, not accuracy. The first version of this script said the loop DEGRADED, because it measured AP on the shrinking pool the algorithm was itself editing; the churn sat at exactly 1.000 for twenty rounds, which is a constant, not a measurement (D33).

    Champ source:batch
    batch_size:
    100
    max_disagreement:
    50
    max_uncertainty:
    30
    random_anchor:
    20
    Champ source:rounds
    20
    Champ source:ap_active_start
    0,0143
    Champ source:ap_active_end
    0,1391
    Champ source:ap_random_start
    0,0238
    Champ source:ap_random_end
    0,0502
    Champ source:active_over_random_x
    2,8
    Champ source:active_beats_random_in_rounds
    18
    Champ source:holdout_churn_final
    0,1111
    Champ source:pi_directive_vindicated
    oui
    Champ source:what_ap_measures_here
    Agreement with the MAJORITY TEACHER, not accuracy. A rising curve is better imitation of three LLMs that overlap each other at Jaccard ~0.5. No amount of imitation crosses that gap, and the audit is the only instrument that can.
    Champ source:anchors_n
    400
    Champ source:anchors_positives
    3
    Champ source:anchors_prevalence
    0,0057
    Champ source:anchors_ci_width
    0,01594
    Champ source:anchors_are_a_usable_evaluation_stream
    non
    Champ source:anchors_demoted_to
    drift sentinels, rescored under the frozen final model, never used to pick a batch
    Champ source:simulation_caveat
    Every query reveals a label that already exists: 16,800 labels over 5,600 works. This is in-sample recycling. It demonstrates the loop's MECHANICS and cannot show that the loop matures on 4.3M unlabeled works, nor validate the v3.1 instrument, under which no screening has run.
    Champ source:superseded_result
    The first version of this experiment measured AP and churn on the POOL of unrevealed works. Active learning REMOVES contested works from the pool by construction, so the pool gets easier every round and the evaluation target moves under the metric. It reported AP FALLING from 0.019 to 0.011 with churn pinned at exactly 1.000 for twenty straight rounds, and the conclusion was going to be 'the loop degrades'. The constant churn is what gave it away. A metric computed on a set the algorithm is actively editing is not a metric. See DEVIATIONS.md D33.
    active_learning · calculé le 2026-07-15T22:52:19Z
    24

    adjudication

    UN JUGE INDÉPENDANT ET AVEUGLE CONCLUT QUE LA GRILLE EST MUETTE POUR 89 % DES TRAVAUX QUI ONT DIVISÉ LES MODÈLES. Fable 5, qui ne fait partie d'aucun des trois volets de tri, a arbitré les 179 travaux contestés en voyant les avis des trois trieurs comme A, B et C dans un ordre aléatoire, sans aucun nom de modèle ; il devait donc déterminer quelle lecture de la GRILLE était correcte plutôt que quel modèle croire. Il ne montre aucune préférence pour un volet (écart de 1,16 fois, khi carré p = 0,493) : ce n'est pas un quatrième vote pour l'un des trieurs, et il les a TOUS LES TROIS désavoués pour 6 travaux. Son verdict sur l'instrument : la grille est MUETTE pour 159 des 179 travaux contestés, et seulement 17 divergences viennent d'un trieur qui applique mal une règle existante. LA FRONTIÈRE DU DOMAINE EST FIXÉE PAR LE TRIEUR, NON PAR L'INSTRUMENT. Les zones de friction qui coûtent le plus d'accord sont methods_dev_vs_study (39 travaux), lis_sts_asymmetry (26) et history_of_science (19). C'est ce qui transforme l'arbitrage en mesure plutôt qu'en départage : non pas 179 réponses, mais un recensement pondéré par fréquence des phrases absentes de la grille et du nombre exact de travaux que chacune coûte. Les NIVEAUX attribués par le juge ne servent à rien et ne produisent aucune estimation : un modèle ne peut pas certifier un modèle, et cet argument vaut aussi pour le juge.

    Énoncé original (findings.json) : AN INDEPENDENT BLINDED JUDGE SAYS THE RUBRIC IS SILENT ON 89% OF THE WORKS THE MODELS FOUGHT OVER. Fable 5, which is not one of the three screened arms, adjudicated all 179 contested works seeing three screener opinions as A/B/C in random order with no model names, so it was asked which reading of the RUBRIC is right rather than which model to trust. It shows no arm preference (spread 1.16x, chi-square p = 0.493), so it is not a fourth vote for one of the screeners, and it overruled ALL THREE on 6 works. Its verdict on the instrument: the rubric is SILENT on 159 of 179 contested works, and only 17 splits are a screener misapplying a rule that actually exists. THE FIELD'S BOUNDARY IS BEING SET BY THE SCREENER, NOT BY THE INSTRUMENT. The seams that cost the most agreement are methods_dev_vs_study (39 works), lis_sts_asymmetry (26), and history_of_science (19). This is what converts adjudication from a tiebreak into a measurement: not 179 answers, but a frequency-weighted census of which sentences the rubric is missing and exactly how many works each one costs. The judge's TIERS are used for nothing and produce no estimate: a model cannot certify a model, and that argument applies to the judge too.

    Champ source:judge
    Fable 5 (xhigh), which is NOT one of the three screened arms
    Champ source:blinded
    oui
    Champ source:blinding
    three screener opinions presented as A/B/C, order randomized per work, no model names in the prompt
    Champ source:contested_works
    179
    Champ source:judge_agrees_with
    • 94
    • 95
    • 109
    Champ source:judge_agrees_with_none
    6
    Champ source:arm_preference_spread
    1,16
    Champ source:arm_preference_chisq_p
    0,493
    Champ source:judge_is_captured
    non
    Champ source:rubric_is_silent
    159
    Champ source:pct_rubric_is_silent
    89
    Champ source:splits_from_screener_error
    17
    Champ source:seam_census
    • 39
    • 26
    • 19
    • 19
    • 17
    • 12
    • 10
    • 8
    • 8
    • 7
    • 6
    • 5
    • 3
    Champ source:adjudicated_tiers
    • 90
    • 23
    • 36
    • 30
    Champ source:adjudicated_in_scope
    59
    Champ source:caveat
    The judge is a MODEL, not a human, and this is NOT a reference standard. It is an independent, blinded, documented fourth reading whose value is the SEAM CENSUS, not the tiers. Its tiers are used for nothing downstream and produce no estimate. The arm-preference test shows the judge is not CAPTURED; it says nothing about whether it is CORRECT, and correctness is exactly what the human audit exists to establish. A model cannot certify a model, which is the whole argument of this project and it applies to the judge too.
    adjudication · calculé le 2026-07-13T11:55:28Z
    25

    canadian_linkage_misnames_itself

    LES DEUX CLAUSES DE « CANADIEN » MESURENT AUTRE CHOSE QUE CE QUE LEUR NOM ANNONCE. Le projet a consacré tous ses efforts à ce que le repérage MANQUE ; voici la première preuve solide de ce qu'il ADMET À TORT. (1) CA-AFF, clause principale de l'estimande PRINCIPALE : 17 466 travaux entrent dans la base à cause d'une « institution » canadienne qui n'existe pas. OpenAlex attribue « Discovery Air (Canada) » à un article brésilien de linguistique, « Musee de la Civilisation » à un article brésilien de sciences alimentaires et « Impact », qui n'est pas une institution mais un échec d'analyse muni d'un identifiant d'institution, à un essai espagnol sur la COVID. La voie CA-AFF contient aussi des milliers de travaux en letton et en indonésien. Ce nombre est une BORNE INFÉRIEURE : les chaînes artéfactuelles testées sont seulement celles que les agents de tri ont remarquées À L'ŒIL en faisant autre chose, et personne n'a parcouru tout le vocabulaire des institutions ; la vraie précision est donc INCONNUE, non simplement non mesurée. (2) ABOUT-CA, l'estimande SECONDAIRE : la voie de repérage signifie « le Canada apparaît dans le texte », tandis que le champ de la grille signifie « le SYSTÈME de recherche canadien est un objet d'étude substantiel » ; dans les strates entièrement fondées sur la première définition, la seconde ne s'active que pour 1,8 % à 3,5 % des travaux. Pire, `about_ca` ne peut être vrai que si le travail porte DÉJÀ sur la recherche, donc il est presque collinéaire avec le niveau : seulement 39 travaux sur 5 600 sont à la fois dans le champ ET consacrés au système de recherche canadien. L'ESTIMANDE SECONDAIRE NE PEUT PAS ÊTRE ESTIMÉE À PARTIR DE LA STRATE CONÇUE POUR ELLE. Quatre agents l'ont découvert indépendamment et ont proposé la même réparation : séparer le champ en `about_ca_system` et `about_ca_topic`, car un booléen ne distingue pas « données canadiennes, conclusion universelle » de « conclusion sur le Canada ».

    Énoncé original (findings.json) : BOTH CLAUSES OF 'CANADIAN' MEASURE SOMETHING OTHER THAN WHAT THEY ARE NAMED. This project has spent all its effort on what retrieval MISSES; this is the first hard evidence about what it wrongly ADMITS. (1) CA-AFF, the PRIMARY estimand's main clause: 17,466 works enter the frame on a Canadian 'institution' that does not exist. OpenAlex hands 'Discovery Air (Canada)' to a Brazilian linguistics paper, 'Musee de la Civilisation' to a Brazilian food-science paper, and 'Impact', which is not an institution at all but a parse failure with an institution id, to a Spanish COVID essay. The CA-AFF route also carries thousands of works in Latvian and Indonesian. That count is a LOWER BOUND: the artifact strings tested are only the ones screening agents noticed BY EYE while doing something else, and nobody has swept the institution vocabulary, so the true precision is UNKNOWN rather than merely unmeasured. (2) ABOUT-CA, the SECONDARY estimand: the retrieval route means 'Canada appears in the text' and the rubric field means 'the Canadian research SYSTEM is a substantive object of study', and in the strata built entirely on the former, the latter fires on 1.8% to 3.5% of works. Worse, `about_ca` cannot be true unless the work is ALREADY about research, so it is nearly collinear with tier: only 39 works in 5,600 are both in-scope AND about the Canadian research system. THE SECONDARY ESTIMAND CANNOT BE ESTIMATED FROM THE STRATUM DESIGNED FOR IT. Four screening agents found this independently and all four proposed the same repair: split the field into `about_ca_system` and `about_ca_topic`, because one boolean cannot separate 'Canadian data, universal claim' from 'a claim about Canada'.

    Champ source:ca_aff_works
    2 714 734
    Champ source:ca_aff_only_institution_is_artifact
    17 466
    Champ source:artifact_strings_tested
    • Impact
    • Discovery Air (Canada)
    • Musée de la Civilisation
    • Encana (Canada)
    • Kellogg's (Canada)
    • The Alberta Paraplegic Foundation
    Champ source:artifact_count_is_a_lower_bound
    oui
    Champ source:ca_aff_non_official_languages
    • 8 734
    • 8 463
    • 5 289
    • 5 160
    • 4 708
    • 1 200
    Champ source:about_ca_pct_in_about_strata_opus
    • 1,8
    • 3,5
    Champ source:about_ca_pct_in_about_strata_gpt
    • 1
    • 2
    Champ source:about_ca_pct_in_about_strata_grok
    • 0,8
    • 1,2
    Champ source:n_screened
    5 600
    Champ source:about_ca_and_in_scope
    39
    Champ source:about_ca_and_out_of_scope
    26
    Champ source:about_ca_nearly_collinear_with_tier
    oui
    Champ source:found_by
    Screening agents, reporting records they were given for an unrelated reason. Four of them independently reported that the ABOUT-CA route and the rubric's about_ca field are different constructs, and all four proposed the same repair without having seen each other's reports.
    Champ source:caveat
    The artifact count is a LOWER BOUND, and deliberately reported as one: the strings tested are only those agents noticed by eye. No sweep of the OpenAlex institution vocabulary has been done, so the true CA-AFF precision is unknown, not merely unmeasured. That is the honest state and it is why the human audit samples the retrieved stratum as well as the non-retrieved one: precision is measured, not assumed.
    canadian_linkage_misnames_itself · calculé le 2026-07-13T11:55:27Z
    26

    classifier

    LE CLASSIFICATEUR NE PEUT PAS ATTRIBUER D'ÉTIQUETTES, ET LES NOMBRES DISENT POURQUOI. Avec pondération du plan, validation hors pli et regroupement par lieu de publication, les têtes enseignantes atteignent une AP de 0,17 à 0,135 pour une AUC d'environ 0,8, avec un gain de 4,68 fois dans le décile supérieur qui est réel mais orienté vers des étiquettes de MACHINE. L'ESTIMANDE DÉPLACE LE SCORE D'UN FACTEUR DE 2,7 (unanimité 0,0783, majorité 0,1364, n'importe lequel 0,2093) : il n'existe donc pas une cible unique, et en choisir une trancherait silencieusement la question que l'audit doit étudier. LE NOYAU UNANIME EST LE PLUS DIFFICILE À APPRENDRE, non le plus facile : si la frontière était une ligne commune entourée de bruit, le noyau d'accord serait la partie séparable, or il ne l'est pas. L'écart entre les têtes prédit les divisions entre enseignants à 3,74 fois le taux de base : signal réel mais faible, présenté comme un gain d'efficacité plutôt que comme une carte. Enfin, le résumé vaut PLUS QUE TOUT CHOIX DE MODÉLISATION (+0,1145 AP, soit 84 % en relatif), mais la base ne peut pas le fournir : works.csv ne contient que le booléen has_abstract, sans texte. Le modèle DÉPLOYABLE est le MOINS BON ; la construction échoue désormais si l'entraînement utilise un champ absent à l'inférence, et cet écart devient une décision budgétaire chiffrée plutôt qu'une hypothèse.

    Énoncé original (findings.json) : THE CLASSIFIER MAY NOT LABEL, AND THE NUMBERS SAY WHY. Design-weighted, out-of-fold, venue-grouped: the teacher heads reach AP 0.17 to 0.135 at AUC ~0.8, with a 4.68x top-decile lift that is real and is lift toward MACHINE labels. THE ESTIMAND MOVES THE SCORE 2.7x (unanimous 0.0783, majority 0.1364, any 0.2093), so there is no 'the' target and choosing one silently settles the question the audit exists to answer. AND THE UNANIMOUS CORE IS THE HARDEST TO LEARN, not the easiest: if the boundary were a shared line with noise around it, the agreed core would be the separable part, and it is not. The between-head spread predicts where the teachers split at 3.74x the base rate: real, weak, and reported as an efficiency gain rather than a map. Finally, the abstract is worth MORE THAN ANY MODELING CHOICE (+0.1145 AP, 84% relative), and the frame cannot serve it: works.csv stores has_abstract, a boolean, and no text. The DEPLOYABLE model is the WORSE one, the build now FAILS if a model is trained on a field inference cannot supply, and the gap is a priced budget decision rather than an assumption.

    Champ source:evaluation
    out-of-fold, venue-grouped, vectorizer fitted on train folds only, DESIGN-WEIGHTED
    Champ source:why_design_weighted
    The 5,600 works are a stratified sample: French carries weight 311, aff_core 1,119. An unweighted AP is an AP over a population that does not exist. Both are reported; the gap is a fact about the design.
    Champ source:ap_design_weighted_by_teacher
    opus:
    0,17
    gpt:
    0,135
    grok:
    0,1417
    Champ source:ap_design_weighted_by_consensus_rule
    unanimous:
    0,0783
    majority:
    0,1364
    any:
    0,2093
    Champ source:estimand_moves_performance_x
    2,7
    Champ source:unanimous_core_is_hardest_to_learn
    oui
    Champ source:top_decile_lift
    opus:
    4,68
    gpt:
    3,93
    grok:
    5,28
    Champ source:ap_by_language_opus
    en:
    0,1757
    fr:
    0,181
    Champ source:ap_by_language_gpt
    en:
    0,1425
    fr:
    0,089
    Champ source:french_is_model_dependent
    GPT's boundary is markedly less learnable in French (AP 0.089 vs 0.1425 in English) while Opus's is not ( 0.181 vs 0.1757 ). 'Character n-grams do the French work' was an assertion; measured per language, it is true for one teacher and false for another.
    Champ source:predicting_teacher_disagreement
    n_contested:
    89
    base_rate:
    0,0159
    average_precision:
    0,0594
    lift_over_base:
    3,74
    verdict:
    Real but weak. It buys an efficiency gain for the audit's disagreement stratum. It is NOT a trustworthy standalone map of the contested region, and shared teacher bias is invisible to it: if all three are wrong the same way the spread is zero and the work looks settled.
    Champ source:payload_skew
    ap_frame_parity:
    0,1364
    ap_with_abstract:
    0,2509
    abstract_is_worth_ap:
    0,1145
    why_the_deployable_model_is_the_worse_one:
    data/db/works.csv stores has_abstract, a boolean, and no abstract text. A model trained with abstracts scores a payload the 4.3M works do not have. Reporting the abstract model's number as the classifier's performance would be reporting a metric for a model that cannot run. The difference is what re-extracting abstracts from the snapshot would buy, and it is a budget decision with a number attached rather than an assumption.
    Champ source:ships_a_category_label
    non
    Champ source:why_no_label
    Distilled from our own metaresearch labels, a student reproduced its own teacher's positive set at Jaccard 0.17 (finding 25). Two independent adversarial reviews reached the same verdict from different directions: per-teacher heads are useful for ALLOCATING HUMAN EFFORT and cosmetic as a solution to the contested boundary. Scores ship. Labels do not.
    Champ source:what_the_classifier_is_actually_for
    Not cost. Finding 13 says the LLM can read every work in the frame for $1,261, so there was never anything to avoid, and a proposal that budgeted the full screen while justifying a cheap substitute for it was describing two studies (D32). It is for: a calibrated score the audit stratifies on; rubric revisions testable in minutes instead of one $1,261 pass each; a frame-wide estimate of where the three teachers would split, which three full passes would cost 3x and ~25 days; and learnability as evidence about the boundary.
    classifier · calculé le 2026-07-15T22:32:09Z
    27

    distillation_ceiling

    UN CLASSIFICATEUR ENTRAÎNÉ SUR LES ÉTIQUETTES DES GRANDS MODÈLES DE LANGAGE NE PEUT MÊME PAS REPRODUIRE SON PROPRE ENSEIGNANT, ENCORE MOINS EN ARBITRER TROIS. Les élèves distillés de chaque modèle atteignent un indice de Jaccard de 0,17 avec leur propre enseignant, alors que les trois enseignants s'accordent ENTRE EUX à 0,50 à 0,56 : le classificateur approxime moins bien Opus que Grok. La distillation d'Opus éloigne aussi de Grok (-0,35) ; elle ne comble donc pas le désaccord entre modèles, elle s'éloigne de tous. LA BOUCLE D'AUTO-APPRENTISSAGE A ENSUITE RÉFUTÉ MA PROPRE PRÉDICTION. J'avais soutenu qu'elle réduirait la classe positive vers un noyau anglophone assuré ; appliquée à 20 000 travaux jamais vus par les modèles, elle ne l'a pas fait : dérive du taux positif de +0,00 point et de la part française de +0,00 point, car la pondération des classes empêche l'effondrement, exactement comme l'examen contradictoire l'avait averti avant l'exécution (« un risque sérieux, PAS UN THÉORÈME »). La boucle a plutôt atteint immédiatement un point fixe et recyclé ses propres étiquettes (224 puis 424 positifs d'entraînement, ensuite plus rien) : elle A CONVERGÉ SANS APPRENDRE, et un point fixe se ressemble de l'intérieur, que ses étiquettes soient justes ou fausses. La valeur réelle du classificateur est la stratification : 48,2 % des travaux dans le champ se trouvent dans son décile supérieur contre une référence aveugle de 10 %, soit un gain de 4,8 fois par unité de codage humain, tandis que les estimations pondérées par le plan restent non biaisées malgré son bruit. LA MACHINE DÉCIDE OÙ LES HUMAINS REGARDENT. ELLE N'ATTRIBUE JAMAIS D'ÉTIQUETTE.

    Énoncé original (findings.json) : A CLASSIFIER TRAINED ON LLM LABELS CANNOT EVEN REPRODUCE ITS OWN TEACHER, LET ALONE ADJUDICATE THREE. Students distilled from each model match their own teacher at Jaccard 0.17, while the three teachers match EACH OTHER at 0.5 to 0.56: the classifier is a worse approximation of Opus than Grok is. Distilling from Opus also moves you AWAY from Grok (-0.35), so distillation does not bridge the models' disagreement, it degrades away from all of them. THE SELF-TRAINING LOOP THEN REFUTED MY OWN PREDICTION. I argued at length that it would shrink the positive class toward the confident anglophone core; run on 20,000 works the models never saw, it did not: positive rate drift +0.00 points, French share drift +0.00 points, because class weighting prevents the collapse, exactly as the adversarial review warned before the run ('a serious risk, NOT A THEOREM'). What the loop DID do is reach a fixed point immediately and recycle its own labels (224 -> 424 training positives, then nothing): it CONVERGED WITHOUT LEARNING, and a fixed point looks identical from the inside whether its labels are right or wrong. What the classifier IS worth is the stratifier: 48.2% of the in-scope works sit in its top decile against a blind baseline of 10%, a 4.8x lift in in-scope works found per unit of human coding effort, and design-weighted estimates stay unbiased however noisy it is. THE MACHINE DECIDES WHERE THE HUMANS LOOK. IT NEVER SUPPLIES A LABEL.

    Champ source:n_three_way_labelled
    5 600
    Champ source:student_vs_own_teacher_jaccard
    • 0,176
    • 0,173
    • 0,17
    Champ source:teacher_vs_teacher_jaccard
    • 0,5
    • 0,5
    • 0,563
    Champ source:bridging_gain
    • -0,348297213622291
    • -0,348214285714286
    • -0,3375
    Champ source:selftrain_rounds
    • 0
    • 1
    • 2
    • 3
    Champ source:selftrain_train_positives
    • 224
    • 424
    • 424
    • 424
    Champ source:selftrain_pool_positive_rate_pct
    • 1
    • 1
    • 1
    • 1
    Champ source:selftrain_french_share_pct
    • 6,5
    • 6,5
    • 6,5
    • 6,5
    Champ source:selftrain_drift_positive_rate_pts
    0
    Champ source:selftrain_drift_french_share_pts
    0
    Champ source:my_shrinkage_prediction_was_refuted
    oui
    Champ source:stratifier_top_decile_recall_pct
    48,2
    Champ source:stratifier_lift_vs_blind
    4,8
    Champ source:used_to_produce_any_estimate
    non
    Champ source:caveat
    The student is TF-IDF word+char n-grams with a cross-validated logistic head: the cheap classifier the question was actually about. A fine-tuned multilingual encoder would raise the student-vs-teacher number and change none of the argument, because the ceiling is set by the LABELS, not by the model class. The self-training loop's non-collapse is CONDITIONAL on class weighting and should not be read as a general safety result: it says the collapse is avoidable, not that the loop is informative. And nothing here produces a prevalence estimate; the classifier is a stratifier, and the human audit is the instrument.
    distillation_ceiling · calculé le 2026-07-15T22:58:15Z
    28

    gemma_gate

    LE MODÈLE GRATUIT A ÉCHOUÉ AU SEUIL, MAIS LE SEUIL VISAIT LE MAUVAIS NIVEAU. Gemma-4-31B a trié 700 travaux selon la grille v3.1 avec 100 % de sorties conformes au schéma ; son indice de Jaccard atteint 0,44 pour la métarecherche et 0,515 pour toute catégorie contre l'union des modèles de pointe, sous le seuil prédéfini de 0,60. IL N'ÉTIQUETTE DONC PAS LA BASE : la règle demeure, car une règle déplacée après l'observation des nombres n'est plus une règle. Toutefois, sur les mêmes travaux, les modèles de pointe ne s'accordent entre eux qu'à 0,384 et 0,36, parce que chaque lot de la boucle est sélectionné POUR le désaccord : le seuil exigeait d'un modèle gratuit ce que les modèles de pointe n'atteignent pas entre eux, et Gemma a dépassé leur accord mutuel sur les deux cibles. Le test 2 est prédéfini sur la strate random_baseline, avec le raisonnement comme seuil : J(gemma, union) >= J(opus, gpt).

    Énoncé original (findings.json) : THE FREE MODEL FAILED THE GATE, AND THE GATE WAS AIMED AT THE WRONG BAR. Gemma-4-31B screened 700 works under rubric v3.1 (100% schema-clean) and scored Jaccard 0.44 on metaresearch and 0.515 on any-category against the frontier union, under the prespecified 0.60 bar, SO IT DOES NOT LABEL THE FRAME: the rule stands because rules that move after the numbers exist are not rules. But the frontier models agree with each other at only 0.384 and 0.36 on the same works, because every loop batch is selected FOR disagreement: the gate demanded of a free model what the frontier models do not achieve among themselves there, and Gemma beat the frontier's own mutual agreement on both targets. Test 2 is prespecified on the random_baseline stratum with the rationale as the bar: J(gemma, union) >= J(opus, gpt).

    Champ source:model
    google/gemma-4-31b-it (free)
    Champ source:n_works
    700
    Champ source:schema_clean_return_rate
    1
    Champ source:jaccard_metaresearch_gemma_vs_union
    0,44
    Champ source:jaccard_metaresearch_opus_vs_gpt
    0,384
    Champ source:jaccard_anycat_gemma_vs_union
    0,515
    Champ source:jaccard_anycat_opus_vs_gpt
    0,36
    Champ source:study_design_exact
    gemma_vs_opus:
    0,643
    gemma_vs_gpt:
    0,614
    opus_vs_gpt_for_reference:
    0,73
    Champ source:prespecified_rule
    Jaccard >= 0.60 vs the frontier union, on metaresearch AND any-category
    Champ source:passed
    non
    Champ source:rule_honored
    oui
    Champ source:gate_was_miscalibrated_because
    the 0.60 bar was set from round-001's 0.44-0.75 pairwise range without registering that loop batches are selected FOR disagreement, which depresses every agreement statistic computed on them. On the same 700 works the frontier models agree with each other at 0.384/0.360: below the bar Gemma was held to, and Gemma beat both numbers.
    Champ source:test2_prespecified
    On the enriched sample's random_baseline stratum (1,500 works, not selected for disagreement), Gemma passes if J(gemma, frontier_union) >= J(opus, gpt) on both metaresearch and any-category. Written before any of those labels exist.
    gemma_gate · calculé le 2026-07-14T00:08:32Z
    29

    instrument_contradicts_itself

    L'INSTRUMENT VERROUILLÉ SE CONTREDIT, ET RIEN NE L'A DÉTECTÉ. La grille et le schéma de sortie définissent DEUX vocabulaires contrôlés DIFFÉRENTS pour le même champ, `genre`, avec seulement 2 valeurs communes (empirical et other) ; 4 termes n'existent que dans la grille et 7 seulement dans le schéma. Chaque trieur a reçu les deux documents avec ordre de respecter les deux, puis a inventé sa propre conciliation : GPT-5.6 et Grok ont suivi la grille et produisent donc 26,9 % et 20,9 % de valeurs ILLÉGALES selon le schéma, tandis qu'Opus a puisé dans les deux listes et émis 13 valeurs distinctes. Le validateur contrôlait `tier` mais jamais `genre`, si bien que 16 800 étiquettes ont réussi tous les contrôles exécutés. Un agent l'a découvert en le mentionnant dans une seule proposition d'un rapport sur autre chose. Rien n'a planté ; la variable avait simplement un sens différent dans chaque volet, et la variance aurait été attribuée aux modèles. Les étiquettes ne sont PAS réparées, car harmoniser les volets après avoir observé leur divergence détruirait la seule preuve de cette divergence ; le champ du genre est déclaré inutilisable et la réparation appartient à l'instrument, à une frontière de version.

    Énoncé original (findings.json) : THE LOCKED INSTRUMENT CONTRADICTS ITSELF, AND NOTHING CAUGHT IT. The rubric and the output schema name TWO DIFFERENT controlled vocabularies for the same field, `genre`, overlapping on 2 values (empirical and other); 4 terms exist only in the rubric and 7 only in the schema. Every screener was handed both and told to obey both, and each invented its own reconciliation: GPT-5.6 and Grok followed the rubric and are therefore 26.9% and 20.9% ILLEGAL against the schema, while Opus drew from both lists at once and emitted 13 distinct values. The validator checked `tier` and never checked `genre`, so 16,800 labels passed every check that ran. It was found by an agent mentioning it in one clause of a report about something else. Nothing crashed; the variable simply meant a different thing in each arm, and the variance would have been attributed to the models. The labels are NOT repaired, because harmonizing the arms after seeing them would destroy the only evidence that they diverged; the genre field is reported as unusable and the fix belongs in the instrument, at a version boundary.

    Champ source:field
    genre
    Champ source:rubric_vocabulary
    • empirical
    • conceptual
    • editorial/commentary
    • policy
    • infrastructure/announcement
    • other
    Champ source:schema_vocabulary
    • empirical
    • review
    • methods
    • commentary
    • editorial
    • protocol
    • dataset
    • software
    • other
    Champ source:shared_values
    • empirical
    • other
    Champ source:n_shared
    2
    Champ source:only_in_rubric
    • conceptual
    • editorial/commentary
    • policy
    • infrastructure/announcement
    Champ source:only_in_schema
    • review
    • methods
    • commentary
    • editorial
    • protocol
    • dataset
    • software
    Champ source:arms
    • opus
    • gpt
    • grok
    Champ source:n_labels
    • 5 600
    • 5 600
    • 5 600
    Champ source:n_distinct_values_by_arm
    • 13
    • 6
    • 12
    Champ source:pct_legal_against_schema
    • 90,8
    • 73,1
    • 79,1
    Champ source:pct_legal_against_rubric
    • 86,1
    • 100
    • 99,8
    Champ source:validator_checked_tier
    oui
    Champ source:validator_checked_genre
    non
    Champ source:labels_repaired
    non
    Champ source:found_by
    An Opus screening agent mentioned it in one clause of a report about something else, while working chunks it had been given for an unrelated reason. It was not looking for this, no check was watching for it, and it had already survived 6,000 labels across three models.
    Champ source:caveat
    The genre variable from this screen is reported as UNUSABLE and is used for nothing. It is not remapped to a common vocabulary: the three arms resolved the contradiction differently, and harmonizing them after the fact would destroy the only evidence that they did.
    instrument_contradicts_itself · calculé le 2026-07-15T22:32:11Z
    30

    the_frame

    La base est CONSTRUITE, non estimée. Les 482 partitions d'un instantané OpenAlex épinglé ont été parcourues et filtrées en continu ; la base canadienne contient 4 299 418 travaux, chacun exactement une fois. Pendant la majeure partie du projet, « la base » comptait 3 507 205 travaux, soit une EXTRAPOLATION d'une seule partition codée en dur dans six scripts ; l'estimation était TROP BASSE DE 18 %, et tous les chiffres de coût, de puissance de l'audit et de taille du domaine calculés à partir d'elle ont changé. La voie décisive : 1 565 226 travaux, soit 36,4 % de la base, sont INVISIBLES À L'AFFILIATION SEULE. Une base fondée sur l'affiliation canadienne contiendrait 2 734 192 travaux et ne les verrait jamais. Voilà le constat 2 à l'échelle de la base ; c'est pourquoi celle-ci réunit quatre voies et pourquoi chaque notice porte la provenance de la voie qui l'a admise.

    Énoncé original (findings.json) : The frame is BUILT, not estimated. All 482 partitions of a pinned OpenAlex snapshot were streamed and filtered, and the Canadian frame holds 4,299,418 works, each exactly once. For most of this project's life 'the frame' was 3,507,205, an EXTRAPOLATION from a single partition, hardcoded in six scripts; the estimate was 18% LOW, and every cost, audit-power and field-size figure computed against it has moved. The route that matters: 1,565,226 works (36.4% of the frame) are INVISIBLE TO AFFILIATION ALONE. A frame built on Canadian affiliation would hold 2,734,192 works and would never see them. That is finding 2 at frame scale, and it is why the frame is a union of four routes and why every record carries the provenance of the route that admitted it.

    Champ source:built_from
    all 482 partitions of a pinned OpenAlex snapshot, streamed and filtered; each work appears exactly once
    Champ source:partitions
    482
    Champ source:works
    4 299 418
    Champ source:prior_estimate
    3 507 205
    Champ source:estimate_was_low_by_pct
    18
    Champ source:estimate_was_an_extrapolation_from_one_partition
    oui
    Champ source:route_ca_aff
    2 734 192
    Champ source:route_about_ca
    1 316 408
    Champ source:route_ca_fund
    788 103
    Champ source:route_ca_venue
    638 161
    Champ source:invisible_to_affiliation
    1 565 226
    Champ source:pct_invisible_to_affiliation
    36,4
    Champ source:with_venue
    3 902 048
    Champ source:pct_with_venue
    90,8
    Champ source:with_abstract
    3 296 301
    Champ source:pct_with_abstract
    76,7
    Champ source:pct_no_abstract
    23,3
    Champ source:french
    237 207
    Champ source:pct_french
    5,5
    Champ source:built_utc
    2026-07-12T17:59:59Z
    Champ source:caveat
    The frame is bounded by OpenAlex. A work whose Canadian link is invisible to the metadata (no affiliation, no funder, no textual mention, no Canadian venue) cannot enter ANY frame by any method, and scholarship indexed by neither OpenAlex nor Erudit is not estimated here. That is a hard boundary of the data, not of this design, and it is stated rather than hidden. Erudit matches ZERO OpenAlex sources (finding 3), so the francophone share reported here (5.5%) measures the pipeline, NOT Canadian scholarship.
    the_frame · calculé le 2026-07-13T11:55:21Z
    31

    v1_to_v2

    ÉCRIRE LES PHRASES MANQUANTES A RÉSOLU 54 % DES DÉSACCORDS CONTRE LESQUELS ELLES AVAIENT ÉTÉ ÉCRITES. Les mêmes trois modèles de pointe ont retrié sous la grille v2 les 179 travaux qui les avaient divisés sous la v1 ; les règles de la v2 provenaient d'un recensement publié et pondéré par fréquence des PHRASES ABSENTES DE LA GRILLE. Parmi les 179 travaux, 96 sont maintenant unanimes, et l'indice de Jaccard par paire pour les ensembles dans le champ passe de 0,19 à 0,36. Cette mesure, presque jamais réalisée, n'existe QUE parce que l'ambiguïté a été énumérée, arbitrée et PUBLIÉE AVANT sa résolution : le recensement est sorti en premier, empêchant d'ajuster discrètement la v2 jusqu'à obtenir de bons chiffres. PORTÉE, énoncée sans l'estomper : ces travaux ont été SÉLECTIONNÉS POUR LEUR DÉSACCORD ; le résultat montre donc que les règles résolvent les cas pour lesquels elles ont été écrites, mais il ne montre PAS que l'accord a augmenté dans toute la base, car la seule régression vers la moyenne déplacerait un ensemble ainsi sélectionné. Le test non biaisé est un retri complet des 5 600 travaux sous la v2, prédéfini, et c'est ce que réalisera le projet financé.

    Énoncé original (findings.json) : WRITING THE MISSING SENTENCES RESOLVED 54% OF THE DISAGREEMENTS THEY WERE WRITTEN AGAINST. The 179 works three frontier models split on under rubric v1 were re-screened, by the same three models, under a v2 whose rules were written against a published, frequency-weighted census of WHICH SENTENCES THE RUBRIC WAS MISSING. 96 of 179 are now unanimous, and the pairwise Jaccard of the in-scope sets moves from 0.19 to 0.36. This is a measurement almost nobody makes, and it is available ONLY because the ambiguity was enumerated, adjudicated and PUBLISHED BEFORE it was resolved: the census went out first, so v2 could not be quietly tuned until the numbers looked good. SCOPE, stated rather than blurred: these works were SELECTED FOR DISAGREEMENT, so this says the rules resolve the cases they were written for; it does NOT say frame-wide agreement rose, because regression to the mean alone would move a set selected this way. The unbiased test is a full re-screen of all 5,600 under v2, prespecified, and it is what the funded work runs.

    Champ source:contested_under_v1
    179
    Champ source:unanimous_under_v2
    96
    Champ source:pct_resolved
    54
    Champ source:still_contested
    83
    Champ source:in_scope_by_model_v1
    • 129
    • 68
    • 53
    Champ source:in_scope_by_model_v2
    • 73
    • 32
    • 44
    Champ source:jaccard_v1_mean
    0,187
    Champ source:jaccard_v2_mean
    0,362
    Champ source:sequence
    lock v1; screen 5,600 works with three models; have an independent blinded judge adjudicate every disagreement AND name the seam that caused it; PUBLISH the seam census; write v2 against the census and lock it BEFORE re-screening; re-screen; report the difference, whatever it is. Publishing the census before resolving it is what stops v2 being tuned until the numbers improve.
    Champ source:caveat
    SELECTED FOR DISAGREEMENT. These 179 works are exactly the ones the three models split on under v1, so this measures whether v2's rules resolve the cases they were written for. It CANNOT support a claim about agreement across the frame: regression to the mean alone would move a set selected this way. The unbiased test is a full re-screen of all 5,600 under v2, prespecified, and it is what the funded work runs.
    v1_to_v2 · calculé le 2026-07-13T11:55:28Z
    32

    zero_probability_region

    LE PLAN D'ÉCHANTILLONNAGE NE POUVAIT PAS ATTEINDRE 12,9 % DE LA BASE, ET AUCUN POIDS NE PEUT LE RÉPARER. Un plan stratifié repose sur une identité, sum(N_h) = N, qui n'avait jamais été vérifiée. Pour 549 370 travaux, soit 12,9 % de la base d'échantillonnage, la probabilité d'inclusion était EXACTEMENT NULLE : 366 856 n'étaient revendiqués par aucun prédicat parce que aff_core excluait tout ce qui portait SUR le Canada tandis que about_only excluait tout ce qui était AFFILIÉ au Canada, laissant entre les deux les travaux qui satisfaisaient les deux critères ; 182 514 autres avaient un prédicat évalué à SQL NULL, que `WHERE p` et `WHERE NOT p` refusent TOUS DEUX de sélectionner, de sorte qu'ils n'appartenaient à aucune strate et n'étaient même pas orphelins. Rien n'a échoué. Chaque strate a renvoyé exactement le nombre demandé, car elle ne peut pas connaître les travaux qu'on ne lui a jamais demandé de considérer, et le plan a produit un échantillon probabiliste impeccable DE 87 % DE LA BASE tout en donnant le nom de la base entière à chaque nombre calculé. Un travail à probabilité nulle n'est pas sous-pondéré ; il est INACCESSIBLE, et les poids du plan sont la garantie centrale du projet. LA CASE SUPPRIMÉE ÉTAIT CELLE DE L'ESTIMANDE SECONDAIRE : 328 912 travaux à affiliation canadienne portant sur le Canada, dans une étude dont l'estimande secondaire est la métarecherche sur le système de recherche canadien. La réparation ajoute le complément exact et sûr face à NULL sous forme de deux strates ; les 7 strates partitionnent maintenant la base par construction (4 255 410 = 4 255 410), ce qui est vérifié à chaque construction. Un modèle contradictoire l'a découvert en additionnant cinq nombres que je lui avais donnés. Le défaut rendait l'échantillon PLUS NET et la variance PLUS PETITE, d'où l'absence de signal d'alarme : les erreurs qui survivent sont celles qui vous flattent.

    Énoncé original (findings.json) : THE SAMPLING DESIGN COULD NOT REACH 12.9% OF THE FRAME, AND NO WEIGHT CAN FIX THAT. A stratified design rests on one identity, sum(N_h) = N. It was never checked. 549,370 works (12.9% of the sampling frame) had an inclusion probability of EXACTLY ZERO: 366,856 that no predicate claimed, because aff_core excluded everything ABOUT Canada while about_only excluded everything AFFILIATED with Canada, so a work that was both fell between them; and 182,514 more whose predicate evaluated to SQL NULL, which `WHERE p` and `WHERE NOT p` BOTH decline to select, so they were in no stratum and were not even orphans. Nothing threw. Every stratum returned exactly the n it asked for, because a stratum cannot know about the works it was never asked about, and the design drew a clean textbook probability sample OF 87% OF THE FRAME while every number computed from it said 'the frame'. A zero-probability work is not underweighted, it is UNREACHABLE, and design weights are the guarantee this project leans on hardest. THE CELL IT DELETED WAS THE SECONDARY ESTIMAND'S: 328,912 Canadian-affiliated works about Canada, in a study whose secondary estimand is 'metaresearch about the Canadian research system'. Repaired by adding the exact NULL-safe complement as two strata; the 7 strata now partition the frame by construction (4,255,410 = 4,255,410), asserted on every build. It was found by an adversarial model adding up five numbers I handed it. The defect made the sample TIDIER and the variance SMALLER, which is why nothing about it felt wrong: the errors that survive are the ones that flatter you.

    Champ source:full_frame
    4 299 418
    Champ source:unscreenable_excluded
    44 008
    Champ source:sampling_frame
    4 255 410
    Champ source:reachable_under_shipped_design
    3 706 040
    Champ source:orphaned_predicate_false
    366 856
    Champ source:invisible_predicate_null
    182 514
    Champ source:zero_probability_works
    549 370
    Champ source:pct_of_sampling_frame
    12,9
    Champ source:aff_and_about_cell
    328 912
    Champ source:strata_after_repair
    7
    Champ source:strata_sum_after_repair
    4 255 410
    Champ source:partition_holds
    oui
    Champ source:found_by
    An adversarial model asked to attack the classifier design. Its first move was to add up the five design weights in a summary table I had handed it: 2000x1119 + 1000x664.2 + 750x536.8 + 750x310.9 + 500x335.8 = 3,705,875, against a frame of 4,299,418. I had never added them up.
    Champ source:caveat
    The repair does not retro-fix numbers produced under the broken design; those estimated a 3.7M subpopulation and were reported under the frame's name, and saying so IS the finding. Every design-weighted figure is now re-derived against the seven-stratum design. The five original predicates are kept byte-for-byte because 5,000 works had already been drawn from them by hash order, and widening a stratum silently re-draws it.
    zero_probability_region · calculé le 2026-07-13T11:55:25Z