MétaCan
Menu
Landscape

The shape of the frame

Counts and trends over the frozen release. Design-based intervals arrive with the human audit; nothing on this page estimates field prevalence.

Works per publication year

2000 to 2025
200020122025
2000 baseline
103,185
2024 peak
314,132
Growth since 2000
×3.0
Works in 2025
300,048

Machine label counts

over 11,048 labeled works
Metaresearch1,914
Bibliometrics923
Research integrity570
Insufficient payload (model declined to judge)536
Scholarly communication407
Science and technology studies392
Open science262
Meta-epidemiology (broad)233
Meta-epidemiology (narrow)60

Machine labels · sparse coverage · An unlabeled work is unknown, not a negative. Label coverage is reported on every query.

Top fields

whole frame · OpenAlex primary field
Medicine932,321
Social Sciences519,440
Engineering372,416
Biochemistry, Genetics and Molecular Biology270,434
Environmental Science246,610

Language mix

whole frame
English 3,896,734 · 90.6%French 237,207 · 5.5%Other languages 165,477 · 3.8%

Landscape

What the machine-labeled subset of the frame looks like: categories, study designs, years, languages. Below it, the frame described by itself.

Read the coverage before the counts. Labels cover 11,048 of the 4,299,418 works in the frame (0.257%). Every count in this section is over that labeled subset only; it says what the labeled works look like, never how much of the frame is in a category. These are direct model labels, unvalidated, and an unlabeled work is NOT a negative.
Labeled works
11,048
works with at least one model label
Label rows
22,696
one per (work, model) pair
Models
gemma · gpt · grok · opus
each work is labeled by up to three

Labeled works by category

A work counts under a category if at least one model applied it; the darker figure requires every model that labeled the work to agree. The gap between the two columns is the models disagreeing, and that gap is a finding, not noise.

CategoryAny modelAll models agree
Metaresearch1,914963
Bibliometrics923808
Research integrity570348
Insufficient payload (model declined to judge)536154
Scholarly communication407163
Science and technology studies392184
Open science262136
Meta-epidemiology (broad)23397
Meta-epidemiology (narrow)6022

Labeled works by study design

Same two readings: any model, and all models in agreement. No design label here is MEDLINE-validated yet; when that validation lands it will be marked explicitly.

Study designAny modelAll models agree
Not applicable3,9811,966
Observational3,5702,145
Other design1,97244
Theoretical or conceptual1,447613
Qualitative1,051536
Bench or experimental843539
Systematic review819467
Simulation or modeling506242
Randomized trial454321
Meta-analysis206124
Non-randomized trial11024
Case report8656

Labeled works by year

Where the labeling rounds have reached so far. This is coverage of the label table, not a property of the field.

YearLabeled works
2000114
2001106
2002126
2003125
2004155
2005199
2006223
2007236
2008244
2009252
2010267
2011311
2012325
2013342
2014364
2015455
2016432
2017498
2018585
2019546
2020697
2021766
2022767
2023909
2024966
20251,038

Labeled works by language

Coverage again: the labeling rounds sample the frame, and the frame is 6% French.

LanguageLabeled works
en9,974
fr839
unknown134
es36
pt16
de12
id7
lv5
ja3
nl3
ro2
it2
tr2
hr1
cs1
uk1
ko1
no1
fa1
ca1
he1
ar1
fi1
zh1
sl1
el1

The frame itself

Everything below is over all 4,299,418 works, labeled or not. Every figure is a query against the database, computed at request time and cached for an hour; nothing here is a number someone typed.

Works in the frame
4,299,418
No Canadian affiliation
1,565,226
36.4% of the frame
No abstract
1,003,117
23.3% of the frame
Retraction notices
1,052
joined from Retraction Watch

Works by year

The frame over time, with the works that carry NO Canadian affiliation drawn underneath. The gap between the two lines is what an affiliation-only frame silently loses: 36.4% of the frame, 1,565,226 works.

Works by route

Why each work is in the frame. The four routes OVERLAP; a work can be admitted by several. These bars therefore sum to 5,476,864, which is 1,177,446 more than the 4,299,418 works in the frame. That overlap is the next chart.

The overlap: exact route combinations

Each work counted once, under the exact set of routes that admitted it. Teal bars are works admitted by a SINGLE route: remove that route from the design and those works vanish from the frame entirely.

Works by field

OpenAlex's primary field, as recorded.

Works by language

French is highlighted. It is 6% of the frame, it is oversampled in the screen on purpose, and it is the language the abstract cascade rescues worst (15.4% recovery against 38.8% for English).

The abstract gap is structural, not noise

Share of works with NO abstract, by type, worst first. 23.3% of the frame has no abstract, and the screen finds HALF as much metaresearch there. If the gap were random, a better index would fix it. It is not random: it is concentrated in types that never carry an abstract at all. "Just screen the works that have abstracts" is therefore a selection on a covariate that predicts the outcome.

The post-publication record has four states, and OpenAlex has a boolean

1,052 works in the frame carry a Retraction Watch notice. The solid bar is what OpenAlex flags; the hatched bar is what it reports as “false”: 143 works whose notice OpenAlex does not carry, which a reader takes to mean “fine”. An expression of concern is not a retraction, and `is_retracted` has no way to say so.

StateWorksOpenAlex flags itOpenAlex reports false
Retraction95590649
Expression of concern52052
Correction34232
Reinstatement11110

Top venues

By work count in the frame.

Top funders

Split from the semicolon-separated funder string. The funder route admits works that carry no Canadian affiliation at all.

Teacher spread over the frame (provisional baseline)

Every one of the 4,299,418 works carries two provisional teacher-head scores, and the spread is how far the two heads sit apart on one work. These figures come from pilot/results/frame_scores.json, the file the scoring run writes; nothing here was typed by hand.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Works scored
4,299,418
every work in the frame, by the two-teacher panel
Mean teacher spread
0.2432
the average disagreement between the two heads
Teacher spread, 99th percentile
0.3963
99% of works sit below this spread
Works where the teachers would split
540
spread above 0.5

Every series here is available as JSON: see the API.