MétaCan
Menu
Back to cohort
Record W4205783425 · doi:10.1002/ail2.62

Deep learning does not replace Bayesian modeling: Comparing research use via citation counting

2022· article· en· W4205783425 on OpenAlexaff
Breck Baldwin

Bibliographic record

VenueApplied AI Letters · 2022
Typearticle
Languageen
FieldComputer Science
TopicExplainable Artificial Intelligence (XAI)
Canadian institutionsPublic Safety Canada
FundersColumbia University
KeywordsArtificial intelligenceDeep learningComputer scienceSurpriseMachine learningCitationDominance (genetics)Data sciencePsychologyWorld Wide Web

Abstract

fetched live from OpenAlex

Abstract One could be excused for assuming that deep learning had or will soon usurp all credible work in reasoning, artificial intelligence, and statistics, but like most “meme” class broad generalizations the concept does not hold up to scrutiny. Memes do not generally matter since the experts will always know better; but in the case of Bayesian software like Stan and PyMC3, even their developers and advocates bemoan the apparent dominance of deep learning as manifested in popular culture, breathtaking performance, and most problematically from funding agency peer review that impacts our ability to further advance the field. The facts, however, do not support the assumed dominance of deep learning in science upon closer examination. This letter simply makes the argument by the crudest of possible metrics, citation count, that once the discipline of Computer Science is subtracted, Bayesian software accounts for nearly a third of research citations. Stan and PyMC3 dominate some fields, PyTorch, Keras, and TensorFlow dominate others with lot of variations in between. Bayesian and deep‐learning approaches are related but very different technologies in goals, implementation, and applicability with little actual overlap‐‐so this is not a surprise. For example, deep learning cannot bring the explainability of applied math/statistics and Bayesian methods do not scale to deep‐learning data sets. While deep‐learning behemoths like Facebook and Google use and support Bayesian efforts, the Bayesian packages scientists actually use are academic/volunteer efforts punching far above their weight class, and they need financial support. It would behoove funders to fully understand the impact and role of Bayesian methods in resource allocation.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.063
metaresearch head score (Gemma)0.447
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Bibliometrics
Consensus categoriesnone
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.955
Threshold uncertainty score0.332

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0630.447
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0020.002
Bibliometrics0.0450.075
Science and technology studies0.0020.003
Scholarly communication0.0120.018
Open science0.0020.006
Research integrity0.0030.002
Insufficient payload (model declined to judge)0.0050.002

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.075
GPT teacher head0.308
Teacher spread0.233 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designObservational
DomainMethods
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations2
Published2022
Admission routes1
Has abstractyes

Explore more

Same venueApplied AI LettersSame topicExplainable Artificial Intelligence (XAI)French-language works237,207