MétaCan
Menu
Back to cohort
Record W4414563984 · doi:10.1101/2025.09.17.676398

Model-based and model-free valuation signals in the human brain vary markedly in their relationship to individual differences in human behavioral control

2025· preprint· en· W4414563984 on OpenAlexaff
Weilun Ding, Jeffrey Cockburn, Julia Simon, Amogh Johri, Scarlet J. Cho, Sarah Oh, Jamie D. Feusner, Reza Tadayonnejad, John P. O’Doherty

Bibliographic record

VenuebioRxiv (Cold Spring Harbor Laboratory) · 2025
Typepreprint
Languageen
FieldComputer Science
TopicCognitive Science and Mapping
Canadian institutionsUniversity of TorontoCentre for Addiction and Mental Health
Fundersnot available
KeywordsReinforcementVentromedial prefrontal cortexAction selectionPrefrontal cortexReinforcement learningVentral striatumMean squared prediction errorNeural activityValuation (finance)

Abstract

fetched live from OpenAlex

Human action selection under reinforcement is thought to rely on two distinct strategies: model-free and model-based reinforcement learning. While behavior in sequential decision-making tasks often reflects a mixture of both, the neural basis of individual differences in their expression remains unclear. To investigate this, we conducted a large-scale fMRI study with 179 participants performing a variant of the two-step task. Using both cluster-defined subgroups and computational parameter estimates, we found that the ventromedial prefrontal cortex encodes model-based and model-free value signals differently depending on individual strategy use. Model-based value signals were strongly linked to the degree of model-based behavioral reliance, whereas model-free signals appeared regardless of model-free behavioral influence. Leveraging the large sample, we also addressed a longstanding debate about whether model-based knowledge is incorporated into reward prediction errors or if such signals are purely model-free. Surprisingly, ventral striatum prediction error activity was better explained by model-based computations, while a middle caudate error signal was more aligned with model-free learning. Moreover, individuals lacking both model-based behavior and model-based neural signals exhibited impaired state prediction errors, suggesting a difficulty in building or updating their internal model of the environment. These findings indicate that model-free signals are ubiquitous across individuals, even in those not behaviorally relying on model-free strategies, while model-based representations appear only in those individuals utilizing such a strategy at the behavioral level, the absence of which may depend in part on underlying difficulties in forming accurate model-based predictions.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.004
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMeta-epidemiology (narrow)
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.282
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0040.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0010.001
Science and technology studies0.0000.000
Scholarly communication0.0010.000
Open science0.0020.001
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.077
GPT teacher head0.290
Teacher spread0.212 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venuebioRxiv (Cold Spring Harbor Laboratory)Same topicCognitive Science and MappingFrench-language works237,207