MétaCan
Menu
Back to cohort
Record W4390828111 · doi:10.1002/cem.3531

Implications of confounding from unmodeled interactions between explanatory variables when using latent variable regression models to make inferences

2024· article· en· W4390828111 on OpenAlexaff
Olav M. Kvalheim, Warren S. Vidar, Tim U. H. Baumeister, Roger G. Linington, Nadja B. Cech

Bibliographic record

VenueJournal of Chemometrics · 2024
Typearticle
Languageen
FieldChemistry
TopicSpectroscopy and Chemometric Analyses
Canadian institutionsSimon Fraser University
FundersNational Center for Complementary and Integrative HealthNational Institutes of HealthOffice of Dietary Supplements
KeywordsLatent variableConfoundingEconometricsRegression analysisStatisticsRegressionVariable (mathematics)Latent variable modelMathematicsLinear regressionComputer science

Abstract

fetched live from OpenAlex

With linear dependency between the explanatory variables, partial least squares (PLS) regression is commonly used for regression analysis. If the response variable correlates to a high degree with the explanatory variables, a model with excellent predictive ability can usually be obtained. Ranking of variable importance is commonly used to interpret the model and sometimes this interpretation guides further experimentation. For instance, when analyzing natural product extracts for bioactivity, an underlying assumption is that the highest ranked compounds represent the best candidates for isolation and further testing. A problem with this approach is that in most cases the number of compounds is larger than the number of samples (and usually much larger) and that the concentrations of the compounds correlate. Furthermore, compounds may interact as synergists or as antagonists. If the modelling process does not account for this possibility, the interpretation can be thoroughly wrong since unmodelled variables that strongly influence the response will give rise to confounding of a first order PLS model and send the experimenter on a wrong track. We show the consequences of this by a practical example from natural product research. Furthermore, we show that by including the possibility of interactions between explanatory variables, visualization using a selectivity ratio plot may provide model interpretation that can be used to make inferences.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.151
metaresearch head score (Gemma)0.360
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Theoretical or conceptual · Consensus signal: Theoretical or conceptual
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.151
Threshold uncertainty score0.800

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.1510.360
Meta-epidemiology (narrow)0.0030.001
Meta-epidemiology (broad)0.0030.004
Bibliometrics0.0020.003
Science and technology studies0.0020.006
Scholarly communication0.0050.006
Open science0.0030.004
Research integrity0.0030.007
Insufficient payload (model declined to judge)0.0030.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.112
GPT teacher head0.363
Teacher spread0.251 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designTheoretical or conceptual
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations2
Published2024
Admission routes1
Has abstractyes

Explore more

Same venueJournal of ChemometricsSame topicSpectroscopy and Chemometric AnalysesFrench-language works237,207