MétaCan
Menu
Back to cohort
Record W4221004076 · doi:10.5194/egusphere-egu22-3713

Fully virtual learning groups - pilot project on Machine Learning for early career researchers

2022· preprint· en· W4221004076 on OpenAlexaff
Julia Mindlin, Priyanka Yadav, Claudia Volosciuk, Valentina Rabanal, Faten Attig Bahar, Gerbrand Koren, Javed Ali, Claude-Michel Nzotungicimpaye

Bibliographic record

Venuenot available
Typepreprint
Languageen
FieldDecision Sciences
TopicScientific Computing and Data Management
Canadian institutionsConcordia University
Fundersnot available
KeywordsCurriculumChemistryMedicineMedical educationMathematics educationPsychologyPedagogy

Abstract

fetched live from OpenAlex

For early career researchers (ECRs), it is of utmost importance to acquire various skills including the application of different methods under the umbrella of data science. However, curricula of scientific degrees do not necessarily always include all relevant methods in the field, and there are also new methodologies emerging. Besides organized training schools, self-organized learning groups are common in universities for collaboratively acquiring new skills. Here, we present a concept that goes beyond in-person meetings and a prescribed curriculum to learn collaboratively, implemented for learning Machine Learning (ML) methods. There is growing interest in ML methods applied to Earth system science. These tools are being incorporated rapidly in the curricula of many scientific degrees, however, there is a generation of ECRs who did not learn to apply or work with ML while obtaining their masters or doctorate degrees and are now interested in filling this hiatus. The Young Earth System Scientists (YESS) Community, a network of ECRs working in Earth system sciences, has organized a learning activity to bring together members of our community who want to apply these methods to their own data and scientific problems and have little or no knowledge on ML. The main goal of this activity was to provide ECRs of our community the opportunity and platform to engage in a guided and collaborative learning process via the participation in small learning groups. The activity was implemented fully virtual. Additionally, the purpose of working in groups was to allow group discussions on how to interpret the results in combination with traditional physics-based methods/knowledge. Each group had a group leader which was in turn exchanging closely with other group leaders about the progress made and challenges encountered while keeping track of their group. The main challenges were working across time-zones, collaborative coding while learning, task distribution that ensured everyone learned from the activity. The activity not only proved to be useful for learning ML concepts, it was also a seedbed for projects which participants wish to continue working on. The skills and lessons learned from the organization included managing different time commitments among group members, working across time zones, learning-tasks distribution, ways to divide people into groups according to their research interests, advancing in knowledge coming from different backgrounds, writing a short proposal, literature review, providing a research project and reading material to stimulate an active learning mindset for students. Here, we show what tools and learning strategies were most successful, results from the research projects and lessons learned that can be useful for other groups, networks or even teachers when designing such learning activities.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.019
metaresearch head score (Gemma)0.015
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Other · Consensus signal: none
Teacher disagreement score0.067
Threshold uncertainty score0.224

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0190.015
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0010.001
Science and technology studies0.0030.002
Scholarly communication0.0030.004
Open science0.0050.009
Research integrity0.0030.004
Insufficient payload (model declined to judge)0.0670.030

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.454
GPT teacher head0.452
Teacher spread0.002 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreOther

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2022
Admission routes1
Has abstractyes

Explore more

Same topicScientific Computing and Data ManagementFrench-language works237,207