icepyx as an icebreaker: starting conversations and building competencies in open science
Bibliographic record
Abstract
AGU 2022 Fall Meeting Presentation (Invited) Abstract: Open science is critical for advancing cryospheric science. Open science improves our ability to conduct adaptive, rigorous research in response to rapidly changing systems, creates community, enhances cooperation, promotes diversity and inclusivity, and reduces computational challenges associated with analyzing increasingly large and complex datasets. Fostering open science practices as the norm requires discussion and compromise, yet initiating and sustaining conversations around this topic can be challenging and is often perceived as sub-par to science objectives. The NASA Transform to Open Science (TOPS) initiative provides training and resources to enable this transition across all NASA missions, aligning well with icepyx’s ongoing work initiated within the cryospheric science community. icepyx is a community and Python software library designed for working with large, complex data products collected on the second Ice, Cloud, and land Elevation Satellite (ICESat-2). As a widely applicable, domain agnostic tool, icepyx provides functionality for data access, visualization, processing, and analysis. Critically, it also serves to facilitate conversations around open science and demonstrates one way that collaborative development is leveraged to build shared tools. An important component of this is not only creating and fostering community, but introducing and building core open science skills in a relevant, applied framework for cryospheric scientists. icepyx also catalyzes cross-disciplinary research by lowering the barrier of entry through widely usable, modular, and extensible tools requiring minimal technical expertise for sustained engagement. This presentation will highlight the people and features behind icepyx and its relevance as a teaching tool at ICESat-2 themed hackweeks and for advancing open science. Plain Language Abstract: The world is changing rapidly, as are the tools we use to study it. This is particularly true for people who study snow and ice. For success, we need everyone to be able to participate and work together. When ideas, methods, and information are shared openly (called open science), anyone can see what is happening and contribute ideas. We won’t end up repeating work that has already been done. This requires learning how to communicate well - even with strangers. It can be scary to enter a new space, but friendly, welcoming, inclusive communities like icepyx are here to show you how! You can think about icepyx like a jigsaw puzzle. Many people know icepyx as a set of software tools (the pieces) for working with satellite data. But behind the code are people (those pieces will not put themselves together!). Each person has a piece of the puzzle. icepyx acts like the puzzle’s edge, providing structure and a shared end goal. The community can solve the puzzle together, piece by piece, to work towards the bigger picture. Our goal is for everyone to be able to contribute and enjoy the experience, answering interesting questions about snow and ice together!
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.012 | 0.012 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.007 | 0.003 |
| Scholarly communication | 0.011 | 0.011 |
| Open science | 0.002 | 0.020 |
| Research integrity | 0.004 | 0.008 |
| Insufficient payload (model declined to judge) | 0.069 | 0.023 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".