2023-2024 High school big data challenge: Leveraging generative ai and data cybersecurity to conserve and foster local biodiversityUnder the Patronage of the Canadian Commission for UNESCO
Bibliographic record
Abstract
The STEM Fellowship High School Big Data Challenge provides students with the unique opportunity of Open Data inquiry into one of the UN Sustainable Development Goals and experiential learning of fundamentals of data analysis – an essential skill set for a young researcher in the digital age. This year, students explore Generative AI and Data Cybersecurity to Conserve and Foster Local Biodiversity and to suggest their own evidence-based solutions following the principles of Open Science. They investigated different topics, ranging from Enhancing Forest Fire Predictions with Sequential Models for Ecosystem Preservation and Public Safety to Leveraging Semantic Segmentation to Perform Wildfire Prediction. We designed an interdisciplinary and agile educational environment, and in-depth learning modules for students as a means of bridging the gap between traditional high school courseware and digital reality and computational science. Students learned how to uncover hidden patterns and trends in structured and unstructured data using a range of data analytics tools and programming languages. Python, R, LaTeX, and machine learning were some of the tools the students learned and used. On behalf of the STEM Fellowship, we extend our sincere congratulations to all students who participated in the challenge, and wish them the best for their future endeavours. We want to express our appreciation to all the mentors and volunteers. This program would not be possible without patronage of CC UNESCO and generous support of our sponsors: RBC Future Launch, Let’s Talk Science, CISCO Networking Academy, Canadian Science Publishing, Schulich Foundation, SciNet at University of Toronto, and the University of Calgary Hunter Hub for Entrepreneurial Thinking. We were privileged to witness first-hand the analytical capabilities of the data-native generation of students, and we are confident they will demonstrate excellence throughout their academic and professional careers.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".