Studying How Cryptocurrency Development Characteristics in GitHub Affect Its Market Price and Developer Sentiment in Stack Overflow Discussions
Bibliographic record
Abstract
Cryptocurrency development has continuous escalation in the past years and holds its presence significantly in open source development.Online collaborative software development platforms such as GitHub offer us an opportunity to observe developer effort, activity and software growth.Cryptocurrency has enabled various applications such as smart contracts, electronically decentralized payments, etc. Since, prices of each cryptocurrency are driven by many factors, we are interested in investigating how various characteristics of cryptocurrency's codebase development affect market capitalization price.Thus, we conduct a study on a panel dataset containing nearly a year of daily observations of development activity, popularity, and market capitalization for over two hundred open source cryptocurrencies.Stack Overflow (SO) remains the most popular Q&A forum for software developers, providing solutions to software related problems.SO is a rich source of user specific information with special emphasis on human emotion.In this study, we mine SO data to explore the hot topics related to cryptocurrency development and study the sentiment of cryptocurrency discussions.Our results demonstrate that 1) the popular cryptocurrencies based on popularity in GitHub are entirely present in the popular cryptocurrency list based on price on CoinMarketCap; 2) Ethereum is at the leading position in terms of popularity, while Bitcoin dominates in price; 3) using Granger causality analysis, we find no convincing evidence of the "predictive" relationship between software development related metrics such as stars, watchers, forks, contributors, commits, and lines of code changes and market price except for Ethereum; 4) developers express a positive sentiment in SO discussions related to popular cryptocurrencies; 5) the extracted keywords from SO discussions related to cryptocurrency prevail the hot topics the developers discuss about.I, Raisul Islam Rashu, would like to express my deep and sincere gratitude to my supervisor, Professor Olga Baysal, for her continuous guidance, patience, insight, helpful discussions throughout my thesis.Her constant encouragement, instructive feedback and support made this work successful.Thanks to the examination committee for their helpful comments.I am very grateful to my family -my parents and my younger brother Saiful Islam Rimon.They gave me immense support by motivating, energizing and keeping
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".