An Integrated Approach to Forecast Bitcoin Price Incorporating Economic Factors, Strategic Commodities, and X Tweets
Bibliographic record
Abstract
Bitcoin is known for its volatility. So, we conducted a diverse market analysis to construct a conclusive market feature dataset and a hybrid model architecture to find better insight and make better forecasts. One key factor that other research overlooked was the market impact of AMD and NVIDIA. These two most prominent tech companies provide consumerlevel GPUs mostly used for cryptocurrency mining. Focusing on these gaps, we conducted an extensive study of the potential markets and their impact on our model training. We provided an extensive data analysis of the impact of crude oil, gold, AMD, NVIDIA, S&P500, NASDAQ, and sentiment analysis of 85 million tweets altogether on Bitcoin price. The volume of tweets that we collected through web scraping is substantial to the other studies. From our comparative analysis, we deduced that A bidirectional LSTM works best for predicting the Bitcoin trend compared to other deep learning and time series models. Thus, we chose Bidirectional LSTM as our base model and used it to build a hybrid architecture ensemble model with Random Forest Regressor. For the proposed model, we have deployed four simultaneous Bayesian Optimized Bidirectional LSTM models, each with its distinct input features, and trained the Random Forest Regressor model using the predictions from those four models. The trained Random Forest model was used to select the best forecast we obtained from the four Bi-LSTM models. Our findings indicate that Bidirectional LSTM predicts more accurately when sentiment analysis and other macroeconomic aspects are incorporated.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".