An Integrated Approach to Forecast Bitcoin Price Incorporating Economic Factors, Strategic Commodities, and X Tweets
Bibliographic record
Abstract
Bitcoin is known for its volatility. So, we conducted a diverse market analysis to construct a conclusive market feature dataset and a hybrid model architecture to find better insight and make better forecasts. One key factor that other research overlooked was the market impact of AMD and NVIDIA. These two most prominent tech companies provide consumerlevel GPUs mostly used for cryptocurrency mining. Focusing on these gaps, we conducted an extensive study of the potential markets and their impact on our model training. We provided an extensive data analysis of the impact of crude oil, gold, AMD, NVIDIA, S&P500, NASDAQ, and sentiment analysis of 85 million tweets altogether on Bitcoin price. The volume of tweets that we collected through web scraping is substantial to the other studies. From our comparative analysis, we deduced that A bidirectional LSTM works best for predicting the Bitcoin trend compared to other deep learning and time series models. Thus, we chose Bidirectional LSTM as our base model and used it to build a hybrid architecture ensemble model with Random Forest Regressor. For the proposed model, we have deployed four simultaneous Bayesian Optimized Bidirectional LSTM models, each with its distinct input features, and trained the Random Forest Regressor model using the predictions from those four models. The trained Random Forest model was used to select the best forecast we obtained from the four Bi-LSTM models. Our findings indicate that Bidirectional LSTM predicts more accurately when sentiment analysis and other macroeconomic aspects are incorporated.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.012 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".