Analysis of COVID-19 blockchain data using social signals for contact tracing
Bibliographic record
Abstract
Communities around the world are adapting to the fast-changing global situation due to the unprecedented effect of COVID-19. The data related to COVID-19 spread is being analyzed for identifying outbreaks and for trying to predict their future movement across geographies with the help of advanced machine learning models. The main challenge faced by researchers was using the centralized data sources, aggregating relevant data, and standardizing them, at a global level. In my thesis, I have first studied the difference between the centralized and blockchain-backed data provided by Mipasa (powered by the IBM Blockchain Platform and the IBM Cloud) and have developed a knowledge graph for COVID-19 cases in USA and Japan. The creation of a knowledge graph helped in predicting the regions which could witness the formation of new clusters. This observation helped in isolating regions and prevented the further spread of the virus. This led to the development of Decentralized Applications. The first decentralized application for COVID-19 symptoms tracking using Blockchain is developed to enhance reliable data collection for training Machine Learning (ML) models. The Blockchain integration in this application helped patients to provide COVID-19 symptoms data with trust. In addition to this, the data was first verified by an entity of the decentralized network (e.g. a COVID-19 testing lab). Then, with the consent of the patient, this data was provided to the centralized system for retraining the ML model. This re-training performed with verified data, updates the ML model and provides accurate results. The data collected from different platforms helped in identifying outbreaks, contact tracing, and the creation of a machine learning model for predicting the future movements of the outbreak to minimize the spread. However, there was delay in taking measures which cost many lives, and many local businesses were shut down around the world. As a solution to this problem, I have developed a Decentralized architecture where the Blockchain Oracle smart contract could access the data outside the Blockchain, the second application. The REST API provided the daily aggregated data from three platforms (Twitter, Google Mobility Data, and COVID-19 cases) for Toronto, Canada. This aggregated data helped participants of the Blockchain ecosystem to initiate the smart contracts early decisions like implementing lockdown, deploying officials for contact tracing, and providing financial support for businesses and individuals affected due to lockdown.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".