vanDisk: An Exploration in Peer-To-Peer Collaborative Back-Up Storage
Bibliographic record
Abstract
As personal computers become an integral part of our daily lives, huge volumes of data need to be reliably managed and archived. Uncorrelated failures within a set of independent personal computers offer the promise of low-cost, reliable data storage. The vanDisk project attempts to realize this promise. The main assumption of our project is that users are willing to donate raw storage space to their peers to increase the reliability of their own data. In our system, users offer a portion of their disks to be used as backup space for other users in exchange for space to store backup copies of their own data, thus decreasing the possibility of catastrophic data loss. A number of characteristics differentiate vanDisk from existing projects that explore this space. First, unlike existing projects that that increase redundancy at the data-block or file level, vanDisk operates at the disk level. This substantially simplifies data management and reduces management overhead at the cost of marginally higher recovery costs from partial failure. Second, all data-related operations are transparently replicated at the data source. Third, our design includes an orthogonal component to manage space and bandwidth. Our system is integrated with Microsoft Windows and offers users a virtual drive that transparently replicates data across multiple machines. As well, a complete, original copy of the data is always available on the user's own system. We have modified TrueCrypt, an open source virtual disk package that offers data confidentiality through encryption, and we have added a new driver layer that redirects and replicates all IO requests to a set of network block device servers offered by the peers to store replicated data. Additionally, we use simple data encoding to offer user-tunable tradeoffs between space overheads, compute overheads, and data reliability.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.011 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.003 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.005 | 0.015 |
| Open science | 0.009 | 0.008 |
| Research integrity | 0.003 | 0.003 |
| Insufficient payload (model declined to judge) | 0.005 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".