Bibliographic record
Abstract
As FPGAs become larger and more complex, productive debugging is becoming more challenging. In this work, we detail a new debugging flow based on hardware checkpointing that provides full visibility and controllability while maintaining reasonable execution speed. Hardware checkpointing is useful not only for debugging but also enables several other capabilities such as live migration, fault recovery, and context switching; however, it has been difficult to achieve for FPGA applications. In this thesis, we overcome the challenges of checkpointing FPGA designs and realize the proposed checkpoint-based debugging flow. First, we propose techniques and wrappers that can safely interrupt a running design to create a consistent (restartable) checkpoint while avoiding hazards such as data loss or deadlock. We also develop approaches and tools that can access buried on-chip state that cannot be directly captured to create complete checkpoints. We next propose a checkpoint-based debugging framework, StateMover, that can seamlessly move the design state back and forth between an FPGA and a simulator, achieving the best of both worlds: speed and observability. Finally, we build a transaction-based co-simulation framework, StateLink, to extend the functionality of the proposed debugging flow to systems that cannot be entirely moved to a simulator such as CPU+FPGA accelerators or datacenter-scale applications. The combination of these tools enables new and productive debugging flows. StateMover allows designs to run at full hardware speed until a region of interest is approached. Then, a checkpoint can be loaded and its execution observed and controlled in a simulator. StateLink allows a designer-selected portion of the system to be moved into a simulation, enabling simulation speedups of up to 25x versus simulating the entire design. We demonstrate on several designs that StateMover and StateLink support can be added to a design with low resource and timing overhead, and illustrate the utility of the flow by debugging a complete Memcached system.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".