Improving SQL query performance on embedded devices using pre-compilation
Bibliographic record
Abstract
Embedded devices are increasingly being used for data collection and on-device data analysis for applications in environmental and infrastructure monitoring, health and wearable computing, and sensor and mobile systems. Processing data on the device rather than transmitting it over a network for analysis reduces energy consumption, network bandwidth usage, and results in more robust and longer functioning devices. A key challenge is enabling efficient data management on devices that may only have a few KBs of memory and limited code space. Previous work has demonstrated that using relational databases is possible on embedded devices with restrictions on the queries that can be processed. In this work, we eliminate one of the key barriers to using relational technology on embedded devices, which is the massive overhead involved in SQL parsing and translation that can take up to 50% of the code and memory resources on the device. Our approach allows developers to continue to use relational APIs and SQL during development which are then pre-compiled when deployed on the device. This produces all of the benefits of relational systems without the on-device overhead and limitations. Experimental results demonstrate that query pre-compiling can reduce query parse times by up to 90% and on-device execution times by up to 50%. The technique is applicable to a wide range of embedded systems and databases.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".