Vehicular infotainment forensics: collecting data and putting it into perspective
Bibliographic record
Abstract
In today???s transportation system, countless numbers of vehicles are on the road and later generations have become mobile computers. Vehicles now have embedded infotainment systems that enable user-friendliness and practicability with functions such as a built-in global positioning system, media playback device and application interface. Smartphones and laptops can connect to them through Bluetooth and WiFi for all sorts of utilities. This enables data flow between a user???s device and the infotainment system and because of this interaction, data remnants are kept on these embedded devices. It is important to determine what type of data is stored long term since this information reflects a user???s activity and potential personal information. In terms of forensics, this data could be used to solve criminal activities if a vehicle was suspected of being an accessory to a crime; raising general awareness about this topic is important due to the potential sensitive information circulated. This main objective of this thesis is to demonstrate what types of information are stored on infotainment systems, how it can be acquired and the implications and contributions of the collected data in relation to the overall field of digital forensics.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.015 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.006 | 0.004 |
| Science and technology studies | 0.002 | 0.003 |
| Scholarly communication | 0.006 | 0.006 |
| Open science | 0.001 | 0.004 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".