A Blind Modeling Tool for Standardized Evaluation of Battery State of Charge Estimation Algorithms
Bibliographic record
Abstract
There are hundreds of approaches to estimating battery state of charge (SOC). It is difficult to compare results reported in different papers because each typically uses a different dataset. While some papers compare multiple SOC estimation algorithms, the author's bias, skill, or effort towards each algorithm may unintentionally skew the results. A standardized way to test and compare methodologies between authors is necessary to allow the best algorithms to stand out. An example in another application area is the National Institute of Standards (NIST) Face Recognition Vendor Test, which compares facial recognition software using a standardized dataset. A similar approach is proposed here for batteries, where data is provided for users to parameterize and train their algorithms. An online tool is provided to subject the algorithms to a wide range of blinded test cases. A high-quality dataset is prepared using battery cells from a prevalent electric vehicle. A total of sixty-four drive cycles are performed at each of six temperatures ranging from -20 °C to 40 °C. The blind modelling tool is demonstrated for one SOC estimation algorithm. It will be made available for researchers to benchmark and compare their algorithms.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".