EEG Integrated Platform Lossless (EEG-IP-L) pre-processing pipeline for objective signal quality assessment incorporating data annotation and blind source separation
Bibliographic record
Abstract
BACKGROUND: The methods available for pre-processing EEG data are rapidly evolving as researchers gain access to vast computational resources; however, the field currently lacks a set of standardized approaches for data characterization, efficient interactive quality control review procedures, and large-scale automated processing that is compatible with High Performance Computing (HPC) resources. NEW METHOD: In this paper we describe an infrastructure for the development of standardized procedures for semi and fully automated pre-processing of EEG data. Our pipeline incorporates several methods to isolate cortical signal from noise, maintain maximal information from raw recordings and provide comprehensive quality control and data visualization. In addition, batch processing procedures are integrated to scale up analyses for processing hundreds or thousands of data sets using HPC clusters. RESULTS: We demonstrate here that by using the EEG Integrated Platform Lossless (EEG-IP-L) pipeline's signal quality annotations, significant increase in data retention is achieved when applying subsequent post-processing ERP segment rejection procedures. Further, we demonstrate that the increase in data retention does not attenuate the ERP signal. CONCLUSIONS: The EEG-IP-L state provides the infrastructure for an integrated platform that includes long-term data storage, minimal data manipulation and maximal signal retention, and flexibility in post processing strategies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.021 | 0.008 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".