Classification of simultaneous multiple partial discharge sources based on probabilistic interpretation using a two-step logistic regression algorithm
Bibliographic record
Abstract
In online condition assessment monitoring of high voltage (HV) insulators, it is often required to identify multiple, simultaneously activated partial discharge (PD) sources that happen in the insulation of the HV apparatus. Phased resolved partial discharge (PRPD) patterns are commonly used to identify PD sources. However, multiple, concurrent PD sources sometimes result in partially overlapped patterns, which make them hard to be identified. In this paper, we develop an accurate, reliable algorithm by constructing a novel two-step logistic regression (LR) model to conduct probabilistic identification of multi-source PDs. To this end, principal component analysis is applied on a database to construct a low dimensional space associated with single-source PDs. Samples of multi-source PDs are then projected onto this space and one-class kernel support vector machine is adopted to distinguish multi-source PDs from single-source ones. Finally, classification is performed by estimating the probability (degree of membership) of each PRPD pattern arising from different multi-source PDs following two rounds of LR modeling. To evaluate the performance of our proposed method, we study a number of multi-source PD models to simulate common defects of Gas-Insulated Switchgear (GIS) in small-scale laboratory test cells with realistic SF6gas condition. Observations are obtained using fingerprints generated by a novel approach from recorded PRPD patterns. Comprehensive performance evaluation of the proposed algorithm and its advantages are conducted and the development of analytical equations is presented. The results of this paper can be used to design a solid basis for an automated multi-source classification system, which facilitates multi-source PD identification in early stages and safe operation of HV apparatus.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.006 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".