Towards automated statistical partial discharge source classification using pattern recognition techniques
Bibliographic record
Abstract
This study presents a comprehensive review of the automated classification in partial discharge (PD) source identification and probabilistic interpretation of the classification results based on the relationship between the variation of the phase‐resolved PD (PRPD) patterns and the source of the PD. The proposed automated classification system consists of modern, high‐performance statistical feature extraction methods and classifier algorithms. Their application in online monitoring and recognition of the PD patterns is investigated based on their low‐processing time and high‐performance evaluation. The application of modern statistical algorithms and pre‐processing methods configured in this automated classification system improves the pattern recognition accuracy of the different PD sources that are suitable to be employed in different high‐voltage (HV) insulation media. To evaluate the performance of the different combinations of the feature extraction/classier pairs, laboratory setups are designed and built that simulate various types of PDs. The test cells include three sources of PD in , two sources of PD in transformer oil, and corona in the air. Data samples for different classes of PD sources are captured under two levels of voltage and two different levels of noise. The results of this study evaluate the suitability of the proposed classification systems for probabilistic source identification in various insulation media. Furthermore, of importance to the problem of the PD source identification is to assign a ‘degree of membership’ to each PRPD pattern, besides assigning a class label to it. Some of the classifier algorithms studied in this study, such as fuzzy classifiers, are not only able to show high classification accuracy rate, but they also calculate the ‘degree of membership’ of a sample to a class of data. This enables probabilistic interpretation of a new PRPD pattern that is being classified. The determination of the degree of membership for future PRPD samples allows safer decision making based on the risk associated with the different sources of PD in HV apparatus.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.004 | 0.002 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".