MétaCan
Menu
Back to cohort

Towards a Universal Vibration Analysis Dataset

2025· article· en· W4410923050 on OpenAlexaff
Mert Sehri, Igor Varejão, Zehui Hua, Vitor Berger Bonella, Adriano Lages dos Santos, Francisco de Assis Boldt, Patrick Dumond, Flávio Miguel Varejão

Bibliographic record

VenueInternational Journal of Prognostics and Health Management · 2025
Typearticle
Languageen
FieldEngineering
TopicStructural Health Monitoring Techniques
Canadian institutionsUniversity of Ottawa
Fundersnot available
KeywordsVibrationComputer scienceData sciencePhysicsAcoustics

Abstract

fetched live from OpenAlex

In the realm of machine learning (ML), particularly in visual computing, ImageNet has established itself as an indispensable resource for transfer learning (TL), enabling the development of highly effective models with reduced training time and data requirements. However, the domain of vibration analysis, which is critical in fields such as predictive maintenance, structural health monitoring, and fault diagnosis, lacks a comparable large-scale, annotated dataset to facilitate similar advancements. To address this gap, we propose a dataset framework that begins with a focus on bearing vibration data as an initial step towards creating a universal dataset for vibration-based spectrogram analysis for all machinery. The initial phase should feature a curated collection of bearing vibration signals, designed to represent a wide array of real-world scenarios, including vibration data of various public bearing datasets. To demonstrate the initial efficacy of this approach, experiments should be conducted using a state-of-the-art deep learning (DL) architecture, showing improvements in model performance when pre-trained on bearing vibration data and fine-tuned on smaller, domain-specific datasets. These findings will illustrate the potential to parallel the success of ImageNet in visual computing, but for vibration analysis. In future iterations, this proposal will evolve to encompass a broader range of vibration signals from multiple types of machinery and sensors, with an emphasis on generating spectrogram-based representations of the data. Multi-sensor data, including signals from accelerometers, microphones, and other devices should be used, ensuring versatility for both domain-specific and generalized applications. They will be incorporated to create a more holistic and comprehensive dataset, enabling the application of advanced sensor fusion techniques in vibration analysis. Each sample will be labeled with detailed metadata, such as machinery type, operational status, and the presence or type of faults, ensuring its utility for supervised and unsupervised learning tasks. This extension will position this work as a universal resource for various industries, enhancing the ability of researchers and practitioners to apply TL to diverse vibration analysis problems. In addition to the dataset, a comprehensive framework for data preprocessing, feature extraction, and model training specific to vibration data should be developed. This framework will standardize methodologies across the research community, fostering collaboration and accelerating progress in predictive maintenance, structural health monitoring, and related fields. In conclusion, this proposal represents a transformative step in vibration analysis, starting with bearing data as its foundation and ultimately evolving into a universal dataset for spectrograms and multi-sensor data for all machinery. By mirroring the success of ImageNet in visual computing, it has the potential to significantly improve the development of intelligent systems in industrial applications, enabling more efficient and reliable operations.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.002
metaresearch head score (Gemma)0.006
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Dataset · Consensus signal: none
Teacher disagreement score0.006
Threshold uncertainty score0.021

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0020.006
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.002
Bibliometrics0.0030.002
Science and technology studies0.0010.001
Scholarly communication0.0010.002
Open science0.0040.004
Research integrity0.0020.003
Insufficient payload (model declined to judge)0.0060.006

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.022
GPT teacher head0.351
Teacher spread0.329 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreDataset

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations3
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueInternational Journal of Prognostics and Health ManagementSame topicStructural Health Monitoring TechniquesFrench-language works237,207