An Explicit Improvement on Generative Adversarial Network-Based Time Series Generation: Applying Synthetic Data to N2O Emission Prediction in Farming
Notice bibliographique
Résumé
<p>Traditionally, time series data augmentation has primarily focused on improving the architecture of Generative Adversarial Network (GAN), with the aim of closely matching the original data distribution while also preserving the dynamic behavior of the original data. However, even state-of-the-art GAN models like TimeGAN fall short in preserving the temporal dynamics present in the original time series due to the absence of first-order difference information. To address this limitation, this study proposes a novel process for generating multivariate time series data. The proposed process comprises four essential modules: a) the GAN module for generating multivariate time series data, b) the sampling module for preserving the first-order difference distribution, c) the smoothing module for refining the generated data, and d) an evaluation module using the Kolmogorov-Smirnov Test (KS-test) and Hilbert-Schmidt Independence Criterion (HSIC), along with other metrics to test the synthetic time series data. This comprehensive approach ensures that the synthetic time series data maintains both the distribution and the dynamic behavior of the original data.</p> <p>We extensively discuss the role of the β factor in the modified Metropolis-Hastings algorithm (in the sampling module), which controls the level of information preservation from the original time series. Our experiments reveal that with small β values, periodic information can be retained effectively. The joint distribution of the first-order difference of the synthetic time series data remains consistent when the same β value is applied in the modified Metropolis-Hastings algorithm. However, we observe that β has no impact on the partial autocorrelation functions. Nevertheless, the generated data from the sampling module maintains the memoryless property of the Markov Chain. Therefore, in the smoothing module, we apply the exponential moving average (EMA) method to simulate the long-term relationships within the original time series, and find that an optimal α value is approximately 0.4 or 0.5. Lastly, we employ the synthetic time series data to train a neural network model developed in another work. Our findings indicate that the neural network model trained on synthetic time series data exhibits performance comparable to that of a model trained on the original data.</p>
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,001 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,001 | 0,001 |
| Science ouverte | 0,002 | 0,003 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».