Statistical Inference for Structured Spatial and Temporal Point Data
Notice bibliographique
Résumé
The availability of large-scale spatial and temporal data has fueled increasing interest in statistical modelling and analysis. With the recent development of data collection and data storage techniques, the observation scopes can sometimes involve a extremely vast range or an explosive amount of cases. Then this always leads to an inevitable focus that there tend to be some heterogeneous properties among observations. Thus, the research was conducted to explain the variability in spatial or temporal data considering the correlation of observations.\n\nWe first considered the intensity estimation problem for large spatial point patterns on complex domains in R2 (e.g., domains with irregular boundaries, sharp concavities, and/or interior holes due to geographic constraints) and linear networks, where many existing spatial point process models suffer from the problems of “leakage" and computation. We proposed an efficient intensity estimation algorithm to estimate the spatially varying intensity function and to study the varying relationship between intensity and explanatory variables on complex domains. The method is built upon a graph regularization technique and hence can be flexibly applied to point patterns on complex domains such as regions with irregular boundaries and holes, or linear networks. An efficient proximal gradient optimization algorithm is proposed to handle large spatial point patterns. Numerical studies were conducted to illustrate the performance of the method. Besides, we apply the method to study and visualize the intensity patterns of the accidents on the Western Australia road network, and the spatial variations in the effects of income, lights condition, and population density on the Toronto homicides occurrences.\n\nIn addition, the spatial inhomogeneity occurred in various scenarios, especially for the data laying in a vast-scale space. we further established a spatially adaptive sampling design approach based in an estimation of the spatially varying underlying contamination distribution. This part of research was motivated by an Arsenic exposure data which were collected through drinking water in private wells across the Iowa state. From the public and environmental health management perspective, it is critical to allocate the limited resources to establish an effective arsenic sampling and testing plan for health risk mitigation. we propose a statistical regularization method to automatically detect spatial clusters of the underlying contamination risk from the currently available private well arsenic testing data in the USA, Iowa. This approach allows us to develop a sampling design method that is adaptive to the changes in the contamination risk across the identified clusters.\n\nFinally, we further looked into the cluster issues in structured temporal point data. How to cluster event sequences from heterogeneous point processes is a challenging task, especially when event sequences are repeatedly observed and associated with multiple event types. To solve this problem, we proposed an efficient model-based clustering framework, based on a novel multivariate mixture of functional point processes (MFPP). The proposed model generated event sequences from a multi-level log-Gaussian Cox process, which allows to uncover complex inner patterns among sequences, by imposing multiple latent random effects. We prove the identifiability of our mixture model and developed an effective semi-parametric Exponential-Solution (ES) algorithm to the proposed model. The effectiveness of the proposed framework is demonstrated through simulation studies and real data analyses.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,000 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,001 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,001 | 0,001 |
| Intégrité de la recherche | 0,000 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,006 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».