Notice bibliographique
Résumé
This work proposes a general structure for inference in modelling discrete data sets using a count time series model with INGARCH models with a log-linear structure and mixed Poisson innovations. To this end, a class of probability distributions will be used, the main objective of which is to model count data over time that present a condition of overdispersion. More specifically, the work presents two particular cases: the Inverse Gaussian Poisson log-linear distribution and the Negative Binomial log-linear distribution, which are obtained by considering cases of unobservable data that follow the Inverse Gaussian and Gamma distributions, respectively. The distributions inserted through the mean have as a common point the fact that they are members of the exponential family of distributions. The iterative maximum likelihood method will be used to estimate the model parameters using the EM algorithm. The performance of the estimators will be evaluated through simulation studies using the Monte Carlo method, considering different sample sizes to evaluate the asymptotic behaviour of these estimators.In the section on applying the proposed model to real data sets, three databases were considered for analysis: the first lists the number of hospitalisations due to alcohol abuse in the state of Paraíba, the second evaluates the same problem, but with the data presented for the state of Piauí and, finally, the database consisting of the number of cases of Campylobacter infections in the province of Quebec in Canada was evaluated, thus closing the section on applications to real data. The simulation data was tested using the two proposed extensions and the comparison model called log-linear Poisson proposed by [9], initially taking into account a graphical analysis of the behaviour of the sample, autocorrelation and partial autocorrelation, the study of simulation by means of convergence taking into account the values obtained and the observation of graphs representing a generalised view of the layout of the simulation data. Subsequently, a reflection was made on its effectiveness through information criteria and the mean square error used in the process of evaluating and choosing the best regression model to adjust the data.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,011 | 0,024 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,002 | 0,003 |
| Bibliométrie | 0,003 | 0,003 |
| Études des sciences et des technologies | 0,001 | 0,002 |
| Communication savante | 0,004 | 0,004 |
| Science ouverte | 0,006 | 0,003 |
| Intégrité de la recherche | 0,003 | 0,003 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,007 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».