MétaCan
Menu
Retour à la cohorte
Enregistrement W2767494814 · doi:10.1111/rssa.12331

Preface to the Papers on ‘Small Area Estimation’

2017· article· en· W2767494814 sur OpenAlexaboutno aff
Nikos Tzavidis, Li‐Chun Zhang, Danny Pfeffermann, Partha Lahiri

Notice bibliographique

RevueJournal of the Royal Statistical Society Series A (Statistics in Society) · 2017
Typearticle
Langueen
DomaineDecision Sciences
Thématiquedemographic modeling and climate adaptation
Établissements canadiensnon disponible
Organismes subventionnairesEconomic and Social Research Council
Mots-clésEstimationComputer scienceMathematicsStatisticsEconomics

Résumé

récupéré en direct d'OpenAlex

Small area estimation (SAE) has been, and still predominantly is, a very fertile area in official and survey statistics research with important theoretical and applied contributions. In recent decades, an increasing number of national statistical institutes and other organizations across the world have recognized the importance of producing small area statistics and their potential use for informing policy decisions. Cutting edge developments in model-based small area methods are used in practice for the production of national statistics. Among many, examples of organizations with research interests in SAE include the US Census Bureau, the UK Office for National Statistics, the World Bank, the Statistical Office of Italy, the Central Bureau of Statistics in Holland, Statistics Canada, the Australian Bureau of Statistics, the Brazilian Statistical Office, the National Council for the Evaluation of Social Development Policy (Consejo Nacional de Evaluación de la Politica de Desarrollo Social) in Mexico and the Ministry for Social Development (Ministerio de Desarrollo Social) in Chile. Over time, users’ needs have surpassed the limits of what can be achieved with traditional SAE methods. For example, in addition to simple linear parameters like averages and proportions, users request the estimation of more complex indicators such as geographically disaggregated measures of deprivation and inequality. In addition, the availability of what is known as ‘big data’, e.g. satellite and mobile phone data, and probabilistically linked administrative and survey data, has created new methodological challenges. Meeting the increasing complexity of users’ needs requires new specialized methodology and software that extend beyond conventional survey operations. This has created new research opportunities and the need for closer collaboration between researchers and practitioners for transferring research into practice and, hence, maximizing the effect of research. The present themed papers in this issue include a selection on SAE. The call for papers was published at the first Latin American International Statistical Institute satellite meeting on SAE that took place at the Pontificia Universidad Católica de Chile in Santiago, Chile, in August 2015. This conference was part of a series of scientific meetings devoted solely to SAE. Starting with the 2001 meeting in Maryland, USA, SAE conferences have taken place in Jyväskylä, Finland (2005), Pisa, Italy (2007), Elche, Spain (2009), Rhine (river cruise) in Germany (2009), Trier, Germany (2011), Bangkok, Thailand (2013), Poznan, Poland (2014), Santiago, Chile (2015), Maastricht, Netherlands (2016), and Paris, France (2017). The next conference will take place in Shanghai, China, June 16th–18th, 2018. The call for papers for this issue attracted many high quality submissions. All manuscripts went through the full peer review process of the journal after which 11 papers were selected for publication. The manuscripts selected put forward new frequentist and Bayesian methodologies on a range of topics including robust and non-parametric methods, poverty mapping, measurement error models and time series models with new tools for model testing and selection also being proposed. The papers include applications in ecology, economics, health and medicine, transportation and education. A major criticism of the use of model-based methods in survey estimation is their reliance on assumptions that are difficult to check and satisfy in practice. However, in many applications, the use of models is deemed necessary for improving the precision of estimates. In recent years part of the small area literature has focused on developing small area methods that are robust to departures from the model assumptions (e.g. Ghosh et al. (2008) and Chambers et al. (2014)). Papers that appear in this issue continue this tradition by considering extensions of popular small area models, which assume Gaussian distributions for the model error terms, e.g. the Fay–Herriot model (Fay and Herriot, 1979). The extensions allow for alternative distributions, which are potentially more suitable for particular types of survey data like business survey data. Moura, Neves and Silva present small area models with skew normal and skew t-distributions for skewed survey data and apply the models to business surveys. Ferrante and Pacei extend a normal, multivariate Fay–Herriot area level model by allowing the random effects and the sampling errors to have skew normal distributions. The model is also applied to business survey data for estimating value added and labour costs. A different form of robustness against failure of the model assumptions is obtained by the use of non-parametric models. Wagner, Münnich, Hill, Stoffels and Udelhoven study the use of non-parametric small area models that use shape-constrained penalized B-splines and apply these models to estimate timber reserves in the German region of Rhineland-Palatinate. Two of the papers that appear in this issue propose new methodologies for estimating non-linear parameters, in particular, income deprivation (poverty) and inequality indicators. This topic generated considerable debate in the small area literature. Marhuenda, Molina, Morales and Rao extend the methodology that was proposed in Molina and Rao (2010) by proposing an empirical best predictor under the twofold nested error regression model and apply the methodology to estimating gender-specific poverty rates in counties of the Spanish region of Valencia. The methodology that is proposed in that paper makes the application of the empirical best methodology more realistic in situations where survey data are collected via multistage clustered designs. Das and Chambers, in contrast, focus on an alternative poverty mapping methodology that has been extensively used by the World Bank (Elbers et al., 2003) and propose robust mean-squared error estimators for poverty estimates produced by the use of this method. The paper by Schmid, Bruckschen, Salvati and Zbiranski presents one of the first attempts to use big data sources—mobile data in their application—as covariate information in area level models. They apply the model to derive gender-specific, subnational benchmarked literacy rates in Senegal. We expect that, in the near future, incorporating such sources of data as covariates when producing small area estimates will become common practice. Despite the challenges and the open research questions, it is encouraging to see a methodology that enables us to do this. Another area of research in SAE with renewed interest is whether or how to account for measurement errors in model covariates. This topic is of particular interest since the ease of access to regularly updated covariate information from large surveys is an important advantage compared with access to census and administrative microdata, which is difficult because of confidentiality constraints. Moreover, covariates derived from big data (e.g. mobile phone data), as in the paper by Schmid and his colleagues, are also likely to be affected by measurement error although quantification of the measurement error in this case is challenging. SAE methods must account for the measurement error in covariates obtained from survey data. In the present issue, Arima, Bell, Datta, Franco and Liseo consider a multivariate Fay–Herriot model where the covariates are assumed to be subject to observation errors. They develop Markov chain Monte Carlo methodology which is applied to estimate US county level poverty rates for school-aged children. Similar in spirit but using a unit level model is the paper by Maples, which presents a methodology that allows for the use of covariate data coming from a large independent survey. The methodology is applied to derive small area estimates of disability. A topic that until recently received relative little attention is model selection and testing (e.g. Datta et al. (2011) and Pfeffermann (2013)). Lombardía, López-Vizcaíno and Rueda propose a mixed generalized Akaike information criterion for model selection. The method is compared with alternative methods, including the conditional Akaike information criterion, using simulations and real labour market and health data. On a related topic, Torkashvand, Jafari Jozani and Torabi propose clustering of small areas based on the Euclidean distance between covariates and propose a statistical test to investigate the homogeneity of between-clusters variance components. Finally, Bollineni-Balabay, van den Brakel, Palm and Boonstra present a paper on another topic that has attracted interest in the SAE literature, namely borrowing strength over time. The paper compares state space models (estimated by use of the Kalman filter combined with a frequentist approach to hyperparameter estimation) with multilevel time series models fitted under the hierarchical Bayesian framework. The application of the methods is to data collected in the Dutch Travel Survey, which has small sample sizes and discontinuities caused by survey redesigns. We thank all the authors for their willingness to submit papers for this issue and for responding to requests for revisions in a timely manner. We are particularly grateful to the Joint Editor for overseeing our work, as Guest Associate Editors, with admirable patience. We also thank the referees. Without their constructive comments this issue would not be possible.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,007
score de la tête « metaresearch » (Gemma)0,072
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: Sans objet
GenreSignal candidat: Éditorial · Signal consensuel: Éditorial
Score de désaccord entre enseignants0,109
Score d'incertitude au seuil0,364

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0070,072
Méta-épidémiologie (sens strict)0,0020,001
Méta-épidémiologie (sens large)0,0010,002
Bibliométrie0,0050,008
Études des sciences et des technologies0,0020,002
Communication savante0,0050,005
Science ouverte0,0020,002
Intégrité de la recherche0,0030,007
Charge utile insuffisante (le modèle a refusé de juger)0,1090,086

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,083
Tête enseignante GPT0,357
Écart entre enseignants0,274 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeSans objet
Domainenon disponible
GenreÉditorial

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2017
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueJournal of the Royal Statistical Society Series A (Statistics in Society)Même sujetdemographic modeling and climate adaptationTravaux en français237 207