MétaCan
Menu
Retour à la cohorte
Enregistrement W4283313359 · doi:10.2196/35797

Strategies and Lessons Learned During Cleaning of Data From Research Panel Participants: Cross-sectional Web-Based Health Behavior Survey Study

2022· article· en· W4283313359 sur OpenAlexvenueno aff
Mariana Arévalo, Naomi C. Brownstein, Junmin Whiting, Cathy D. Meade, Clement K. Gwede, Susan T. Vadaparampil, Kristin J Tillery, Jessica Y. Islam, Anna R. Giuliano, Shannon M. Christy

Notice bibliographique

RevueJMIR Formative Research · 2022
Typearticle
Langueen
DomaineSocial Sciences
ThématiqueSurvey Methodology and Nonresponse
Établissements canadiensnon disponible
Organismes subventionnairesNational Center for Advancing Translational SciencesUniversity of South CarolinaSouth Carolina Clinical and Translational Research Institute, Medical University of South CarolinaNational Cancer InstituteNational Institutes of HealthMoffitt Cancer CenterMedical University of South Carolina
Mots-clésChecklistData qualityThe InternetSurvey data collectionWeb applicationQuality (philosophy)PsychologyData scienceComputer scienceMedical educationWorld Wide WebMedicineBusinessMarketing

Résumé

récupéré en direct d'OpenAlex

BACKGROUND: The use of web-based methods to collect population-based health behavior data has burgeoned over the past two decades. Researchers have used web-based platforms and research panels to study a myriad of topics. Data cleaning prior to statistical analysis of web-based survey data is an important step for data integrity. However, the data cleaning processes used by research teams are often not reported. OBJECTIVE: The objectives of this manuscript are to describe the use of a systematic approach to clean the data collected via a web-based platform from panelists and to share lessons learned with other research teams to promote high-quality data cleaning process improvements. METHODS: Data for this web-based survey study were collected from a research panel that is available for scientific and marketing research. Participants (N=4000) were panelists recruited either directly or through verified partners of the research panel, were aged 18 to 45 years, were living in the United States, had proficiency in the English language, and had access to the internet. Eligible participants completed a health behavior survey via Qualtrics. Informed by recommendations from the literature, our interdisciplinary research team developed and implemented a systematic and sequential plan to inform data cleaning processes. This included the following: (1) reviewing survey completion speed, (2) identifying consecutive responses, (3) identifying cases with contradictory responses, and (4) assessing the quality of open-ended responses. Implementation of these strategies is described in detail, and the Checklist for E-Survey Data Integrity is offered as a tool for other investigators. RESULTS: Data cleaning procedures resulted in the removal of 1278 out of 4000 (31.95%) response records, which failed one or more data quality checks. First, approximately one-sixth of records (n=648, 16.20%) were removed because respondents completed the survey unrealistically quickly (ie, <10 minutes). Next, 7.30% (n=292) of records were removed because they contained evidence of consecutive responses. A total of 4.68% (n=187) of records were subsequently removed due to instances of conflicting responses. Finally, a total of 3.78% (n=151) of records were removed due to poor-quality open-ended responses. Thus, after these data cleaning steps, the final sample contained 2722 responses, representing 68.05% of the original sample. CONCLUSIONS: Examining data integrity and promoting transparency of data cleaning reporting is imperative for web-based survey research. Ensuring a high quality of data both prior to and following data collection is important. Our systematic approach helped eliminate records flagged as being of questionable quality. Data cleaning and management procedures should be reported more frequently, and systematic approaches should be adopted as standards of good practice in this type of research.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,462
score de la tête « metaresearch » (Gemma)0,483
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesMétarecherche
Catégories consensuellesMétarecherche
DomaineSignal candidat: Méthodes · Signal consensuel: Méthodes
Devis d'étudeSignal candidat: Observationnel · Signal consensuel: aucune
GenreSignal candidat: Empirique · Signal consensuel: aucune
Score de désaccord entre enseignants0,538
Score d'incertitude au seuil0,664

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,4620,483
Méta-épidémiologie (sens strict)0,0030,003
Méta-épidémiologie (sens large)0,0030,004
Bibliométrie0,0070,006
Études des sciences et des technologies0,0070,005
Communication savante0,0090,010
Science ouverte0,0090,010
Intégrité de la recherche0,0040,007
Charge utile insuffisante (le modèle a refusé de juger)0,0030,002

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,894
Tête enseignante GPT0,684
Écart entre enseignants0,210 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; l’étiquette directe de Gemma et le classifieur distillé Codex s’accordent sur ce qui est montré ici.

Devis d'étudeObservationnel
DomaineMéthodes
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations35
Publié2022
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueJMIR Formative ResearchMême sujetSurvey Methodology and NonresponseTravaux en français237 207