MétaCan
Menu
Back to cohort
Record W2898372466 · doi:10.2196/10995

Evaluation of Electronic and Paper-Pen Data Capturing Tools for Data Quality in a Public Health Survey in a Health and Demographic Surveillance Site, Ethiopia: Randomized Controlled Crossover Health Care Information Technology Evaluation

2018· article· en· W2898372466 on OpenAlexvenueno aff
Atinkut Alamirrew Zeleke, Adina Demissie, Fabian Otto‐Sobotka, Marc Wilken, Myriam Lipprandt, Binyam Tilahun, Rainer Röhrig

Bibliographic record

VenueJMIR mhealth and uhealth · 2018
Typearticle
Languageen
FieldSocial Sciences
TopicSurvey Methodology and Nonresponse
Canadian institutionsnot available
Fundersnot available
KeywordsData collectionPopularityHealth information technologyData qualityPublic health surveillancePublic healthHealth careElectronic data captureElectronic dataQuality (philosophy)Computer scienceMedicineData scienceBusinessPsychologyNursingDatabaseMarketingStatistics

Abstract

fetched live from OpenAlex

BACKGROUND: Periodic demographic health surveillance and surveys are the main sources of health information in developing countries. Conducting a survey requires extensive use of paper-pen and manual work and lengthy processes to generate the required information. Despite the rise of popularity in using electronic data collection systems to alleviate the problems, sufficient evidence is not available to support the use of electronic data capture (EDC) tools in interviewer-administered data collection processes. OBJECTIVE: This study aimed to compare data quality parameters in the data collected using mobile electronic and standard paper-based data capture tools in one of the health and demographic surveillance sites in northwest Ethiopia. METHODS: A randomized controlled crossover health care information technology evaluation was conducted from May 10, 2016, to June 3, 2016, in a demographic and surveillance site. A total of 12 interviewers, as 2 individuals (one of them with a tablet computer and the other with a paper-based questionnaire) in 6 groups were assigned in the 6 towns of the surveillance premises. Data collectors switched the data collection method based on computer-generated random order. Data were cleaned using a MySQL program and transferred to SPSS (IBM SPSS Statistics for Windows, Version 24.0) and R statistical software (R version 3.4.3, the R Foundation for Statistical Computing Platform) for analysis. Descriptive and mixed ordinal logistic analyses were employed. The qualitative interview audio record from the system users was transcribed, coded, categorized, and linked to the International Organization for Standardization 9241-part 10 dialogue principles for system usability. The usability of this open data kit-based system was assessed using quantitative System Usability Scale (SUS) and matching of qualitative data with the isometric dialogue principles. RESULTS: From the submitted 1246 complete records of questionnaires in each tool, 41.89% (522/1246) of the paper and pen data capture (PPDC) and 30.89% (385/1246) of the EDC tool questionnaires had one or more types of data quality errors. The overall error rates were 1.67% and 0.60% for PPDC and EDC, respectively. The chances of more errors on the PPDC tool were multiplied by 1.015 for each additional question in the interview compared with EDC. The SUS score of the data collectors was 85.6. In the qualitative data response mapping, EDC had more positive suitability of task responses with few error tolerance characteristics. CONCLUSIONS: EDC possessed significantly better data quality and efficiency compared with PPDC, explained with fewer errors, instant data submission, and easy handling. The EDC proved to be a usable data collection tool in the rural study setting. Implementation organization needs to consider consistent power source, decent internet connection, standby technical support, and security assurance for the mobile device users for planning full-fledged implementation and integration of the system in the surveillance site.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.049
metaresearch head score (Gemma)0.038
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Randomized trial · Consensus signal: Randomized trial
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.951
Threshold uncertainty score0.260

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0490.038
Meta-epidemiology (narrow)0.0020.002
Meta-epidemiology (broad)0.0030.003
Bibliometrics0.0020.002
Science and technology studies0.0020.002
Scholarly communication0.0020.002
Open science0.0020.002
Research integrity0.0030.002
Insufficient payload (model declined to judge)0.0060.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.468
GPT teacher head0.557
Teacher spread0.089 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designRandomized trial
DomainMethods
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations43
Published2018
Admission routes1
Has abstractyes

Explore more

Same venueJMIR mhealth and uhealthSame topicSurvey Methodology and NonresponseFrench-language works237,207