Identifying bioethical issues in biostatistical consulting: findings from a US national pilot survey of biostatisticians
Bibliographic record
Abstract
OBJECTIVES: The overall purposes of this first US national pilot study were to (1) test the feasibility of online administration of the Bioethical Issues in Biostatistical Consulting (BIBC) Questionnaire to a random sample of American Statistical Association (ASA) members; (2) determine the prevalence and relative severity of a broad array of bioethical violations requests that are presented to biostatisticians by investigators seeking biostatistical consultations; and (3) establish the sample size needed for a full-size phase II study. DESIGN: A descriptive survey as approved and endorsed by the ASA. PARTICIPANTS: Administered to a randomly drawn sample of 112 professional biostatisticians who were ASA members. PRIMARY AND SECONDARY OUTCOME MEASURES: The 18 bioethical violations were first ranked by perceived severity scores, then categorised into three perceived severity subcategories in order to identify seven 'top tier concern violations' and seven 'second tier concern violations'. RESULTS: Methodologically, this phase I pilot study demonstrated that the BIBC Questionnaire, as administered online to a random sample of ASA members, served to identify bioethical violations that occurred during biostatistical consultations, and provided data needed to establish the sample size needed for a full-scale phase II study. The No. 1 top tier concern was 'remove or alter some data records in order to better support the research hypothesis'. The No. 2 top tier concern was 'interpret the statistical findings based on expectation, not based on actual results'. In total, 14 of the 18 BIBC Questionnaire items, as judged by a combination of 'severity of violation' and 'frequency of occurrence over past 5 years', were rated by biostatisticians as 'top tier' or 'second tier' bioethical concerns. CONCLUSION: This pilot study gives clear evidence that researchers make requests of their biostatistical consultants that are not only rated as severe violations, but further that these requests occur quite frequently.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.030 | 0.112 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".