WikiBuild: A New Online Collaboration Process For Multistakeholder Tool Development and Consensus Building
Bibliographic record
Abstract
BACKGROUND: Production of media such as patient education tools requires methods that can integrate multiple stakeholder perspectives. Existing consensus techniques are poorly suited to design of visual media, can be expensive and logistically demanding, and are subject to caveats arising from group dynamics such as participant hierarchies. OBJECTIVE: Our objective was to develop a method that enables multistakeholder tool building while averting these difficulties. METHODS: We developed a wiki-inspired method and tested this through the collaborative design of an asthma action plan (AAP). In the development stage, we developed the Web-based tool by (1) establishing AAP content and format options, (2) building a Web-based application capable of representing each content and format permutation, (3) testing this tool among stakeholders, and (4) revising this tool based on stakeholder feedback. In the wiki stage, groups of participants used the revised tool in three separate 1-week "wiki" periods during which each group collaboratively authored an AAP by making multiple online selections. RESULTS: In the development stage, we recruited 16 participants (9/16 male) (4 pulmonologists, 4 primary care physicians, 3 certified asthma educators, and 5 patients) for system testing. The mean System Usability Scale (SUS) score for the tool used in testing was 72.2 (SD 10.2). In the wiki stage, we recruited 41 participants (15/41 male) (9 pulmonologists, 6 primary care physicians, 5 certified asthma educators, and 21 patients) from diverse locations. The mean SUS score for the revised tool was 75.9 (SD 19.6). Users made 872, 466, and 599 successful changes to the AAP in weeks 1, 2, and 3, respectively. The site was used actively for a mean of 32.0 hours per week, of which 3.1 hours per week (9.7%) constituted synchronous multiuser use (2-4 users at the same time). Participants averaged 23 (SD 33) minutes of login time and made 7.7 (SD 15) changes to the AAP per day. Among participants, 28/35 (80%) were satisfied with the final AAP, and only 3/34 (9%) perceived interstakeholder group hierarchies. CONCLUSION: Use of a wiki-inspired method allowed for effective collaborative design of content and format aspects of an AAP while minimizing logistical requirements, maximizing geographical representation, and mitigating hierarchical group dynamics. Our method faced unique software and hardware challenges, and raises certain questions regarding its effect on group functioning. Potential uses of our method are broad, and further studies are required.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.049 | 0.074 |
| Meta-epidemiology (narrow) | 0.003 | 0.002 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.006 | 0.002 |
| Science and technology studies | 0.003 | 0.004 |
| Scholarly communication | 0.009 | 0.013 |
| Open science | 0.006 | 0.019 |
| Research integrity | 0.004 | 0.004 |
| Insufficient payload (model declined to judge) | 0.012 | 0.005 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".