Validity Assessment of an Experimental e-Shopping Interface and Survey in a Multicultural Context
Bibliographic record
Abstract
Within a context of rapidly expanding international Internet penetration, organizations must offer multicultural versions of their information systems. In multicultural countries such as Canada, enterprises recognize the benefits of conducting e-business in a multilingual environment. As well, more and more IS research is conducted within a multicultural context. However, previous research has highlighted the difficulties of IS internationalization: differences in verbal style and word interpretation greatly increase the risk of miscommunication. In the case of multicultural IS research, miscommunication problems threaten the validity of the results. Internet-based shopping interfaces offer the opportunity of selling products and services across national and geographical boundaries. However, these interfaces must be adapted to multicultural settings. A research about the cognitive fit of two e-shopping interfaces was conducted in a multicultural context. For each of eight experimental conditions, two equivalent versions, in English and in French, were prepared. The subjects shopped online for a television set then answered a survey about their shopping experience. 561 consumers participated in the first study: 399 answered in English and 162 in French. The second study, a protocol analysis, required 39 consumers (13 English-speaking and 26 French-speaking) to complete the same online experiment while performing a think-aloud protocol. To ensure that results could be pooled across the two cultural contexts, great care was taken to ensure that a) each cultural version of the interface was valid and controlled for confounding variables, b) each cultural version of the survey met strict criteria of convergent and discriminating validity, but also c) that the two cultural versions were in every respect equivalent to each other. This third point was critical to the success of the experiment. To achieve this goal, 8 pre-tests were conducted. This paper presents the methodology used to ensure experimental interface and survey validity as well as their perfect equivalence. This research methodology represents a contribution to IS researchers as well as practitioners who internationalize web-based applications and must ensure their cultural context equivalence.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.114 | 0.175 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.003 | 0.004 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".