L'efficacité de l'échantillonnage passif pour obtenir un portrait représentatif de l'électorat: Le cas de<i>Vote au pluriel – Québec</i>
Bibliographic record
Abstract
Résumé On estime généralement que l'utilisation d'une méthode d’échantillonnage actif pour la conduite de sondage Internet est plus efficace que l'utilisation d’échantillonnage passif, autant en ce qui a trait au recrutement des participants que de la représentativité de l’échantillon. Dans le premier cas, des firmes spécialisées contactent directement des répondants, alors que dans le deuxième cas les participants doivent se rendre par eux-mêmes sur une plate-forme de sondage en ligne. Dans cet article, nous évaluons empiriquement l'efficacité relative d'une stratégie d’échantillonnage passif à l'aide d'une campagne médiatique mise en place dans le cadre de l'expérience WebVote au pluriel – Québec. Cette stratégie avait pour objectif d'augmenter le nombre de participants et d'améliorer la représentativité de l’échantillon obtenu par échantillonnage passif. Nos résultats suggèrent que la stratégie a eu un impact significatif, mais limité, sur le recrutement des participants. De plus, la représentativité de l’échantillon complet comprend d'importantes lacunes, surtout lorsque celui-ci est comparé à un échantillon Web obtenu par échantillonnage actif. Nos analyses démontrent cependant que l'utilisation des médias traditionnels, dans le cadre d'une stratégie plus large, pourrait diminuer l'impact négatif dudigital divide, qui nuit à la représentativité des échantillons récoltés sur le Web.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.010 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.003 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.021 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".