Analyse du marché du travail à l’aide des données de Google Trends
Bibliographic record
Abstract
In this report, we evaluate the relevance of weekly Google search query data for current and next month prediction on several labour market variables in Canada and Quebec. Several types of mixed-frequency models are considered and their performance is evaluated in an out-of-sample forecasting exercise spanning the period 2014M09 - 2019M09. Google Trends improve the accuracy of forecasts of the employment rate, hours worked and unemployment rate. The availability of this data in high frequency is crucial. Their contribution is important especially during the first two weeks of the month, so when Labor Force Survey data are not yet available for the last month. Dans ce rapport, nous évaluons la pertinence des données hebdomadaires des requêtes faites sur le moteur de recherche de Google au niveau de la prédiction du mois courant et du prochain mois sur plusieurs variables du marché d’emploi au Canada et au Québec. Plusieurs types de modèles en fréquence mixte sont considérés et leur performance est évaluée dans un exercice de prévision hors échantillon s’étalant sur la période 2014M09 - 2019M09. Les Google Trends améliorent la précision des prévisions du taux d’emploi, des heures travaillées et du taux de chômage. La disponibilité de ces données en haute fréquence est cruciale. Leur apport est important surtout durant les deux premières semaines du mois, donc lorsque les données de l’Enquête sur la population active ne sont pas encore disponibles pour le dernier mois.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.002 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".