LET’S HEAR IT ONCE MORE FOR THE UNSUNG COACH: ON THE EFFICIENCY OF COACHING FOR THE UNIFIED STATE EXAM
Bibliographic record
Abstract
Introduction. With the introduction of a new attestation procedure of school graduates in the form of the Unified State Examination (USE), coaching has gained widespread acceptance in Russia. By some estimates, between a quarter and a half of school graduates have recourse to one-to-one coaching when preparing for the USE. However, the question of the efficiency of such lessons remains open. Recently, various publications in professional periodicals and the media have begun to appear, which cast doubt on the benefits of coaching. The authors of these publications are specialists of the Higher School of Economics (HSE). According to their studies, additional lessons in preparing for the USE, including those with coaches, have a very little effect. Theaimof the research was to discuss the validity of the HSE specialists’ arguments concerning the low efficiency of coaching activities. Methodology and research methods. In the course of studying the problem, a comprehensive research methodology was applied, including approaches for comparative and statistical analysis of data and materials published by the HSE, Federal Institute of Pedagogical Measurements (FIPI) and Federal Testing Centre. Results and scientific novelty. An analysis of scientific works published by the HSE specialists showed that their conclusions with regard to the claimed low efficiency of additional lessons in preparation for the USE are unsubstantiated due to the presence of gross methodological errors in the calculations. Firstly, the students’ initial level of knowledge prior to lessons with a coach was miscalculated, with the final school grades in Russian language and mathematics being taken as the initial level instead of the average score of the certificate. Secondly, the specialists ignored the fact that the final grade “two” does not exist in the school attestation system. In this regard, the models used by the HSE specialists’ did not allow the progress in training from the school grade “three” to the USE “three” evaluation to be adequately recognised. Thirdly, the determination of the efficiency of coaching was made without taking the specific character of different teaching disciplines into account. Thus, the reliance on formal mathematical procedures to the detriment of content problem analysis led the specialists of the HSE to snap judgements that do not reflect the true situation. Practical significance.The authors believe that the observations provided in this paper will help education specialists to adjust approaches when determining the efficiency of additional lessons during USE preparation.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.026 | 0.217 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.003 |
| Scholarly communication | 0.003 | 0.004 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.011 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".