The GEMS Exams in Israel—Between Center and Periphery
Bibliographic record
Abstract
The relationship between input and output is one of the leading topics in modern educational discourse. In the current study, we focus on the GEMS (Growth and Effectiveness Measures for Schools) exams in Israel, which constitute a measure of a school’s scholastic achievements and its academic-social climate. The GEMS is the equivalent of exams such as the TIMSS and the PISA, used in other countries. The GEMS is a school supervisory tool of major importance operated by Israel’s Ministry of Education for improving scholastic achievements and academic-social climate in schools. As an objective indicator, GEMS scores open the field of education to competition. Data on all schools in Israel whose students take the GEMS also appear on the National Authority for Measurement and Assessment in Education (RAMA) website. This study aims to examine the reasons for the disparities in the GEMS results between Israel’s center and periphery and explores whether they can be reduced. Studies published on this issue in the last 16 years that explored these disparities, which are reflected in the extent of parental involvement and students’ educational deficits, were conducted on behalf of RAMA and under its supervision, and some were not sufficiently critical in their review of the efficacy of the GEMS exams. Identifying and understanding these causes is a significant step toward reducing the disparities which have important implications for future acquisition of a secondary and tertiary education. The research findings offer a practical contribution for policy makers in the educational system, while identifying elements of positive change in the schools.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".