The fastest 24-hour ultramarathoners are from Eastern Europe
Bibliographic record
Abstract
Ultramarathon running is of increasing popularity, where the time-limited 24-hour run is one of the most popular events. Although we have a high scientific knowledge about different topics for this specific race format, we do not know where the best 24-hour runners originate from and where the fastest races are held. The purpose of the present study was to investigate the origin of these runners and the fastest race locations. A machine learning model based on the XG Boost algorithm was built to predict running speed based on the athlete´s age, gender, country of origin and the country where the race takes place. Model explainability tools were used to investigate how each independent variable would influence the predicted running speed. A sample of 171,358 race records from 63,514 unique runners from 73 countries participating in 24-hour races held in 57 countries between 1807 and 2022 was analyzed. Most of the athletes originated from the USA, France, Germany, Great Britain, Italy, Japan, Russia, Australia, Austria, and Canada. Tunisian athletes achieved the fastest average running speed, followed by runners from Russia, Latvia, Lithuania, Island, Croatia, Slovenia, and Israel. Regarding the country of the event, the ranking looks quite similar to the participation by the athlete, suggesting a high correlation between the country of origin and the country of the event. The fastest 24-hour races are recorded in Israel, Romania, Korea, the Netherlands, Russia, and Taiwan. On average, men were 0.4 km/h faster than women, and the fastest runners belonged to age groups 35-39, 40-44, and 45-49 years. In summary, the 24-hour race format is spread over the world, and the fastest athletes mainly originate from Eastern Europe, while the fastest races were organized in European and Asian countries.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.007 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".