The relative age effect in soccer: a case study of the São Paulo Football Club
Bibliographic record
Abstract
The aim of this study was to compare the birth-date distribution of youth athletes of a high-level Brazilian soccer club with the general population of the same age group. In a cross-sectional study, the birth date of 341 youth athletes (under 10-20) was compared with a reference population (live births that occurred in São Paulo state in the same age group; n = 5,480,868). The subjects were divided into quarters of birth: 1st = January-March; 2nd = April-June; 3rd = July-September; 4th = October-December. The chisquare test (χ2) was used to compare the expected (reference population) and observed (athletes) distributions. It was detected a significant difference between the expected distribution and observed distribution (χ2= 29.53; p<0.0001), with a higher percentage of athletes born in the 1st quarter (47.5%) and a lower percentage in the 4th quarter (8.8%). The present results confirm the occurrence of the relative age effect (RAE) during the player selection process in a top-level Brazilian soccer club. The occurrence of this phenomenon during the selection and development of young athletes needs to be taken into account and should be analyzed carefully in order to minimize the loss of potential youth soccer talent. Further studies are required to identify the determinants of RAE and to establish preventive strategies that ensure a more efficient selection process of young soccer players.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".