The developmental pathways of major league baseball players and their influence on career performance
Bibliographic record
Abstract
Performance and developmental data of 2,291 American-born Major League Baseball (MLB) players who debuted between 1990 and 2010 were amassed using baseball-reference.com. Two performance indicators: career games played and wins above replacement (WAR; player's total contributions in wins) were coupled with pre-draft data to determine the influence of developmental pathways on career success. Non-linearity of athlete development (Gulbin et al., 2013) was prevalent as 17 qualitatively different pathways to MLB were identified through draft information. When distilled, analyses reveal 63% of the athletes started their career directly after attending a four-year institution (23% high school, 13% junior college) and 79% did not sign or were not selected as high school draft picks. There were statistically significant differences in career MLB (F (2, 2,288) = 3.63, p < .05) and Minor League Baseball (MiLB) (F (2, 2,228) = 9.07, p < .001) games played with athletes drafted directly from high school averaging 48 to 50 more MLB and 66 to 77 MiLB games than those drafted from a junior college or four-year institution. No statistically significant differences between career WAR metrics were observed, but the difficulty of obtaining career success via this metric was noted as only 48.4% of athletes in this sample achieved a positive WAR. The collection of milestone data and additional performance indicators is needed to understand the variation within and between pathways, which may have important implications for improving talent identification accuracy (Koz et al., 2012) and developmental programs.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".