Bibliographic record
Abstract
The gross job gains and gross job loss statistics from the BLS Business Employment Dynamics (BED) program measure the large gross job flows that underlie the quarterly net change in employment. In the fourth quarter of 2004, employment grew by 869,000 jobs. This growth is the sum of 8.1 million gross job gains from opening and expanding establishments, and 7.2 million gross job losses from contracting and closing establishments. The new BED data have captured the attention of economists and policymakers across the country, and these data are becoming a major contributor to our understanding of employment growth and business cycles in the U.S. economy. Following the initial release of the BED data in September 2003, the BED data series expanded in May 2004 with the release of industry statistics. The BLS then began work on tabulations by size class. The production of size-class statistics is a complex task involving several economic and statistical issues. Although it is trivial to classify a business into a size class in any given quarter, it is difficult to classify a business into a size class for a longitudinal analysis of employment growth. Several different classifications exist, and many of these possible classifications have appealing theoretical and statistical properties. Furthermore, these alternative classification methodologies result in sharply different portraits of employment growth by size class. In this article, we discuss the alternative statistical methodologies that the BLS considered for creating size class tabulations from the Business Employment Dynamics data. Our primary focus is on four methodologies: quarterly base-sizing, annual base-sizing, mean-sizing, and dynamic-sizing. We discuss the evaluation criteria that BLS considered for choosing its official size class methodology.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.024 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.006 | 0.008 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.016 | 0.005 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".