Bibliographic record
Abstract
Epidemiology is ‘the study of the distribution and determinants of health-related states or events in specified populations, and the application of this study to the control of health problems’ (Last, 1995). Historically, the science of epidemiology began with the study of outbreaks of infectious disease. It has since progressed in parallel with a shift, in the Western world, from infectious disease to chronic diseases as the major causes of morbidity and mortality. New branches have developed, such as clinical epidemiology, which again have spawned the concept of evidence-based medicine. The modern neurologist must now have both an understanding of the traditional concepts of epidemiology, such as incidence and prevalence of major diseases, and also a solid understanding of clinical research methods and how results of clinical trials apply to their patients. This chapter is designed, with brevity in mind, to provide an initial overview of these fundamentals. Population-based research in neurological disease Populations and sampling There is no substitute for good natural history data. In some cases, society has legislated that all instances of a disease be reported. Both rare diseases, such as rabies or previously more common diseases such as poliomyelitis, are reportable in most jurisdictions in the western world. Legislated data collection results in population-based information. It is no coincidence that these examples are both infectious diseases. With the shift to chronic or degenerative diseases such as atherosclerosis, cancer, arthritis as the leading killing and disabling illnesses, we have not made the same commitments to collecting data as with infectious diseases. We rely upon extrapolation from much smaller samples drawn from the population. If one could study every human being, it would be unnecessary to understand sampling. Even in very large studies, researchers can only study a tiny proportion of the population; pragmatism dictates it. The principle of sampling is to select from the population a truly representative group for study. A population is any group of persons described as generally or as specifically as appropriate. One may study the entire population of the Western hemisphere, the population of North American First Nations peoples, or the population born in a particular year or years, e.g. the ‘baby boomers’. A sampling unit is the basic unit of sampling, e.g. individual, family, city, etc. This entire list of sampling units is called the sampling frame. The sample is derived from the sampling frame.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.009 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.002 | 0.010 |
| Scholarly communication | 0.007 | 0.006 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.004 | 0.006 |
| Insufficient payload (model declined to judge) | 0.017 | 0.012 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".