Bibliographic record
Abstract
Data analysis encompasses the use of statistical methods to describe data, test hypotheses and estimate measures of effect such as risk, relative risk and survival probabilities. As a post-graduate student in Canada, I had two excellent teachers who introduced me to the world of epidemiology and statistics. I still remember their words and have paraphrased some of them here to explain how I developed my approach to clinical research. An approach to your own research Ask yourself: “What is my hypothesis?” “What am I really trying to show?” “Can I simplify things?” Don’t ask too many questions - keep it simple. From the hypothesis, you can create a table of what you expect the results will look like and start to think about what numbers will go into the table. First the purpose of the study is stated - often as a research question phrased in the form of a testable hypothesis. Then the study is designed to collect the appropriate data. Once a study has been completed, the data are entered into an electronic database, spreadsheet or statistical package for analysis. Data are normally entered using the columns for the variables and entering the cases in rows such that the data for each study subject is entered in one row. A recent study (de Risio and others 2008) investigated the association of clinical and magnetic resonance imaging (MRI) findings with outcome in dogs with presumed ischaemic myelopathy. This study is classified as a retrospective case series (Cardwell 2008) since the exposures (clinical and MRI findings) were recorded at presentation at a referral hospital and outcome of interest (clinical outcome) was evaluated at a later time. Another study (Koffas and others 2008) used colour M-mode tissue Doppler imaging to detect differences in the myocardium of healthy cats and cats with hypertrophic cardiomyopathy. This study is classified as a cross-sectional study (Cardwell 2008) since the exposure (disease status) and outcome of interest (measurements from tissue Doppler imaging) were evaluated at the same time. Variables are defined in terms of the type of data they represent and are classified as either categorical (discrete) or continuous. The NOIR system is commonly used to define the type of data as nominal, ordinal, interval or ratio (Table 1). A variable can also be considered to be dependent or independent. The value of a dependent variable depends on (or can be predicted by) the value of some other variable. Thus, a dependent variable is also called an outcome or response variable. One of the dependent or response variables that was examined in the ischaemic myelopathy study was the outcome of the cases. Outcome was defined as being either successful or unsuccessful (a dichotomous categorical variable). The dependent or response variable that was examined in the echocardiography study included tissue Doppler imaging measurements such as myocardial velocity gradient and mean myocardial velocities (continuous variables). Three steps from research question to statistical model Consider the type (NOIR) of the dependent(outcome or response) variable. Categorical binary/dichotomous (nominal) e.g. alive or dead at the end of the study 2 categories (ordinal) e.g. disease absent (0) or present (1) >2 categories (nominal) e.g. blood group (A, B, AB) ranked categories (ordinal) e.g. cancer stage (I, II, III) Continuous interval e.g. Glasgow Coma Scale score (1-18) ratio e.g. red blood cell count Consider the type (NOIR) of independent (exposure or predictor) variable(s) as above. Choose the statistical test(s) appropriate to the type of dependent and independent variables that you plan to include in your study (Table 2). An independent variable is defined as an explanatory variable that is measured and hypothesised to be associated with an outcome of interest (dependent variable). Thus, an independent variable is also called an exposure or predictor variable. In the ischaemic myelopathy study, the independent or exposure variables included: neuroanatomic location of the lesion, treatment prior to referral, upper vs lower motor neuron signs on presentation and whether the lesion was symmetrical or not (nominal categorical variables with two or more unordered categories). In the analysis of data from the echocardiography study, the main independent or exposure variable was the cardiac disease status of the cats (normal or affected with hypertrophic cardiomyopathy), a binary categorical variable. Additional independent variables that were included in the analysis included the R-R interval, age and weight (all continuous variables). Based on the classification of the independent and dependent variables, there are four basic types of data sets that can occur. For each of these there are different methods of statistical analysis available. The statistical models presented in Table 2 include methods for evaluating how much of an outcome occurred or whether or not an event occurred. Therefore, in the ischaemic myelopathy study a contingency table or cross-tabulation with chi-square or Fisher’s exact test is the appropriate approach to analysis to examine the effect of each of the independent variables mentioned above with the dichotomous categorical outcome variable (successful or unsuccessful outcome). In the echocardiography study with a categorical main exposure variable, several continuous independent variables and a continuous outcome variable, linear regression was an appropriate approach to the analysis. This study found statistically significant differences in several of the tissue Doppler imaging measurements. This approach can be extended to all types of data, including “messy” data that might include the presence of repeated measurements on individual animals or the occurrence of unbalanced data sets due to missing data. The ability to extend this approach allows us to consider some of the more advanced methods of statistical analysis such as multiple regression (with ≥2 independent variables that can be a mix of continuous and categorical), mixed or multi-level models (with ≥2 levels of measurements that need to be taken into account, such as when looking at kittens within litters or clinicians within a practice) and survival analysis. Vicki Adams graduated from the Western College of Veterinary Medicine in Saskatoon in 1990 and went on to complete a one-year small animal internship at the University of Minnesota. After seven years in general and emergency small animal practice, she returned to the University of Saskatchewan to do research. Having obtained an MSc in the epidemiology of rabies in wildlife, Vicki completed a PhD in small animal epidemiology investigating owner compliance with veterinary recommendations and prescribed medications. Vicki started working at the Animal Health Trust in January 2003 and is currently Head of the Small Animal Epidemiology Unit. With grateful thanks to Drs Carl Ribble and John Campbell for their very wise words.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.124 | 0.339 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.004 | 0.004 |
| Bibliometrics | 0.009 | 0.010 |
| Science and technology studies | 0.002 | 0.004 |
| Scholarly communication | 0.009 | 0.006 |
| Open science | 0.005 | 0.005 |
| Research integrity | 0.003 | 0.008 |
| Insufficient payload (model declined to judge) | 0.123 | 0.033 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".