Cohort Profile: DOC*X: a nationwide Danish occupational cohort with eXposure data – an open research resource
Bibliographic record
Abstract
Occupational exposures contribute substantially to morbidity and mortality worldwide1,2 and are of pivotal importance for sick leave, disability and preterm exit from the labour market regardless of the aetiology of disease.3,4 Premature exit from the labour market is becoming more important as increasing life expectancy is calling for an extension of working life into older ages.5 To comply with demands for knowledge on occupational determinants for maintaining a healthy workforce, there is a need for building large nationwide databases that integrate prospective data on occupational exposures, with data on health, work participation and labour market exits. Health research in Denmark has unique opportunities to establish such databases due to well-established nationwide administrative and health registers (e.g. Danish registers on personal labour market affiliation6 and The Danish National Patient Register7) and the possibility to link information on an individual level across registers by the personal identification number established on April 2, 1968.8 Nevertheless, occupational epidemiology has not yet fully capitalized on these possibilities, partly because occupational exposures are not included in the administrative registers on labour market affiliation.6 The Danish Occupational Cohort with eXposure data (DOC*X) aims to fill this gap by including a series of job exposure matrices (JEMs) that to the largest extent possible builds on uniform principles and metrics across different occupational exposures. A JEM is a cross tabulation of jobs or job groups; in our case occupational codes according to DISCO-88 (the Danish version of the International Standard Classification of Occupations, ISCO, from 1988) and corresponding exposure estimates.9,10 Occupational exposure assessment by means of population-based JEMs was described in the early 1980s.11–13 Since then, JEMs have gained increasing acceptance in the scientific community10,14,15 and have a great potential for cost-effective exposure assessment in large-scale studies.16 The DOC*X database benefits from progress made in constructing JEMs during the past decade and includes psychosocial, mechanical, physical and chemical exposures, and lifestyle factors. Since the exposures are assessed at the same level of detail (DISCO-88 code level), the DOC*X database is designed to provide joint estimates of risks related to exposures that are most often analysed separately. In short, the objective of DOC*X is to provide a research resource that enables cost-effective studies of the impact of occupational exposures on disease and disability in a nationwide setting, by using administrative and health registers together with JEMs. The cohort includes all persons gainfully employed and living in Denmark in 1970 and from 1976 to 2015, regardless of nationality. The period 1971–1975 is not covered as Statistics Denmark changed from survey-based to register-based censuses and did not carry out censuses in these years. This amounts to approximately 6.4 million persons, whose occupational history was followed for a median of 15 years (range 1–41 years). Between 2.0 and 2.9 million persons were in employment yearly, with a gradual increase from the earlier to the later years. Further details on source registers and labour market attachment are provided in the ‘What has been measured?’ section. Key characteristics of the cohort are provided in Table 1. Description of the cohort in the DOC*X database Rounded to nearest thousand. DISCO-88 is the Danish version of the International Classification of Occupations (ISCO). Missing DISCO-88 code even though the person was registered as in employment (employed or self-employed). The 1970 registration of occupation used a classification system that is not generally harmonizable to DISCO-88. DB-07 (Dansk Branchekode 2007) is the Danish extended 6-digit version of the ‘Nomenclature statistique des Activités économiques dans la Communauté Européenne’. 1st to 9th deciles. Employment status and occupation have been followed through the 1970 census17 and by annual registrations from 1976 to 20156 for all persons aged at least 16 years. Further annual registrations will be added to the database in the future. In parallel, continuous registration of migrations, disappearances and deaths enables precise timing of such events within the entire cohort from 1970 and onwards.8 The DOC*X database is made up of three main parts: (i) The Occupation and Industry Register, (ii) the JEMs, and (iii) register-based information on demographics, education, income, migration, retirement, death, hospitalizations and transfer payments. (i) The Occupation and Industry Register. This register was compiled for DOC*X. It contains yearly registrations of employment status from the 1970 census,17 which was based on mandatory self-reported information, and from the Employment Classification Module 1976 to 2015.6,18 Both sources are from Statistics Denmark. The Employment Classification Module is based on various collections of administrative data including self-report to the civil registration authorities, union membership, public employers, private companies and tax records. Self-report and union membership were predominantly used earlier in the period, whereas tax records and information from public and private employers were used increasingly later in the period, as it became mandatory for companies with 10 or more employees to report individual-level information on occupation. Assigning occupational codes was carried out by trained coders at Statistics Denmark in the earlier part of the period but is increasingly reliant on direct employer reports through automated systems. Employment status is divided into three categories: employed, self-employed and non-employed. The two former categories are linked with information on occupation and industry. The source registers used several different classifications of occupation and industry. The 1970 census used a distinct classification developed by Statistics Denmark. The Employment Classification Module has used three classifications. (1) 1976–1992: a scheme developed by Statistics Denmark based on ISCO-68, (2) 1993–2009: DISCO-88, and (3) 2010–2015: DISCO-08 The DISCO-88 and -08 are Danish versions of the International Standard Classification of Occupations (ISCO) from 1988 and 2008, respectively.19 The DISCO system classifies occupations into 10 major groups (1st digit level: 0, Armed forces; 1, Legislators, senior officials, and managers; 2, Professionals; 3, Technicians and associate professionals; … 8, Plant and machine operators, and assemblers; 9, Elementary occupations) and 372 groups at the 4th and most detailed level (e.g. 1210, Directors and chief executives; 2221, Medical doctors; 3231, Nursing; 7412, Bakers; and 9313, Building construction labourers). The DISCO classification is based on a combination of education/skill and job content. To enable studies of long-term exposures and studies with long-term follow-up, we have introduced a harmonized coding of occupations according to DISCO-88 as it was the primary classification for the longest period and is more detailed than DISCO-08.19 The harmonization was performed code by code, aiming to keep the reclassification as detailed as possible. DOC*X also contains harmonized industry codes according to ‘Dansk Branchekode 2007’ (DB-07), the Danish extended 6-digit version of the ‘Nomenclature statistique des Activités économiques dans la Communauté Européenne’ (NACE).20 Industry codes were drawn directly from the source registers, with harmonization carried out by Statistics Denmark in advance. The individual-level codes are based on industry coding of companies, with all employees in a given company given the common company code. Industry coding of companies is mandatory. Additionally, the register retains occupation and industry codes according to the original classifications. Figure 1 gives an overview of the percentage of employed or self-employed people with harmonized DISCO-88 codes by year. On average, 79% had a DISCO-88 code at the major group level (1st digit level, i.e. the 10 major DISCO-88 categories), with somewhat higher proportions of persons with codes for the later years. The percentage with DISCO-88 codes was lower at the 2nd–4th digit levels, particularly in the 1980s, but shows the same temporal trends. The main reason for missing DISCO codes is that only companies with more than 10 employees are required to report occupational information. In general, men had a lower coverage of DISCO-88 codes (on average at the 4th digit level: men 64%, women 71%). This is primarily an effect of varying registration practices across the somewhat gender segregated Danish labour market. Men are more frequently employed as e.g. ‘Craft and related trades workers’ and ‘Plants and machine operators and assemblers’, whereas women constitute the majorities in ‘Clerks’ and ‘Service workers and shop and market sales’, where companies are typically larger and thus required to report occupational titles. Percentage of persons in employment with DISCO-88 codes from 1976 to 2015 at 1st–4th digit level of classification. Employment status is measured from the year of the 16th birthday or from the year of first registered employment, whichever comes latest, until the last year of registered employment, or 2015, whichever comes first. With future yearly updates, the Occupation and Industry Register will continue to include up-to-date information on employment and occupation. (ii) The job exposure matrices. The DOC*X database contains measures of a wide variety of occupational exposures provided by JEMs. Table 2 presents an overview of the occupational exposure matrices that are so far included in DOC*X. To enable linkage between the Occupation and Industry Register and the JEMs, all exposure estimates are classified according to DISCO-88 and DB-07. Additionally, the included JEMs may find use in combination with other databases with similar coding of occupation and industry. Job exposure matrices included in the DOC*X database (examples of contents) The psychosocial JEM includes six exposures (Table 2)21 that are based on data from a questionnaire survey of a random sample of 15 207 Danish employees in 2012: The Work Environment and Health in Denmark cohort study.22 For continuous exposures, the DISCO group specific mean levels (and standard deviations) of scale values were generated in a random intercept multilevel model using best linear unbiased prediction (BLUP) estimators. For dichotomous exposures, the DISCO code specific probability of exposures was generated using a logit model with DISCO code and age as predictors. Continuous and dichotomous exposure estimates were generated separately for men and women, and age was modelled as a spline according to the sex-specific quartiles of the age-distribution in the population. The JEM thus contains specific exposure estimates according to occupation, sex and one-year age categories. The mechanical JEMs (the Shoulder JEM23 and the Lower Body JEM24) include 18 exposures such as time per day spent working with upper arm elevation >90°23 and total load lifted per day24 (Table 2). These JEMs have been constructed based on expert ratings. The estimates of upper arm elevation and repetition in a range of jobs in the Shoulder JEM have been validated against inclinometer measurements23 and the Shoulder JEM has shown good predictive validity with respect to surgery for subacromial impingement syndrome.25–28 The Lower Body JEM was originally evaluated by two external experts, who in general agreed with the ranking of exposures,24 and the JEM has shown good predictive validity in studies of inguinal hernia repair,29–31 total hip replacement32 and surgery for varicose veins.33 Both mechanical JEMs have shown good predictive validity with respect to sickness absence and permanent work disability in relation to upper- and lower-body pain.34 The JEM on physical exertion and body position covers an index score based on self-report in the domains sitting, standing/walking, kneeling, lifted arms, repetitive arm movements, twisted back, lifting/carrying and pushing/pulling. Sex- and age-specific measures have been computed by BLUP estimation (Table 2).21 The physical JEMs cover whole-body vibration from the Lower Body JEM24 and noise measured as A-weighted sound level in decibels (dB) (Table 2). Noise exposure estimates are based on full-shift measurements using personal samplers in a range of occupations35 in combination with expert ratings. The particulate airborne exposure JEMs provide data on exposure to mineral dust, organic dust, and fumes and vapours, with exposure intensity categorized as no, medium, and high average exposure (Table 2). This JEM is similar to the ALOHA JEM.36,37 Furthermore, a wood dust JEM with quantified airborne exposure to inhalable wood dust is available based on 12 704 dust measurements collected between 1978 and 2007 from wood-related industries in Denmark, Finland, the UK, France, Norway and The Netherlands.38 An endotoxin JEM with quantified airborne exposure to inhalable endotoxin (a major constituent in organic dust) is also available and is based on 3350 dust measurements collected between 1992 and 2008 from agricultural industries in Denmark, Germany, Canada, Norway and The Netherlands.39 The chemical JEMs provide calendar time specific (1945–59, 1960–74, 1975–84, 1985–94) exposure intensity and prevalence estimates for potential and confirmed carcinogenic exposures (asbestos, quartz, wood-dust, diesel-exhaust, formaldehyde and selected organic solvents) that have historically been prevalent in Danish workplaces (Table 2). The JEM is based on the template of a Finnish JEM (FINJEM40) but includes Danish exposure measurements where possible.10 The original JEM used occupational codes according to NYK (Nordisk Yrkes Klassifikation), which corresponds to ISCO-1958 and has been translated into DISCO-88. The lifestyle JEMs cover lifestyle factors associated with both job/industry and health/disability. Lifestyle factors covered are smoking, alcohol consumption, body mass index, leisure-time physical activity and intake of fruit and vegetables (Table 2).41 The JEMs are based on a collection of representative Danish surveys with in total almost 300 000 participants covering the period from 1973 to 2013, where information on lifestyle factors was collected in consistent ways combined with information on occupation at the time of survey participation. The lifestyle JEMs are computed in a random intercept multilevel model using BLUP estimators by DISCO-88 code. These JEMs are included to facilitate adjustment for potentially confounding lifestyle factors, which are rarely available in nationwide registers. Whereas inclusion of socio-economic status may be used as a proxy for lifestyle factors, socio-economic status (in a Danish register setting) is dependent on occupation, and thus not a good choice when studying occupational health. (iii) Supplementary data. The third part of the DOC*X database consists of register-based sociodemographic, health and transfer payment data that can be linked with the Occupation and Industry Register on an individual level. These data include sex, date of birth, education, income, migration, death, hospitalizations, transfer payments and labour market exit (Figure 2). A large part of the supplementary data are yearly status data, e.g. age, place of living, family structure (single, cohabitating, married, number of children), income and education, which are found in registers at Statistics Denmark (BEF, FAM, FAIK, UDD).42 Some data are registered continuously with dates, e.g. migration, retirement and death, which are also found in registers at Statistics Denmark (VNDR, AKM, DODE).42 Supplementary data in the DOC*X database with years of availability. Health outcomes are included in the DOC*X database by individual-level linkage to the National Patient Register7,43 and the Causes of Death Register.44 Data on birth outcomes from the Medical Birth Register45 are also available in the DOC*X database together with linkage between parents and children. Data from the Prescription Register46 and specialized registers on specific diseases (e.g. cancer47 and heart disease48) may be included on a study-by-study basis. Data on transfer payments (e.g. sickness absence, disability pension, age pensioning and early retirement) are included from the Danish National Register on Public Transfer Payments (The DREAM database) starting in 1991.49–51 These supplementary data enable categorization of periods of unemployment due to various reasons and studies of return to work. Using the DOC*X database, we have analysed the rate of permanent retirement from the labour market (disability pension, age pensioning and early retirement using data from the Danish National Register on Public Transfer Payments) according to major DISCO-88 groups in the Occupation and Industry Register (Figure 3), and according to occupational exposure to total load lifted per day by use of the Lower Body JEM24 (Table 3; Figure 4). We calculated sex- and time-period- (in 5-year groups) specific age-standardized retirement rate ratios, using the average sex-specific retirement rates in the entire period as reference. Major group ‘0, Armed forces’ was excluded, as the pensioning scheme of the armed forces differs from the rest of the labour market, and ‘Skilled agricultural and fishery workers’ was excluded women due to in retirement rates across DISCO-88 major groups were the period, with higher rates in DISCO-88 major groups but a of from the We a similar of of in occupations according to quartiles of total load lifted per day (Figure 4). retirement rate (and in Denmark according to DISCO-88 major groups to 2015 for men and women DISCO-88 is the Danish version of the International Classification of Occupations version retirement rate (and in Denmark in DISCO-88 groups according to total load lifted per day quantified by The Lower Body Job to 2015 for men and women DISCO-88 is the Danish version of the International Classification of Occupations version Percentage of total load lifted per day according to The Lower Body JEM exposure by calendar year periods The Lower Body which covers the period from is also to the of with exposure status the period covered by the DOC*X The not up to due to to DISCO code. of research that has DOC*X JEMs to nationwide data, are studies of disease in relation to noise at the and studies of in relation to exposure to formaldehyde and also in the of the mechanical The main are the nationwide inclusion of the major part of the Danish workforce, the more than years coverage of the Occupation and Industry Register, and the of JEMs. In the of supplementary data with almost covering the entire cohort enables detailed of between occupational exposures and disease and disability outcomes with adjustment for other potentially factors including The main are the registration of DISCO codes in the Occupation and Industry Register, with up to missing data in time particularly for companies with where report of occupational is not and of for DISCO as even the most detailed 4th digit level may include several different which may different exposures. These may be by of DISCO-88 codes with high and of in combination with industry which are also part of the DOC*X To these we work to include information from the Register of the Danish Supplementary a mandatory employer employment period and industry information, but occupational on all Danish employees from The validity of the DOC*X occupational the DISCO-88 classification using self-reported information on occupation as standard is in a for A tabulation of the between DISCO-88 codes and self-reported occupation will be available for specific DISCO-88 codes at the DOC*X to in occupations with the most and DISCO-88 the registers carry information on lifestyle factors, which may be important of the between occupational exposures and health JEMs on lifestyle factors were constructed as an in the database to and such potential exposure assessment may be by JEMs not of exposure within job which may be with the between job This may be in a large cohort DOC*X by occupations with large of the exposure of large between job groups to the within job by occupational is to out by the as into related DISCO-88 codes in similar exposure Furthermore, the exposure assessment that is in JEMs and as the main part of the will be than the of estimates due to of exposure is or in the JEM estimates measures of average exposure at the DISCO group level. This means that will be with a where individual exposure levels are but this may not be given the large sample of the DOC*X to use data is through the DOC*X at Further information on of the DOC*X database, and on to is available at the DOC*X data in the DOC*X database are available through at Statistics Denmark standard The DOC*X database was by from the Danish Environment and from the Danish of The cohort was to for register-based research in occupational health, with a nationwide occupational classification of the major part of the Danish in 1970 and from 1976 to The cohort covers the major part of the Danish working from age 16 in the period from 1970 to 2015 approximately 6.4 million The register-based of the cohort yearly on occupations and continuous on most outcomes from 1970 or to 2015 and The cohort contains yearly information on employment occupation and industry in a harmonized coding across into occupational exposures is by means of job exposure matrices. for be to or at The of Occupational and Denmark. The DOC*X contains information on to individual-level data, which is available through at Statistics Denmark standard
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".