Code for replication: Absenteeism, Productivity, and Relational Contracts Inside the Firm
Bibliographic record
Abstract
Achyuta Adhvaryu, Jean-François Gauthier, Anant Nyshadham and Jorge Tamayo This document describes the instructions for: (1) accessing the datasets used in the manuscript, (2) the program files used to replicate the analysis of this study, (3) steps to replicate results. The program files replicate not just the tables and figures in the main paper, but also all tables and figures in the appendices. 1. Data Availability and Provenance Statement The data used in this study were provided to us by the firm partner (Shahi Exports Pvt. Ltd.) on an exclusive and proprietary basis. We have entered into a data-sharing agreement with the firm to allow access to the de-identified data upon request for replication purposes. Please email datarequest@shahi.co.in to submit a request. Once the request for replication data is received, we can share the anonymized data after the requester signs an agreement with the firm that the data will not be shared further and will be used only for the replication purpose. We obtained two types of data which we combined to generate the analysis file: (1) Administrative data containing information about production, absenteeism, and productivity by line. (2) Responses to survey from line supervisors. For the main regression analysis, we combine line-level information to create a dyadic dataset called “raw_reg.dta”. For supporting tables and figures it is often more convenient to use the line level data. All the datasets can be found in the “Data” folder upon request. 2. Code Description All codes are in the “Code” folder. PPML_reg_final.do runs all the main PPML regressions of the results section and those in appendix. We have enclosed our ado files in case changes in the ppmlhdfe function is updated by the time of replication. Note that ppmlhdfe is a form of maximum likelihood command and therefore may produce slightly different results when run on a different number of cores. For replication purposes, the .do file runs the command on a single core. Final_support_tables.do execute all other tables of the paper and appendix. Final_figures.do generates all figures of the paper. Figures 8 and 9 produce the simulation figures. The code for Figures 8 and 9 runs Simulation.do. The simulations and Figure I1, rely on the estimation of a monotonic polynomial regression which is done in Matlab via the code Simulations_monotonic_polynomial_fit.m. Simulation.do is called in the code of Figure 8 and 9. It runs all the simulations. Simulations_monotonic_polynomial_fit.m estimates the monotonic polynomial production function based on the data created in the code of figure I1. 3. Steps PPML_reg_final.do, Final_support_tables.do, and Simulations_monotonic_polynomial_fit.m can be run independently. Final_figures.do executes Simulation.do. Hence, the latter doesn’t need to be run independently.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.019 | 0.206 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.001 | 0.005 |
| Bibliometrics | 0.005 | 0.007 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.004 | 0.005 |
| Open science | 0.002 | 0.004 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.323 | 0.107 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".