Tobit Estimation with Unknown Point of Censoring with an Application to Milk Market Participation in the Ethiopian Highlands
Bibliographic record
Abstract
Data augmentation is a powerful technique for estimating models with latent or missing data, but applications in agricultural economics have thus far been few. This paper showcases the technique in an application to data on milk market participation in the Ethiopian highlands. There, a key impediment to economic development is an apparently low rate of market participation. Consequently, economic interest centers on the “locations” of nonparticipants in relation to the market and their “reservation values” across covariates. These quantities are of policy interest because they provide measures of the additional inputs necessary in order for nonparticipants to enter the market. One quantity of primary interest is the minimum amount of surplus milk (the “minimum efficient scale of operations”) that the household must acquire before market participation becomes feasible. We estimate this quantity through routine application of data augmentation and Gibbs sampling applied to a random‐censored Tobit regression. Incorporating random censoring affects markedly the marketable‐surplus requirements of the household, but only slightly the covariates requirements estimates and, generally, leads to more plausible policy estimates than the estimates obtained from the zero‐censored formulation. Augmenter le nombre de données est une technique puissante permettant d'estimer les modèles pour lesquels les données manquent ou sont latentes, mais, jusqu'à présent, on n'y a guère recouru en économie rurale. L'article que void illustre une application de cette technique aux données se rapportant à la participation des producteurs au marché laitier dans les hauts plateaux de l'Éthiopie. Le faible taux de participation apparent constitue un obstacle de taille au développemenl économique de la région. Par conséquent, pour l'économiste, il est inléressanl de savoir où les « non‐participants » se trouvent par rapport au marché et de jauger leurs « réserves » en fonction des autres variables. Les valeurs de ce genre présenlenl de l'intérêt au niveau de l'élaboration des politiques, parce qu'elles donnent une idée des intrants supplémenlaires qui convaincraient les non‐participants de faire leur entrée sur le marché. Une valeur d'un intérêt primordial est le volume minimal de lait excédentaire («échelle minimale d'efficacité») que le ménage doit acquérir avant de participer au marché. On a eslimé cette valeur en appliquant la technique d'augmentation des données de la façon habituelle et en appliquant l'échanlillonnage de Gibbs à une régression Tobit censurée au hasard. L'intégration d'une censuration aléaloire affecte de maniére appréciable le volume de lait excédentaire commercialisable dont le ménage a besoin, mais ne louche que légèrement la valeur estimative des covariables requises et, en général, débouche sur des estimations plus plausibles que celles oblenues sans censuration en ce qui concerne l'élaboration des politiques.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.010 | 0.033 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".