Datasets from the work: "Functional single nucleotide polymorphisms in CACNA2D3 and other autophagy-related genes are associated with leprosy among Brazilians."
Bibliographic record
Abstract
The data deposited here come from the work entitled "Functional single nucleotide polymorphisms in CACNA2D3 and other autophagy-related genes are associated with leprosy among Brazilians." by Isabela Espasandin1*, Cynthia Chester Cardoso2, Thyago Leal-Calvo3, Mayara Abud Mendes1, Ana Carla Pereira Latini4, Samuel Henrique Malcher de Castro5, André Luiz Leturiondo5, Ohanna C. L. Bezerra6, Roberta Olmo Pinheiro1, Anna Maria Sales1, Milton Ozório Moraes1†¶ and Fernanda de Souza Gomes Kehdy1¶ 1 Leprosy Laboratory, Oswaldo Cruz Institute/FIOCRUZ, Rio de Janeiro, Rio de Janeiro, Brazil2 Molecular Virology Laboratory, Department of Genetics, Federal University of Rio de Janeiro, Rio de Janeiro, Rio de Janeiro, Brazil3 University of California Berkeley, Innovative Genomics Institute, Berkeley, California, USA4 Research and Teaching Division, Lauro de Souza Lima Institute, Bauru, São Paulo, Brazil5 Laboratory of Molecular Biology, Alfredo da Matta Hospital Foundation, Manaus, Amazonas, Brazil6 Dalla Lana School of Public Health, University of Toronto, Toronto, Ontario, Canada* Corresponding authorEmail: isabela.espasandin@gmail.com† Deceased¶ These authors contributed equally to this work and are sharing the last authorship. The zipped folder contains 5 files:- the dataset used in the case-control association analysis for the population of Rio de Janeiro, containing all the variables mentioned in the article (genotypes_dataset_riodejaneiro.csv);- the dataset used in the case-control association analysis for the population of Manaus, containing all the variables mentioned in the article (genotypes_dataset_manaus.csv);- the dataset used in the case-control association analysis for the population of Rondonopolis, containing all the variables mentioned in the article (genotypes_dataset_rondonopolis.csv);- the dataset used for the expression analyses involving skin biopsies, as described and containing all the variables mentioned in the article (expression_RNAseq_skinbiopsies_data.xlsx);- the dataset used for the expression analyses involving whole blood samples (paxgene), as described and containing all the variables mentioned in the article (expression_qPCR_paxgene_totalblood_data.xlsx);- the dataset used for the expression analyses involving macrophages before and after infection with live M. leprae, as described and containing all the variables mentioned in the article (expression_macrophages_CACNA2D3_data.xlsx);
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.020 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.003 | 0.006 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.003 | 0.004 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.054 | 0.019 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".