Notice bibliographique
Résumé
The chromatin immunoprecipitation followed by high throughput sequencing (ChIP-seq) method, initially introduced a decade ago, is widely used by the scientific community to detect protein/DNA binding and histone modifications across the genome in various cell lines. Every experiment is prone to noise and bias, and ChIP-seq experiments are no exception. To alleviate bias, incorporation of control datasets in ChIP-seq analysis is an essential step. The controls are used to detect background signal, whilst the ChIP-seq experiment captures the true binding or histone modification signal. However, a recurrent issue is the existence of noise and bias in the controls themselves, as well as different types of bias in ChIP-seq experiments. Thus, depending on which controls are used, peak calling can produce different results (i.e., binding site positions) for the same ChIP-seq experiment. Consequently, generating "smart" controls, which model the non-signal effect for a specific ChIP-seq experiment, could enhance contrast and thus increase the reliability and reproducibility of the results. Our analysis aims to improve our understanding of ChIP-seq controls and their biases. We use unsupervised clustering and dimensionality reduction techniques to compare 160 controls for the K562 cell line in the ENCODE project, finding distincting groupings of controls which correlate to experimental characteristics. To customize a control for each ChIP-seq experiment, we use LASSO regression to fit a sparse set of controls to each of 500 ChIP-seq experiments (again, from ENCODE data for the K562 cell line). We look at how many controls are selected, which controls are used per ChIP-seq experiment, and how they are related to the different ChIP-seq experiment characteristics. Perhaps most surprisingly, we find that the LASSO models are not particularly sparse, often including half of the possible controls to model any given ChIP-seq. Cross-validation as well as testing with smaller sets of candidate controls proves that such large numbers of controls are beneficial for modeling ChIP-seq background distributions. We also observe clusters of ChIP-seq experiments that tend to rely on clusters of controls, and we look at the experimental characteristics that tend to cause a given control to be useful in modeling the background of a given ChIP-seq experiment. Through these analyses, we attempt to answer largely-unstudied questions regarding how much control data and of what types are useful in ChIP-seq analysis, and how suitable controls can be matched to ChIP-seq datasets.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,011 | 0,030 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,001 |
| Méta-épidémiologie (sens large) | 0,002 | 0,002 |
| Bibliométrie | 0,002 | 0,002 |
| Études des sciences et des technologies | 0,001 | 0,002 |
| Communication savante | 0,003 | 0,001 |
| Science ouverte | 0,002 | 0,001 |
| Intégrité de la recherche | 0,001 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,003 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».