Genotype Imputation and Reference Panel: A Systematic Evaluation
Notice bibliographique
Résumé
Abstract Here, 622 imputations were conducted with 394 customized reference panels for Han Chinese and European populations. Besides validating the fact that the imputation accuracy could always benefit from the increased panel size when the reference panel was population-specific, the results brought two new thoughts as follows. First, when the haplotype size of reference panel was fixed, the imputation accuracy of common and low-frequency variants (MAF>0.5%) decreased while the population-diversity of reference panel increased, but for rare variants (MAF<0.5%), a fraction of diversity (<20%) of panel could improve the imputation accuracy. Second, when the haplotype size of reference panel was increased with extra population-diverse samples, the imputation accuracy of common variants (MAF>5%) for European population could always benefit from the expanding sample size. But for Han Chinese population, the accuracy of all imputed variants reached the highest when reference panel contained a fraction of extra diverse sample (15%∼21%). In addition, we evaluated the existing reference panels such as the HRC and 1000G Phase3 and CONVERGE. For European population, HRC was the best reference panel. For Han Chinese population, we proposed an optimum constituent ratio for the Han Chinese imputation if researchers would like to customize their own sequenced reference panel, but a high quality and large-scale Chinese reference panel was still needed. Our findings could be generalized to the other populations with conservative genome, a tool was provided to investigate other populations of interest ( https://github.com/Abyss-bai/reference-panel-reconstruction ). Highlights (Key points) A total of 394 reference panels were designed and customized by three strategies, and large-scale genotype imputations were performed with these panels for systematic evaluation in Han Chinese and European populations. The accuracy of imputed variants reached the highest when reference panel contains a fraction of extra diverse sample (15%∼21%) for Han Chinese population, if the haplotype size of the reference panel was increased with extra samples, which is the most common cases. The imputation accuracy showed the different trends between Han Chinese and European populations. In a sense, the European genome may more diverse than Han Chinese genome by itself. Existing reference panels were not the best choice for Chinese imputation, a high quality and large-scale Chinese reference panel was still needed.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,218 | 0,348 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,001 |
| Méta-épidémiologie (sens large) | 0,006 | 0,010 |
| Bibliométrie | 0,005 | 0,008 |
| Études des sciences et des technologies | 0,001 | 0,001 |
| Communication savante | 0,003 | 0,002 |
| Science ouverte | 0,004 | 0,002 |
| Intégrité de la recherche | 0,002 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,004 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».