MétaCan
Menu
Retour à la cohorte
Enregistrement W1531547509 · doi:10.1090/dimacs/055/09

On computing the nearest neighbor interchange distance

2000· book-chapter· en· W1531547509 sur OpenAlexaboutno aff
Bhaskar Dasgupta, Xin He, Tao Jiang, Ming Li, John Tromp, Louxin Zhang

Notice bibliographique

RevueDIMACS series in discrete mathematics and theoretical computer science · 2000
Typebook-chapter
Langueen
DomaineBiochemistry, Genetics and Molecular Biology
ThématiqueGenomics and Phylogenetic Studies
Établissements canadiensnon disponible
Organismes subventionnairesnon disponible
Mots-clésk-nearest neighbors algorithmComputer scienceNearest neighbor searchArtificial intelligence

Résumé

récupéré en direct d'OpenAlex

In the practice of molecular evolution, different phylogenetic trees for the same group of species are often produced either by procedures that use diverse optimality criteria [24] or from different genes [15, 16, 17, 18, 14]. Comparing these trees to find their similarities (e.g. agreement or consensus) and dissimilarities, i.e. distance, is thus an important issue in computational molecular biology. The nearest neighbor interchange (nni) distance [29, 28, 34, 3, 6, 2, 19, 20, 23, 33, 22, 21, 26] is a natural distance metric that has been extensively studied. Despite its many appealing aspects such as simplicity and sensitivity to tree topologies, computing this distance has remained very challenging, and many algorithmic and complexity issues about computing this distance have remained unresolved. This paper studies the complexity and efficient approximation algorithms for computing the nni distance and a natural extension of this distance on weighted phylogenies. The following results answer many open questions about the nni distance posed in the literature. 1. Computing the nni distance between two labeled trees is NP-complete. This solves a 25 year old open question appearing again and again in, for example, [29, 34, 3, 6, 2, 19, 20, 23, 22, 21, 26]. 2. Computing the nni distance between two unlabeled trees is also NPcomplete. This answers an open question in [3] for which an erroneous proof appeared in [23]. 3. Biological applications motivate us to extend the nni distance to weighted phylogenies, where edge weights indicate the time-span of evolution along each edge. We present an O(n2) time approximation algorithm for computing the nni distance on weighted phylogenies with a performance ratio of 4 logn+ 4, where n is the number of leaves in the phylogenies. We also observe that the nni distance is in fact identical to the linear-cost subtree-transfer distance on unweighted phylogenies discussed in [4, 5]. Some consequences of this observation are also discussed. 1991 Mathematics Subject Classification. Primary 68Q17, 68W40; Secondary 68Q25. The results reported here also form a subset of the results that appeared in Proc. 8th Annual ACM-SIAM Symposium on Discrete Algorithms, 1997, pp. 427-436 [4]. The remaining results of the conference paper which do not appear in this paper appeared separately in Algorithmica, Vol. 25, No. 2, pp. 176-195, 1999. The first author was supported by an CGAT (Canadian Genome Analysis and Technology) grant. The second author was supported in part by CGAT and NSF grant 9205982. The third author was supported in part by NSERC Operating Grant OGP0046613 and CGAT. The fourth author was supported by NSERC Operating Grant OGP0046506 and CGAT. The fifth author was supported by an NSERC International Fellowship and CGAT. . Work done while the first author was at University of Waterloo and McMaster University, the second author was visiting at University of Waterloo, the third author was visiting University of Washington, and the fifth and the sixth authors were at University of Waterloo.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,000
score de la tête « metaresearch » (Gemma)0,000
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesÉtudes des sciences et des technologies
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Théorique ou conceptuel · Signal consensuel: Théorique ou conceptuel
GenreSignal candidat: Empirique · Signal consensuel: aucune
Score de désaccord entre enseignants0,824
Score d'incertitude au seuil0,999

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0000,000
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0000,000
Bibliométrie0,0000,000
Études des sciences et des technologies0,0000,004
Communication savante0,0000,000
Science ouverte0,0010,001
Intégrité de la recherche0,0000,000
Charge utile insuffisante (le modèle a refusé de juger)0,0000,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,009
Tête enseignante GPT0,232
Écart entre enseignants0,223 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.

Devis d'étudeThéorique ou conceptuel
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations66
Publié2000
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueDIMACS series in discrete mathematics and theoretical computer scienceMême sujetGenomics and Phylogenetic StudiesTravaux en français237 207