MétaCan
Menu
Back to cohort
Record W1531547509 · doi:10.1090/dimacs/055/09

On computing the nearest neighbor interchange distance

2000· book-chapter· en· W1531547509 on OpenAlexaboutno aff
Bhaskar Dasgupta, Xin He, Tao Jiang, Ming Li, John Tromp, Louxin Zhang

Bibliographic record

VenueDIMACS series in discrete mathematics and theoretical computer science · 2000
Typebook-chapter
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicGenomics and Phylogenetic Studies
Canadian institutionsnot available
Fundersnot available
Keywordsk-nearest neighbors algorithmComputer scienceNearest neighbor searchArtificial intelligence

Abstract

fetched live from OpenAlex

In the practice of molecular evolution, different phylogenetic trees for the same group of species are often produced either by procedures that use diverse optimality criteria [24] or from different genes [15, 16, 17, 18, 14]. Comparing these trees to find their similarities (e.g. agreement or consensus) and dissimilarities, i.e. distance, is thus an important issue in computational molecular biology. The nearest neighbor interchange (nni) distance [29, 28, 34, 3, 6, 2, 19, 20, 23, 33, 22, 21, 26] is a natural distance metric that has been extensively studied. Despite its many appealing aspects such as simplicity and sensitivity to tree topologies, computing this distance has remained very challenging, and many algorithmic and complexity issues about computing this distance have remained unresolved. This paper studies the complexity and efficient approximation algorithms for computing the nni distance and a natural extension of this distance on weighted phylogenies. The following results answer many open questions about the nni distance posed in the literature. 1. Computing the nni distance between two labeled trees is NP-complete. This solves a 25 year old open question appearing again and again in, for example, [29, 34, 3, 6, 2, 19, 20, 23, 22, 21, 26]. 2. Computing the nni distance between two unlabeled trees is also NPcomplete. This answers an open question in [3] for which an erroneous proof appeared in [23]. 3. Biological applications motivate us to extend the nni distance to weighted phylogenies, where edge weights indicate the time-span of evolution along each edge. We present an O(n2) time approximation algorithm for computing the nni distance on weighted phylogenies with a performance ratio of 4 logn+ 4, where n is the number of leaves in the phylogenies. We also observe that the nni distance is in fact identical to the linear-cost subtree-transfer distance on unweighted phylogenies discussed in [4, 5]. Some consequences of this observation are also discussed. 1991 Mathematics Subject Classification. Primary 68Q17, 68W40; Secondary 68Q25. The results reported here also form a subset of the results that appeared in Proc. 8th Annual ACM-SIAM Symposium on Discrete Algorithms, 1997, pp. 427-436 [4]. The remaining results of the conference paper which do not appear in this paper appeared separately in Algorithmica, Vol. 25, No. 2, pp. 176-195, 1999. The first author was supported by an CGAT (Canadian Genome Analysis and Technology) grant. The second author was supported in part by CGAT and NSF grant 9205982. The third author was supported in part by NSERC Operating Grant OGP0046613 and CGAT. The fourth author was supported by NSERC Operating Grant OGP0046506 and CGAT. The fifth author was supported by an NSERC International Fellowship and CGAT. . Work done while the first author was at University of Waterloo and McMaster University, the second author was visiting at University of Waterloo, the third author was visiting University of Washington, and the fifth and the sixth authors were at University of Waterloo.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesScience and technology studies
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Theoretical or conceptual · Consensus signal: Theoretical or conceptual
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.824
Threshold uncertainty score0.999

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.004
Scholarly communication0.0000.000
Open science0.0010.001
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.009
GPT teacher head0.232
Teacher spread0.223 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designTheoretical or conceptual
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations66
Published2000
Admission routes1
Has abstractyes

Explore more

Same venueDIMACS series in discrete mathematics and theoretical computer scienceSame topicGenomics and Phylogenetic StudiesFrench-language works237,207