MétaCan
Menu
Back to cohort
Record W2340406802

Dna watson-crick complementarity in computer science

2010· article· en· W2340406802 on OpenAlexaff
Shinnosuke Seki

Bibliographic record

Venuenot available
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicDNA and Biological Computing
Canadian institutionsWestern University
Fundersnot available
KeywordsMolecular Structure of Nucleic Acids: A Structure for Deoxyribose Nucleic AcidCombinatoricsMathematicsComplementarity (molecular biology)PrefixDiscrete mathematicsDNABase pairGeneticsPhilosophyBiology
DOInot available

Abstract

fetched live from OpenAlex

Genetic information is encoded over the four nucleotide alphabet {A, C, G, T} in the form of DNA helix (double-stranded structure). This structure consists of DNA strands with opposite orientation (called Watson and Crick strands), bonded via the Watson-Crick complementarity A-T, C-G. During DNA replication, each of these strands serves as a template for the reproduction of the complementary strand so as to produce two identical copies of the original DNA helix. Thus, we can say that the Watson and Crick strands are equivalent with respect to the information they encode. The Watson-Crick complementarity is mathematically modeled as an antimorphic involution t. Hence, we can formalize the above-mentioned equivalence by the equivalence between a word and its image under t. This generalization enables us to extend the notions of periodicity and power (repetition) to those of pseudo-periodicity and pseudo-power. We call any word in u{u,t(u)}* a pseudo-power of u. With the notion of pseudo-power, we extend two problems of significance which involve power of words, that is, the Fine and Wilf's theorem and the Lyndon-Schutzenberger equation. The first theorem answers the question of how long prefix a pseudo-power of u and that of v should share to imply that u and v are pseudo-powers of some common word. Onto the length of this prefix, we provide an upper bound 2 max(|u|, |v|) + min(|u|, |v|) – gcd(|u|, |v|), and later improve it slightly. We also investigate its lower bound by constructing words u, v which cannot be written as pseudo-powers of a common word, but some of whose pseudo-powers can share a prefix of length quite close to the upper bound. The extended Lyndon-Schutzenberger equation is of the form au,qu =bv,q vg w,qw , where α(u, t(u)) ∈ {u, t(u)}e, β( v, t(v)) ∈ {v, t( v)}n, and γ(w, t( w)) ∈ {w, t(w)} m for some e, n, m ≥ 1. We ask the question of under what conditions on e, n, m, this equation implies that u, v, w ∈ {t, t(t)} + for some word t. The strongest condition we obtained so far is e ≥ 4, m, n ≥ 3. Keywords: Watson-Crick complementarity, antimorphic involution, Fine and Wilf's theorem, Lyndon-Schutzenberger equation

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.002
metaresearch head score (Gemma)0.008
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Theoretical or conceptual · Consensus signal: Theoretical or conceptual
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.004
Threshold uncertainty score0.020

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0020.008
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0020.003
Science and technology studies0.0010.005
Scholarly communication0.0030.006
Open science0.0010.002
Research integrity0.0020.005
Insufficient payload (model declined to judge)0.0040.002

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.013
GPT teacher head0.272
Teacher spread0.259 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designTheoretical or conceptual
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2010
Admission routes1
Has abstractyes

Explore more

Same topicDNA and Biological ComputingFrench-language works237,207