A parallel graph-based approach for protein sequence motif discovery
Bibliographic record
Abstract
This thesis pro','ides a parallel, graph based approach to discover conservecl regìons such as motifs ìn Protein sequences.The motif discover.vproblem has gainecl lot of significa'ce i'r biologlcal scìence o'er the past decade.Recently, r,arious approaches have been used srrccessl,lll' to discor¡er motifs.some of theur are basecl on proba- ìrilistic appr-oach and the otirers on a combinatorial approach.Tiris thesis rolìou,s a graph-based approach to solve this problern, in partìcular.using the icìea of de Bruijr gaphs The de Bruij'graph has been successfully arìoptecì i. rlìe past to solve prob- lens such as locai multiple aligrment and DNA flagment assembly.The proposed algorithni ha¡nesses the power of the de Bruìju graph to <ììscover the corse¡rred re- gior.rs in a proteiìì sequence.The sequentia.lalgorithm has z0% matcires of the rrotil.s*ith the r'fEME and 65% patter'rratches with the Gibbs motif sampler.The algo- rìthrn ìs redesigned and parallelized on the high perfbrma.ncecomputers a'ailable on the Wcstern Canada Rese¡rr"h Grid (WestGricl).Perfbrmance analysis was urade on a pure distributed memorl' luachine using o¡l¡' message passing ancì on a hybricl r¡a- chinr: using shared and clistributed access space.Experiments shou,ed that the hyìrrid i.rpÌementation runs 3 times a"s fast às the pure distributecì memorv implenentafion.continuous guidar:rce and encouragenent to develop tìris rvork.We acknou'ìerìge the partial support fiom Natural Science and Engineering R.e- search Councìl(NSERC) of Canada.\tly sincere thanks to the managenent of \\:cstgricl lbr allowing us to u¡ork on tllcir systems.Finally, a plofound gratìtrrclc goes to aÌl of urv family ntembels who proviclecl contjnuous encouragement and emotional support to acconplish thjs endeavor.lll
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.004 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.008 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".