Bibliographic record
Abstract
1. Abstract In this paper I preliminarily examine Genome Canada's data release policy, (1) and patenting as a form of data release. Through a summary of several interesting case studies, I illustrate a sample of data release approaches that are currently in use by various consortia. I note that Genome Canada through their policy continues to try and strike a balance between the values of science and commercialization. For instance, data release policies in general have different impacts, one being the active free flow of scientific information. As well, there are a range of concerns about publicly disclosing such information and how effective these measures actually are when we think about long term goals surrounding translation and commercialization. 2. Introduction The MORGEN (2) G[E.sup.3]LS team is looking at the implications of data release policies in general, as well as specifically upon the products of genomic research, for the MORGEN project. The MORGEN project includes a study of the regulation of gene expression and how it specifies organogenesis, through studying development of the heart, liver and pancreas in the mouse embryo. This project follows on a previous project, the Mouse Atlas of Gene Expression, that deposited research results to a website. (3) Many such policies have emphasized the need for specific tailoring of data release for areas such as gene expression analysis. The Genome Canada data release policy specifically applicable to the MORGEN project notes a range of potential avenues for data release including: publishing; patents; researcher's publicly and freely accessible sources (website); data archives; open source archives; and depositing into strain collections. In order to better understand the alternatives provided by Genome Canada, we are also considering several policies set forth by other funding agencies. 3. Patenting as a form of Data Release Patenting is listed as a form of data release. The Genome Canada data release policy states that data must be released no later than the date that the patent (including provisional patents) has been applied for, and there is a provision to apply for delays beyond this. As such, patenting as a form of data release is actually keyed toward the actual application/provisional stage rather than the stage of issued patent. In this respect, there is a difference between the Genome Canada policy and the policy of similar funding agencies. In Genome Canada's policy, you can patent if you want, and there are delay possibilities built in to let you delay the release of data beyond the time of patent application. As well, if you are a proponent of the open science ideals, you can simply release to the public domain. In contrast to Genome Canada, the NIH Data Sharing Policy (4) does not call patenting a form of data release and puts clear limits on delays. Several timely examples of data release that our group is in the process of developing further are summarized in the following. 4. Potential Case Studies: Patenting & Data Release MORGEN: One form of data release used by the MORGEN project is depositing data to a publicly available website. MORGEN's website uses a Creative Commons (5) attribution 2.5 license. This license only provides copyright protection and does not protect any patent rights. Before disclosing to a publicly available website without first filing a provisional or a full patent, a number of issues surrounding disclosures and patentability must be considered. SARS: The patenting of the SARS virus has arguably been held out as an example of preserving access to genetic materials via the patent system. Provisional patents were filed by University of Hong Kong, the U.S. Centers for Disease Control and Prevention, and the British Columbia Cancer Agency [BCCA]. (6) In the context of BCCA's situation, a provisional patent was filed and data was also released to a publicly available website. …
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".