A call for transparency in data reporting
Notice bibliographique
Résumé
“It is a capital mistake to theorize before one has data. Insensibly one begins to twist facts to suit theories, instead of theories to suit facts.” ─ Sir Arthur Conan Doyle in 'A Scandal in Bohemia' The Adventures of Sherlock Holmes A physician from the early twentieth century transported in time would hardly recognize or comprehend modern medical literature. The good doctor would be used to penning her/his personal experience in eloquent sentences. An elegant turn of phrase would be used to convince readers. We have moved far beyond this. While we lament the fact that the beauty of informal language has given way to the rigid formality of modern scientific communication, our focus in this editorial is more to do with the data within. Today, tightly structured articles bristle with statistics and use complex tables and graphs to present as much of the data as possible in summary form. Educated readers can interpret these and make their own judgments on the validity and weight of the conclusion reached. The data generating these statistics is often massive involving numerous parameters in large numbers of patients. The very aim of this exercise is to reduce the 'literature' in medical literature and attempt to increase the science; however the outcome might often be the reverse. The ongoing COVID-19 pandemic has brought great scrutiny upon published medical literature. Faced with a rapidly spreading, previously un-encountered problem, never has it been more imperative to produce good quality scientific evidence on diagnosis, prevention, and intervention with great rapidity. Social media ensured wide dissemination of the findings often before any form of peer review; papers were dissected by unofficial reviewers, many biased for or against the conclusions. The desire to be first to discover new aspects of the disease often led to premature submissions, and journals hurried to publish potentially practice and life-altering data. The stately measured submission and peer review process of many weeks and months became a frenzied race over a few days. The biggest faux-pas were possibly by two leading journals publishing papers[12] based on data provided by a medical informatics company – Surgisphere. When the company could not share the data for independent review of its veracity, it ultimately led to the retraction of these articles. This caused immense damage to the very basis of scientific research as these retractions were extensively covered by the lay press, and the reputation of these leading journals suffered. Journal editors take pride in putting across medical literature in the most transparent fashion. The loss of literary style of issues of the past has been stimulated by the need for more transparent data presentation. Each type of article will have its own checklist to ensure that the structure is as complete as possible. The EQUATOR (Enhancing the QUAlity and Transparency Of health Research) network site lists all these. The International Committee of Medical Journal Editors (ICMJE) have ensured that transparency is maintained by standard reporting for their journals. They also require that all clinical trials be listed in a public registry along with its protocol to ensure transparency of methodology. The ICMJE statement on data sharing is the next step.[3] Every article submitted after July 2018 and any clinical trial after January 2019 is required to submit a data sharing plan. This explicitly states how much of the data is freely accessible and for what duration of time, as well as any supplemental information like the protocol and data analysis plan. Provision for secondary use of data is to be mentioned. Articles depending on secondary data are encouraged to collaborate or at the least, acknowledge the original work. When journals were printed exclusively on paper, data sharing would have been a distant dream. Today, all databases are maintained electronically. Ultimate transparency is provided when the actual data recorded for each and every patient is available so it is easy to see what was collected and what is missing. Any flaws in the presentation of the truth in the summary statistics and by the statistical analysis technique would be exposed. Further, data fabrication would be a particular challenge if the full database was exposed for all to see. The presence of raw data would facilitate meta-analysis. The Surgisphere data was put into doubt post-publication by researchers who questioned the number of cases in Australia in the time period reported and the drug availability in various countries. These are specific facts which may not have been available to the reviewers of the article. The retraction was precipitated by authors not being able to independently verify the data and analysis. This is an extreme case where even the first author had no transparent access to the raw data. Implementation of data transparency would have obviated this. If all the recorded parameters in the multitude of papers published in the pandemic were accessible to all, it would be easy to pool data and arrive at conclusions more rapidly. Similarly, it would be easy to discount poorly collected, incomplete data where confounders were ignored. The need for complete transparency would ensure that appropriate and ethical utilization of resources could be verified; matching with the prospectively published methodology would also ensure against selective reporting. Open scrutiny would also tighten data collection and presentation, and premature publication may be prevented. Would there be harm in declaring the data in a transparent manner? Clearly, if there were parameters recorded where patient identity could be revealed, this would be a problem. However, masking these details in a database is now quite easy. Certain commercial entities would worry about intellectual property and the revelation of 'trade secrets'. Similarly, in academic research, the lack of exclusivity for the primary researchers to mine the data further might disincentivize original research. However, in its purest sense, science should be open and transparent for the greater benefit of all. Privately owned pharmaceutical firms like Novartis, Merck, and AstraZeneca have their own policy encouraging data transparency. If research has been done with public funds, the public would have the absolute right to access the data so collected, and therefore, there can be no objection to data transparency. Several organizations including the Wellcome Trust, the Centers for Disease Control and Prevention, and the Canadian Institutes of Health Research actively encourage the use of common data repositories for researchers to share data. In summary, the benefits of transparency and sharing of data in the public domain far outweigh the perceived drawbacks. Modern science and researchers should embrace the philosophy of open data for the larger good of scientific progress and humanity itself. Not only should research be done well, it should also be seen to be done well. “In God we trust.. all others need to bring data.” - WE Deming (attribution disputed)
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,009 | 0,202 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,001 | 0,000 |
| Intégrité de la recherche | 0,001 | 0,009 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».