MétaCan
Menu
Back to cohort
Record W2998981416 · doi:10.1017/s1431927619015332

Proliferation of Faulty Materials Data Analysis in the Literature

2020· article· en· W2998981416 on OpenAlexaff
Matthew R. Linford, Vincent S. Smentkowski, John T. Grant, C. R. Brundle, Peter M. A. Sherwood, Mark C. Biesinger, Jeff Terry, Kateryna Artyushkova, Alberto Herrera‐Gómez, S. Tougaard, William Skinner, Jean Jacques Strodiot, C. F. McConville, Christopher D. Easton, Thomas R. Gengenbach, George H. Major, Paul Dietrich, Andreas Thißen, Mark Engelhard, C. J. Powell, Karen J. Gaskell, Donald R. Baer

Bibliographic record

VenueMicroscopy and Microanalysis · 2020
Typearticle
Languageen
FieldEngineering
TopicMineral Processing and Grinding
Canadian institutionsWestern University
Fundersnot available
KeywordsMaterials scienceComputer scienceForensic engineeringEngineering

Abstract

fetched live from OpenAlex

As a group of subject-matter experts in X-ray photoelectron spectroscopy (XPS) and other material characterization techniques from different countries and institutions, we write this document to raise awareness of an epidemic of poor and incorrect materials data analysis in the literature. This issue is a growing problem with many causes and very undesirable consequences. It contributes to what has been called a “reproducibility crisis”, which is a recent concern of the U.S. National Academies of Science (Baker, 2016; Harris, 2017; NASE&M, 2019). Over the past decade material analysis techniques have matured to the point that dedicated expert operators are often not considered to be necessary to collect and analyze data, especially when the samples are perceived as simple or routine. The tools in this growing arsenal, including XPS, are now used in academia, industry, and government laboratories to provide both compositional information and a mechanistic understanding of a wide variety of materials. This situation, coupled with increased accessibility of the equipment, improved instrument reliability, and the promise of useful data, has resulted in significant growth in the number of researchers using these characterization tools and reporting material analysis data. Although many of the resulting papers are of high quality, especially in journals that focus on materials characterization, others are unsatisfactory. In an ongoing analysis of XPS data in journals that emphasize next generation materials, we find that about 30% of the analyses are completely incorrect (Linford and Major, 2019). Thus, for some applications, inappropriate data analysis has reached a critical stage, making it difficult for researchers lacking the relevant expertise to find and readily identify reliable examples of what would be considered good-quality data analysis. The errors we are observing in the literature are not limited to journals that may be deemed to be of lower impact—they regularly appear in what are identified as upper-tier/high-impact-factor journals. It is not uncommon to similarly find that 20–30% of the analyses of data from other material characterization techniques are also incorrect (Chirico et al., 2013; Park et al., 2017). The consequences of this issue are significantly greater than merely having a few poorly executed figures in otherwise good papers. Results and conclusions in a study hinge on the data collected and analyzed. If the characterization of a material is incorrect, an entire work may be fundamentally flawed. In some areas, the proliferation of advanced analytical instruments appears to have exceeded the world's supply of expertise necessary to collect, interpret, and review the results obtained from them. Some sub-disciplines in science only require a single analytical/measurement tool or just a few tools for a complete analysis of their systems. In contrast, materials analysis generally requires multiple advanced-characterization techniques to obtain an appropriate understanding of a new thin film or material (Baer & Gilmore, 2018). These techniques typically require an understanding of the physics and chemistry behind them, can be performed in multiple modes, and often require detailed first-principles and/or established empirical/semi-empirical modeling for their data reduction. Furthermore, each technique is supported by an extensive literature written by experts. Because of the need for information from these methods, the burden placed on materials researchers is heavy. In addition to a requirement to develop novel materials, they must characterize them at a high level with multiple analytical tools. Of course, not every materials problem requires advanced data analysis. Many important quality-control and device-failure problems have been solved by a basic application of one or more pieces of modern characterization equipment. However, in mature industries and fields, advances are more often made through the development of a detailed, comprehensive understanding of materials. In these cases, inadequate data collection and unsatisfactory analysis impede progress. This epidemic has a plethora of consequences. When too much of the literature is corrupted by poorly collected and poorly interpreted results, a resource that was designed to further the cause of research is compromised. Unfortunately, the literature does not have a generally accepted mechanism for identifying studies (or portions thereof) of questionable value, so incorrect results may influence the thinking, direction, and future research of other scientists and engineers. Anecdotally, we note that, as analysts, we are often asked to reproduce or follow a protocol from the literature that is fundamentally flawed. Incorrect precedent is sometimes cited in the literature, which perpetuates errors. Results from materials characterization influence business decisions, and graduate students and researchers who ought to be able to learn from the literature are misinformed. In our view, the fact that many incorrect analyses are appearing in the literature is a systemic problem. Researchers are under intense pressure to publish, and without some change to the system, they will most likely continue to do their work as they have in the past. Instrument manufacturers are, at least indirectly, if unintentionally, complicit; they have developed high-quality and easy-to-operate systems that may mislead customers into believing that data collection and analysis from their instruments are straightforward endeavors. This certainly may be the case for some routine samples, but not for all materials. Moderately to very complex materials such as nanoparticles, nanostructured, and two- or three-dimensional materials, catalysts, and anisotropic or graded materials, require a more nuanced approach. Reviewers and editors of manuscripts are often experts in the synthesis and/or development of a particular type of material, and in this sense are appropriately chosen to evaluate certain classes of manuscripts. However, they often do not possess a detailed understanding of all the analytical methods that may have been used to characterize the new materials described in the documents they review. Thus, the structure, traditions, constraints, and pressures of the current scientific endeavor often lead to the publication of faulty or misleading data analysis (Baer & Gilmore, 2018). A partial solution that some of us are applying to the review process is to review only portions of papers for which we have the needed expertise and to clearly inform the editors of the areas where we were not qualified to provide a needed evaluation. This is only a preliminary suggestion. For evaluation purposes it will be useful to have authors include more detailed information about characterization in the supplemental information sections allowed by many journals. As this discussion and the analysis of this problem have progressed, our consensus of a solution has become multifaceted, with emphasis on each level of stakeholder. A more detailed analysis of the current problems in XPS data analysis, along with specific actions that can be taken to address the issues, is forthcoming. Some of us are in the process of writing a series of guides, tutorials, and recommended protocol articles on XPS that are being published in the Journal of Vacuum Science and Technology (Shah et al., 2018; Baer et al., 2019). We believe that these will be an aid to those who wish to acquaint themselves with the technique, so that they can avoid some of the common pitfalls in XPS data analysis and reporting. Many of us have also been involved in developing documentary standards for XPS and other surface-analysis methods that have been published by ASTM International and the International Organization for Standardization (ISO). High-quality surface-analysis data have also been published in Surface Science Spectra. The guides and papers we are developing will include lists of common errors made in XPS data analysis, recommendations to all the stakeholders in this issue, and a more detailed, quantitative analysis of the problem. Similar guidance and reference data, for example, ASTM and ISO standards, have also been developed for other material characterization techniques. We commend the efforts of other groups of experts that similarly teach appropriate data analysis for their methods and call upon the scientific community to pay greater heed to the more accurate work-up and publication of instrumental data. We close by reiterating that while the focus of this document has been on XPS, similar trends and problems are being noted in all areas of materials characterization. Dr. C. J. Powell contributed to this communication in a personal capacity. The views expressed are his own and do not necessarily represent the views of the National Institute of Standards and Technology or the United States Government.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.051
metaresearch head score (Gemma)0.279
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.949
Threshold uncertainty score0.270

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0510.279
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0330.022
Science and technology studies0.0040.007
Scholarly communication0.0090.010
Open science0.0040.005
Research integrity0.0040.006
Insufficient payload (model declined to judge)0.0050.004

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.019
GPT teacher head0.268
Teacher spread0.248 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designObservational
DomainMethods
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations86
Published2020
Admission routes1
Has abstractyes

Explore more

Same venueMicroscopy and MicroanalysisSame topicMineral Processing and GrindingFrench-language works237,207