msamDB: Towards addressing data-scarcity challenges in L-PBF additive manufacturing
Bibliographic record
Abstract
Data science techniques, particularly machine learning (ML), have proven to be valuable tools in PBF-LM research. While ML can rapidly model the large process parameter space of PBF-LM, their efficacy is dependent on large, informative and diverse training datasets. However, scarcity in the development and availability of such datasets is an on-going challenge. This work outlines the on-going progress to address this challenge through the development of a database platform, tentatively named msamDB (multi-scale additive manufacturing database). This platform, specifically created to manage PBF-LM academic research data, is a modular, extensible and scalable database that can promote data-sharing among researchers. The initial architecture of msamDB focuses on surface roughness data generated throughout the PBF-LM lifecycle. This work highlights the findings and challenges encountered in the design, implementation and pilot data population stages of msamDB. In its current stage, msamDB data spans data from approximately 30 builds, multiple research and industry studies, 3 different powder materials and a broad range of process parameters. Data has been collected from various stages such as powder characterization, build planning, process parameter selection, surface characterization, etc. In reference to surface roughness measurements, the database currently has more than 1000 data points across various surface orientations. This work represents first known effort to curate research PBF-LM data at scale for PBF-LM. The potential impact of such a database is to promote federated data for PBF-LM researchers, which allows for data-driven model development to have increased usability.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.003 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.004 | 0.003 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".