MétaCan
Menu
← Back to cohort
Record W6981062255

Development of a Scalable Machining Feature Recognition System

2023· dissertation· en· W6981062255 on OpenAlexaff

Bibliographic record

VenueUWSpace (University of Waterloo) · 2023
Typedissertation
Languageen
FieldEconomics, Econometrics and Finance
TopicHealthcare Policy and Management
Canadian institutionsUniversity of Waterloo
Fundersnot available
KeywordsOverfittingCADFeature (linguistics)Pattern recognition (psychology)CrossoverDropout (neural networks)Feature recognitionScalabilityMachining
DOInot available

Abstract

fetched live from OpenAlex

In this thesis, various pre-processing and training techniques were applied to improve \nthe performance of a model trained with an existing machining feature recognition approach \nby Yeo et al. using a smaller dataset that more effectively mimics the complexity of \nCAD models used in industry. \n \nA GUI tool was developed to tag faces in CAD models with the corresponding machining \nfeatures which would be necessary to resolve those faces. Using the encoding algorithm \noutlined by Yeo et al., a tool was developed to generate feature vectors from tagged \nCAD models. Two CAD datasets were compiled. First, a dataset of generic CAD models \nwas filtered from a larger dataset compiled by Koch et al., selecting those models which \ncould be manufactured using a 3-axis CNC machine. Second, a dataset of real-world CAD \nfiles used in CNC manufacturing was compiled from models contributed by individuals \nfrom the Unviersity of Waterloo, Hurco Inc. and Perfecto Inc. \n \nUsing the first dataset, three potential improvements to the feature recognition algo- \nrithm developed by Yeo et al. were explored: the incorporation of dropout to improve model \nstability and accuracy, the incorporation of ID3 tree pre-classification to reduce training \ntime by reducing the size of the deep learning dataset without impacting classification \naccuracy, and the incorporation of crossover data generation to improve classification ac- \ncuracy by reducing overfitting due to insufficient training data. It was determined that \nincorporating dropout improved the stability of the model and improved 5-fold cross val- \nidation accuracy. Further, it was determined that incorporating a 2-deep ID3 decision \ntree pre-classification marginally improved classification performance and was effective in \nreducing the size of deep learning training dataset. Crossover data generation did not \nimprove model performance, and so was rejected. Using the model trained on the generic \nCAD dataset, and incorporating 10% dropout and a 2-deep ID3 tree, models from the \nreal-world dataset were classified. This classifier was effective in classifying some simple \nfeatures, but had poor accuracy overall. To improve this accuracy, an incremental learning \ntechnique was applied. The generic model was re-trained using samples from the real-world \ndataset, which improved the classification accuracy of the system.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.001
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.009
Threshold uncertainty score0.029

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0000.001
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0010.001
Science and technology studies0.0000.000
Scholarly communication0.0010.002
Open science0.0030.001
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0090.008

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.040
GPT teacher head0.220
Teacher spread0.180 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2023
Admission routes1
Has abstractyes

Explore more

Same venueUWSpace (University of Waterloo)→Same topicHealthcare Policy and Management→French-language works237,207→