Improving the performance of neural networks through parallel processing in the cell broadband engine
Bibliographic record
Abstract
This Thesis focuses on the exploration of parallelization approaches for improving the performance of ANN.A main goal of this Thesis is to define the routes for the parallel computation of this problem using the multi-core Cell Broadband Engine.In particular, a new design for parallel tracing of the gradient descent algorithm showed the feasibility for efficient finding of viable solutions for the approximations of 2D non-linear functions and the predictions of ID time series by neural networks.One objective was to identify the parameters of the gradient descent algorithm which can be used for parallelization of the tasks in 2D function approximation and ID time series prediction in terms of speed and accuracy of the delivered solutions, and to obtain fast convergence to the optimal solutions.Specifically, for a 2D function approximation, the entrapment in the local minima has been addressed via parallel tracing of the converging trajectories, while verifying the optimality of the solutions.For a ID function approximation, the original task involves multiple-input-multiple-output multi-dimensional neural networks and thus is challenging for the gradient descent algorithm, posing problems of speed and convergence.In this case, the goal was set to verify the efficiency of the splitting of the multiple-steps forecasting task into several sub-tasks with various forecasting horizons in order to achieve fast and accurate forecasting solutions.The sub-tasks with various forecasting horizons extracted from the complex task would require a simpler type of multiple-input-single-output neural networks.The objective was to demonstrate the improved efficiency of such approach in parallel computing environment to reach fast and accurate solutions with the gradient descent algorithm.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".