Bibliographic record
Abstract
In this thesis, we develop various methods for the purpose of data denoising. We propose a method for Mean Square Error (MSE) estimation in Soft Thresholding. The MSE estimator is based on Minimum Noiseless Data Length (MNDL). Our simulation results show that this MSE estimate is a valuable comparison measure for different soft thresholding methods. Two denoising methods are proposed for analog domain: Mean Square Error EstiMation (MSEEM) which minimizes the worst case MSE estimate, and Noise Invalidation Denoising (NIDe) method which is based on the newly prosposed idea of noise signature. While MSEEM shown to be the optimum denoising method for non-sparse signals, NIDe approach outperforms the other well known denoising methods in presence of colored noise. In digital domain we address two interesting problems: 1) simultaneous denoising and quantization method, 2) denoising a digital signal in digital domain. For problem one, we propose a new method that generalizes the idea of dead zone estimation to a multi-level noise removal. An example of this method is shown for hyperspectral image denoising and compression. A digital domain denoising approach pioneers in answering the second problem with only one prior knowledge on the desired signal, that it is digital. The method provides the optimum reconstruction levels in the MSE sense. One of the critical steps of denoising process is the noise variance estimation. As a part of this thesis, we propose a novel noise variance estimation method for BayesShrink that outperforms conventional MAD-based noise variance estimation. Although BayesShrink is one of the most efficient denoising methods, no analytical analysis is available for it. Here, we study Bayes estimators for General Gaussian Distribued (GGD) data and provide the theoretical justification for BayesShrink. This study enables us to generalize the BayesShrink threshold to Generalized BayesShrink which outperforms the BayesShrink itself.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.003 | 0.001 |
| Open science | 0.002 | 0.008 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".