Improved hidden Markov model partial tracking for additive synthesis using time-frequency analysis
Bibliographic record
Abstract
Additive synthesis models are popular for musical applications because they offer precise control over the temporal-spectral evolution of sound. The main difficulty in additive modelling lies in the estimation and tracking of model parameters for sounds with time-varying frequency content (this is especially common in music). In the work presented here, a modification is proposed to the hidden Markov model developed by Depalle et al. which improves the estimation and tracking of partial frequencies for additive synthesis. The Wigner-Ville distribution is used in conjunction with the Hough transform in order to re-estimate the frequency and chirp rate from the short-term spectrum. These estimates are then used to formulate a new objective function to score state transitions in the hidden Markov model. The system performance is evaluated using a suite of real and synthetic test signals, and compared to state-of-the-art analysis tools. A clear improvement is demonstrated for highly non-stationary monophonic and polyphonic sounds. Additionally, crossing partials are detected to a limited extent, although further work is still needed to address this problem. This work identifies many strategies to improve additive modelling and suggests several directions for future research.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".