University College London
Bayesian modelling of music: algorithmic advances and experimental studies of shift-invariant sparse coding
Abstract
dc:description.abstractIn order to perform many signal processing tasks such as classification,<br/>pattern recognition and coding, it is helpful to specify a signal model in<br/>terms of meaningful signal structures. In general, designing such a model<br/>is complicated and for many signals it is not feasible to specify the appropriate<br/>structure. Adaptive models overcome this problem by learning<br/>structures from a set of signals. Such adaptive models need to be general<br/>enough, so that they can represent relevant structures. However, more<br/>general models often require additional constraints to guide the learning<br/>procedure.<br/><br/>In this thesis a sparse coding model is used to model time-series. Relevant<br/>features can often occur at arbitrary locations and the model has to be<br/>able to reflect this uncertainty, which is achieved using a shift-invariant<br/>sparse coding formulation. In order to learn model parameters, we use<br/>Bayesian statistical methods, however, analytic solutions to this learning<br/>problem are not available and approximations have to be introduced. In<br/>this thesis we study three approximations, one based on an analytical<br/>integral approximation and two based on Monte Carlo approximations.<br/>But even with these approximations, a solution to the learning problem<br/>is computationally too expensive for the applications under investigation.<br/>Therefore, we introduce further approximations by subset selection.<br/><br/>Music signals are highly structured time-series and offer an ideal testbed<br/>for the studied model. We show the emergence of note- and score-like features<br/>from a polyphonic piano recording and compare the results to those<br/>obtained with a different model suggested in the literature. Furthermore,<br/>we show that the model finds structures that can be assigned to an individual<br/>source in a mixture. This is shown with an example of a mixture<br/>containing guitar and vocal parts for which blind source separation can<br/>be performed based on the shift-invariant sparse coding model.
Degree
thesis:*- Name dc:type.qualificationname
- Ph.D.
- Level dc:type.qualificationlevel
- doctoral
- Grantor dc:publisher.institution
- University College London
- Year dc:date.issued
- 2006
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Blumensath, Thomas
- Advisor dc:contributor.advisor
-
- Davies, Mike