Download A Comparison of Extended Source-Filter Models for Musical Signal Reconstruction
Recently, we have witnessed an increasing use of the sourcefilter model in music analysis, which is achieved by integrating the source filter model into a non-negative matrix factorisation (NMF) framework or statistical models. The combination of the source-filter model and NMF framework reduces the number of free parameters needed and makes the model more flexible to extend. This paper compares four extended source-filter models: the source-filter-decay (SFD) model, the NMF with timefrequency activations (NMF-ARMA) model, the multi-excitation (ME) model and the source-filter model based on β-divergence (SFbeta model). The first two models represent the time-varying spectra by adding a loss filter and a time-varying filter, respectively. The latter two are extended by using multiple excitations and including a scale factor, respectively. The models are tested using sounds of 15 instruments from the RWC Music Database. Performance is evaluated based on the relative reconstruction error. The results show that the NMF-ARMA model outperforms other models, but uses the largest set of parameters.
Download The development of an online course in DSP eartraining
The authors present a collaborative effort on establishing an online course in DSP eartraining. The paper reports from a preliminary workshop that covered a large range of topics such as eartraining in music education, terminology for sound characterization, e-learning, automated tutoring, DSP techniques, music examples and audio programming. An initial design of the web application is presented as a rich content database with flexible views to allow customized online presentations. Technical risks have already been mitigated through prototyping.
Download Piano-SSM: Diagonal State Space Models for Efficient Midi-to-Raw Audio Synthesis
Deep State Space Models (SSMs) have shown remarkable performance in long-sequence reasoning tasks, such as raw audio classification, and audio generation. This paper introduces PianoSSM, an end-to-end deep SSM neural network architecture designed to synthesize raw piano audio directly from MIDI input. The network requires no intermediate representations or domainspecific expert knowledge, simplifying training and improving accessibility. Quantitative evaluations on the MAESTRO dataset show that Piano-SSM achieves a Multi-Scale Spectral Loss (MSSL) of 7.02 at 16kHz, outperforming DDSP-Piano v1 with a MSSL of 7.09. At 24kHz, Piano-SSM maintains competitive performance with an MSSL of 6.75, closely matching DDSP-Piano v2’s result of 6.58. Evaluations on the MAPS dataset achieve an MSSL score of 8.23, which demonstrates the generalization capability even when training with very limited data. Further analysis highlights Piano-SSM’s ability to train on high sampling-rate audio while synthesizing audio at lower sampling rates, explicitly linking performance loss to aliasing effects. Additionally, the proposed model facilitates real-time causal inference through a custom C++17 header-only implementation. Using an Intel Core i712700 processor at 4.5GHz, with single core inference, allows synthesizing one second of audio at 44.1kHz in 0.44s with a workload of 23.1GFLOPS/s and an 10.1µs input/output delay with the largest network. While the smallest network at 16kHz only needs 0.04s with 2.3GFLOP/s and 2.6µs input/output delay. These results underscore Piano-SSM’s practical utility and efficiency in real-time audio synthesis applications.
Download Timbre-Constrained Recursive Time-Varying Analysis for Musical Note Separation
Note separation in music signal processing becomes difficult when there are overlapping partials from co-existing notes produced by either the same or different musical instruments. In order to deal with this problem, it is necessary to involve certain invariant features of musical instrument sounds into the separation processing. For example, the timbre of a note of a musical instrument may be used as one possible invariant feature. In this paper, a timbre estimate is used to represent this feature such that it becomes a constraint when note separation is performed on a mixture signal. To demonstrate the proposed method, a timedependent recursive regularization analysis is employed. Spectral envelopes of different notes are estimated and a modified parameter update strategy is applied to the recursive regularization process. The experiment results show that the flaws due to the overlapping partial problem can be effectively reduced through the proposed approach.
Download Computer Synthesis Of Bird Songs And Calls
Bird songs are fascinating acoustic phenomena. In this paper, we first examine the acoustic mechanisms of sound production in birds and contrast them with their anatomy. Next, we describe the simulation of a one-mass source together with a simple transmission line model for a psittacine bird. In concluding, we discuss future areas for research.
Download Musical Sound Effects in the SAS Model
Spectral models provide general representations of sound in which many audio effects can be performed in a very natural and musically expressive way. Based on additive synthesis, these models control many sinusoidal oscillators via a huge number of model parameters which are only remotely related to musical parameters as perceived by a listener. The Structured Additive Synthesis (SAS) sound model has the flexibility of additive synthesis while addressing this problem. It consists of a complete abstraction of sounds according to only four parameters: amplitude, frequency, color, and warping. Since there is a close correspondence between the SAS model parameters and perception, the control of the audio effects gets simplified. Many effects thus become accessible not only to engineers, but also to musicians and composers. But some effects are impossible to achieve in the SAS model. In fact structuring the sound representation imposes limitations not only on the sounds that can be represented, but also on the effects that can be performed on these sounds. We demonstrate these relations between models and effects for a variety of models from temporal to SAS, going through well-known spectral models.
Download The Image-Source Reverberation Model in an N-Dimensional Space
The image method is generalized to geometries with an arbitrary number of spatial dimensions. n-dimensional (n-D) acoustics is discussed, and an algorithm for n-D room impulse response calculations is presented. Synthesized room impulse responses (RIRs) from n-D rooms are presented. RIR characteristics are discussed, and computational considerations are examined.
Download Stable Structures for Nonlinear Biquad Filters
Biquad filters are a common tool for filter design. In this writing, we develop two structures for creating biquad filters with nonlinear elements. We provide conditions for the guaranteed stability of the nonlinear filters, and derive expressions for instantaneous pole analysis. Finally, we examine example filters built with these nonlinear structures, and show how the first nonlinear structure can be used in the context of analog modelling.
Download Non-linear effects modeling for polyphonic piano transcription
Download Time Varying Frequency Warping: Results And Experiments
Dispersive tapped delay lines are attractive structures for altering the frequency content of a signal. In previous papers we showed that in the case of a homogeneous line with first order all-pass sections the signal formed by the output samples of the chain of delays at a given time is equivalent to compute the Laguerre transform of the input signal. However, most musical signals require a time-varying frequency modification in order to be properly processed. Vibrato in musical instruments or voice intonation in the case of vocal sounds may be modeled as small and slow pitch variations. Simulations of these effects require techniques for time- varying pitch and/or brightness modification that are very useful for sound processing. In our experiments the basis for time-varying frequency warping is a time-varying version of the Laguerre transformation. The corre- sponding implementation structure is obtained as a dispersive tapped delay line, where each of the frequency dependent delay element has its own phase response. Thus, time-varying warping results in a space-varying, inhomogeneous, propagation structure. We show that time-varying frequency warping may be associated to expansion over biorthogonal sets generalizing the discrete Laguerre basis. Slow time-varying characteristics lead to slowly varying parameter sequences. The corresponding sound transformation does not suffer from discontinuities typical of delay lines based on unit delays.