Download A jump start for NMF with N-FINDR and NNLS
Nonnegative Matrix Factorization is a popular tool for the analysis of audio spectrograms. It is usually initialized with random data, after which it iteratively converges to a local optimum. In this paper we show that N-FINDR and NNLS, popular techniques for dictionary and activation matrix learning in remote sensing, prove useful to create a better starting point for NMF. This reduces the number of iterations necessary to come to a decomposition of similar quality. Adapting algorithms from the hyperspectral image unmixing and remote sensing communities, provides an interesting direction for future research in audio spectrogram factorization.
Download Simulating Microphone Bleed and Tom-tom Resonance in Multisampled Drum Workstations
In recent years multisampled drum workstations have become increasingly popular. They offer an alternative to recording a full drum kit if a producer, engineer or amateur lacks the equipment, money, space or knowledge to produce a quality recording. These drum workstations strive for realism, often recording up to a hundred different velocity hits of the same drum, including recordings from all microphones for each drum hit and including bleed between these microphones. This paper describes research undertaken to investigate if it is possible to simulate the snare and kick drum bleed into the tom-tom microphones and the subsequent resonance of the tom-tom that is caused, with the aim of reducing the amount of audio data that needs to be stored. A listening test was performed asking participants to identify the real recording from a simulation. The results were not statistically significant to reject the hypothesis that subjects were unable to distinguish the difference between the real and simulated recordings. This suggests listeners were unable to identify the real recording in the majority of cases.
Download The wave digital reed: A passive formulation
In this short paper, we address the numerical simulation of the single reed excitation mechanism. In particular, we discuss a formalism for approaching the lumped nonlinearity inherent in such a model using a circuit model and the application of wave digital filters (WDFs), which are of interest in that they allow simple stability verification, a property which is not generally guaranteed if one employs straightforward numerical methods. We present first a standard reed model, then its circuit representation, then finally the associated wave digital network. We then enter into some implementation issues, such as the solution of nonlinear algebraic equations, and the removal of delay-free loops, and present simulation results.
Download Spherical Decomposition of Arbitrary Scattering Geometries for Virtual Acoustic Environments
A method is proposed to encode the acoustic scattering of objects for virtual acoustic applications through a multiple-input and multiple-output framework. The scattering is encoded as a matrix in the spherical harmonic domain, and can be re-used and manipulated (rotated, scaled and translated) to synthesize various sound scenes. The proposed method is applied and validated using Boundary Element Method simulations which shows accurate results between references and synthesis. The method is compatible with existing frameworks such as Ambisonics and image source methods.
Download Stationary/transient Audio Separation Using Convolutional Autoencoders
Extraction of stationary and transient components from audio has many potential applications to audio effects for audio content production. In this paper we explore stationary/transient separation using convolutional autoencoders. We propose two novel unsupervised algorithms for individual and and joint separation. We describe our implementation and show examples. Our results show promise for the use of convolutional autoencoders in the extraction of sparse components from audio spectrograms, particularly using monophonic sounds.
Download High-level musical control paradigms for Digital Signal Processing
No matter how complex DSP algorithms are and how rich sonic processes they produce, the issue of their control immediately arises when they are used by musicians, independently on their knowledge of the underlying mathematics or their degree of familiarity with the design of digital instruments. This text will analyze the problem of the control of DSP modules from a compositional standpoint. An implementation of some paradigms in a Lisp-based environment (omChroma) will also be concisely discussed.
Download Pinna Morphological Parameters Influencing HRTF Sets
Head-Related Transfer Functions (HRTFs) are one of the main aspects of binaural rendering. By definition, these functions express the deep linkage that exists between hearing and morphology especially of the torso, head and ears. Although the perceptive effects of HRTFs is undeniable, the exact influence of the human morphology is still unclear. Its reduction into few anthropometric measurements have led to numerous studies aiming at establishing a ranking of these parameters. However, no consensus has yet been set. In this paper, we study the influence of the anthropometric measurements of the ear, as defined by the CIPIC database, on the HRTFs. This is done through the computation of HRTFs by Fast Multipole Boundary Element Method (FM-BEM) from a parametric model of torso, head and ears. Their variations are measured with 4 different spectral metrics over 4 frequency bands spanning from 0 to 16kHz. Our contribution is the establishment of a ranking of the selected parameters and a comparison to what has already been obtained by the community. Additionally, a discussion over the relevance of each approach is conducted, especially when it relies on the CIPIC data, as well as a discussion over the CIPIC database limitations.
Download Wave Field Synthesis - Generation and Reproduction of Natural Sound Environments
Since the early days of stereo good spatial sound impression had been limited to a small region, the so-called sweet spot. About 15 years ago the concept of wave field synthesis (WFS) solving this problem has been invented at TU Delft, but due to its computational complexity it has not been used outside universities and research institutes. Today the progress of microelectronics makes a variety of applications of WFS possible, like themed environments, cinemas, and exhibition spaces. This paper will highlight the basics of WFS and discuss some of the solutions beyond the basics to make it work in applications.
Download Extensions and Applications of Modal Dispersive Filters
Dispersive delay and comb filters, implemented as a parallel sum of high-Q mode filters tuned to provide a desired frequency-dependent delay characteristic, have advantages over dispersive filters that are implemented using cascade or frequency-domain architectures. Here we present techniques for designing the modal filter parameters for music and audio applications. Through examples, we show that this parallel structure is conducive to interactive and time-varying modifications, and we introduce extensions to the basic model.
Download Flexible Real-Time Reverberation Synthesis With Accurate Parameter Control
Reverberation is one of the most important effects used in audio production. Although nowadays numerous real-time implementations of artificial reverberation algorithms are available, many of them depend on a database of recorded or pre-synthesized room impulse responses, which are convolved with the input signal. Implementations that use an algorithmic approach are more flexible but do not let the users have full control over the produced sound, allowing only a few selected parameters to be altered. The realtime implementation of an artificial reverberation synthesizer presented in this study introduces an audio plugin based on a feedback delay network (FDN), which lets the user have full and detailed insight into the produced reverb. It allows for control of reverberation time in ten octave bands, simultaneously allowing adjusting the feedback matrix type and delay-line lengths. The proposed plugin explores various FDN setups, showing that the lowest useful order for high-quality sound is 16, and that in the case of a Householder matrix the implementation strongly affects the resulting reverberation. Experimenting with delay lengths and distribution demonstrates that choosing too wide or too narrow a length range is disadvantageous to the synthesized sound quality. The study also discusses CPU usage for different FDN orders and plugin states.