Download The Shape of RemiXXXes to Come: Audio Texture Synthesis with Time-frequency Scattering This article explains how to apply time–frequency scattering, a convolutional operator extracting modulations in the time–frequency domain at different rates and scales, to the re-synthesis and manipulation of audio textures. After implementing phase retrieval in the scattering network by gradient backpropagation, we introduce scale-rate DAFx, a class of audio transformations expressed in the domain of time–frequency scattering coefficients. One example of scale-rate DAFx is chirp rate inversion, which causes each sonic event to be locally reversed in time while leaving the arrow of time globally unchanged. Over the past two years, our work has led to the creation of four electroacoustic pieces: FAVN; Modulator (Scattering Transform); Experimental Palimpsest; Inspection (Maida Vale Project) and Inspection II; as well as XAllegroX (Hecker Scattering.m Sequence), a remix of Lorenzo Senni’s XAllegroX, released by Warp Records on a vinyl entitled The Shape of RemiXXXes to Come.
Download Sparse and Structured Decompositions of Audio Signals in Overcomplete Spaces We investigate the notion of “sparse decompositions” of audio signals in overcomplete spaces, ie when the number of basis functions is greater than the number of signal samples. We show that, with a low degree of overcompleteness (typically 2 or 3 times), it is possible to get good approximation of the signal that are sparse, provided that some “structural” information is taken into account, ie the localization of significant coefficients that appears to form clusters. This is illustrated with decompositions on a union of local cosines (MDCT) and discrete wavelets (DWT), that are shown to perform well on percussive signals, a class of signals that is difficult to sparsely represent on pure (local) Fourier bases. Finally, the obtained clusters of individuals atoms are shown to carry higher levels of information, such as a parametrization of partials or attacks, and this is potentially useful in an information retrieval context.
Download Statistical Measures of Early Reflections of Room Impulse Responses An impulse response of an enclosed reverberant space is composed of three basic components: the direct sound, early reflections and late reverberation. While the direct sound is a single event that can be easily identified, the division between the early reflections and late reverberation is less obvious as there is a gradual transition between the two. This paper explores two statistical measures that can aid in determining a point in time where the early reflections have transitioned into late reverberation. These metrics exploit the similarities between late reverberation and Gaussian noise that are not commonly found in early reflections. Unlike other measures, these need no prior knowledge about the rooms such as geometry or volume.
Download Digital Emulation of Distortion Effects by Wave and Phase Shaping Methods This paper will consider wave (amplitude) and phase signal shaping techniques for the digital emulation of distortion effect processing. We examine how to determine the Wave- and Phaseshaping functions with harmonic amplitude and phase data. Three distortion effects units are used to provide test data. The action of the Wave- and Phase- shaping functions derived for these effects is demonstrated with the assistance of a superresolution frequency-domain analysis technique.
Download Scalable Spectral Reflections In Conic Sections The object of this project is to present a novel digital audio effect based on a real-time, windowed block-based FFT and inverse FFT. The effect is achieved by mirroring the spectrum, producing a sound effect ranging from a purer rendition of the original, through a rougher one, to a sound unrecognisable from the original. A mirror taking the shape of a conic section is constructed between certain partials, and the modified spectrum is created by reflecting the original spectrum in this mirror. The user can select the type and continuously vary the amount of curvature, typically ‘roughening’ the input sound quite gratifyingly. We demonstrate the system with live real-time audio via microphone.
Download A hierarchical approach to automatic musical genre classification A system for the automatic classification of audio signals according to audio category is presented. The signals are recognized as speech, background noise and one of 13 musical genres. A large number of audio features are evaluated for their suitability in such a classification task, including well-known physical and perceptual features, audio descriptors defined in the MPEG-7 standard, as well as new features proposed in this work. These are selected with regard to their ability to distinguish between a given set of audio types and to their robustness to noise and bandwidth changes. In contrast to previous systems, the feature selection and the classification process itself are carried out in a hierarchical way. This is motivated by the numerous advantages of such a tree-like structure, which include easy expansion capabilities, flexibility in the design of genre-dependent features and the ability to reduce the probability of costly errors. The resulting application is evaluated with respect to classification accuracy and computational costs.
Download Spectral Analysis of Stochastic Wavetable Synthesis Dynamic Stochastic Wavetable Synthesis (DSWS) is a sound synthesis and processing technique that uses probabilistic waveform synthesis techniques invented by Iannis Xenakis as a modulation/ distortion effect applied to a wavetable oscillator. The stochastic manipulation of the wavetable provides a means to creating signals with rich, dynamic spectra. In the present work, the DSWS technique is compared to other fundamental sound synthesis techniques such as frequency modulation synthesis. Additionally, several extensions of the DSWS technique are proposed.
Download Modeling of the Carbon Microphone Nonlinearity for a Vintage Telephone Sound Effect The telephone sound effect is widely used in music, television and the film industry. This paper presents a digital model of the carbon microphone nonlinearity which can be used to produce a vintage telephone sound effect. The model is constructed based on measurements taken from a real carbon microphone. The proposed model is a modified version of the sandwich model previously used for nonlinear telephone handset modeling. Each distortion component can be modeled individually based on the desired features. The computational efficiency can be increased by lumping the spectral processing of the individual distortion components together. The model incorporates a filtered noise source to model the self-induced noise generated by the carbon microphones. The model has also an input level depended noise generator for additional sound quality degradation. The proposed model can be used in various ways in the digital modeling of the vintage telephone sound.
Download Expressive Piano Performance Rendering from Unpaired Data Recent advances in data-driven expressive performance rendering have enabled automatic models to reproduce the characteristics and the variability of human performances of musical compositions. However, these models need to be trained with aligned pairs of scores and performances and they rely notably on score-specific markings, which limits their scope of application. This work tackles the piano performance rendering task in a low-informed setting by only considering the score note information and without aligned data. The proposed model relies on an adversarial training where the basic score notes properties are modified in order to reproduce the expressive qualities contained in a dataset of real performances. First results for unaligned score-to-performance rendering are presented through a conducted listening test. While the interpretation quality is not on par with highly-supervised methods and human renditions, our method shows promising results for transferring realistic expressivity into scores.
Download GPU-Based Spectral Model Synthesis for Real-Time Sound Rendering The timbre of an instrument is usually represented by sinusoids plus noise. Spectral modeling synthesis (SMS) is an audio synthesis technique which can create musical timbre and give control over the frequency and amplitude. Additive synthesis and LPC synthesis are usually applied for synthesizing sinusoids and residuals, respectively. However, it takes fairly large computing power while implementing the algorithms. The purpose of this paper is to present GPU-based techniques of implementing SMS for real-time audio processing by using parallelism and programmability in graphics pipeline. The performance is compared to CPU-based implementations.