Download Feature design for the classification of audio effect units by input/output measurements
Virtual analog modeling is an important field of digital audio signal processing. It allows to recreate the tonal characteristics of real-world sound sources or to impress the specific sound of a certain analog device upon a digital signal on a software basis. Automatic virtual analog modeling using black-box system identification based on input/output (I/O) measurements is an emerging approach, which can be greatly enhanced by specific pre-processing methods suggesting the best-fitting model to be optimized in the actual identification process. In this work, several features based on specific test signals are presented allowing to categorize instrument effect units into classes of effects, like distortion, compression, modulations and similar categories. The categorization of analog effect units is especially challenging due to the wide variety of these effects. For each device, I/O measurements are performed and a set of features is calculated to allow the classification. The features are computed for several effect units to evaluate their applicability using a basic classifier based on pattern matching.
Download Hierarchical Organization and Visualization of Drum Sample Libraries
Drum samples are an important ingredient for many styles of music. Large libraries of drum sounds are readily available. However, their value is limited by the ways in which users can explore them to retrieve sounds. Available organization schemes rely on cumbersome manual classification. In this paper, we present a new approach for automatically structuring and visualizing large sample libraries through audio signal analysis. In particular, we present a hierarchical user interface for efficient exploration and retrieval based on a computational model of similarity and self-organizing maps.
Download Vivos Voco: A survey of recent research on voice transformations at IRCAM
IRCAM has a long experience in analysis, synthesis and transformation of voice. Natural voice transformations are of great interest for many applications and can be combine with text-to-speech system, leading to a powerful creation tool. We present research conducted at IRCAM on voice transformations for the last few years. Transformations can be achieved in a global way by modifying pitch, spectral envelope, durations etc. While it sacrifices the possibility to attain a specific target voice, the approach allows the production of new voices of a high degree of naturalness with different gender and age, modified vocal quality, or another speech style. These transformations can be applied in realtime using ircamTools TR A X.Transformation can also be done in a more specific way in order to transform a voice towards the voice of a target speaker. Finally, we present some recent research on the transformation of expressivity.
Download Software modules for HRTF based dynamic spatialisation
This paper describes the object oriented design and development of software modules intended to enhance multimedia presentations with sound sources spatialisation, and environmental effects (reverberation), allowing dynamic reconfiguration of the input sound parameters. Implementations have been carried out on a PC platform, on top of the Win32 API. The resulting modules (in fact C++ classes) have later been integrated into a working application for demonstration purposes.
Download Automatic Polyphonic Piano Note Extraction Using Fuzzy Logic in a Blackboard System
This paper presents a piano transcription system that transforms audio into MIDI format. Human knowledge and psychoacoustic models are implemented in a blackboard architecture, which allows the adding of knowledge with a top-down approach. The analysis is adapted to the information acquired. This technique is referred to as a prediction-driven approach, and it attempts to simulate the adaptation and prediction process taking place in human auditory perception. In this paper we describe the implementation of Polyphonic Note Recognition using a Fuzzy Inference System (FIS) as part of the Knowledge sources in a Blackboard system. The performance of the transcription system shows how polyphonic music transcription is still an unsolved problem, with a success of 45% according to the Dixon formula. However if we consider only the transcribed notes the success increases to 74%. Moreover, the results obtained in the paper presented in [1], show how the transcription can be used with success in a retrieval system, encouraging the authors to develop this technique for more accurate transcription results.
Download On Restoring Prematurely Truncated Sine Sweep Room Impulse Response Measurements
When measuring room impulse responses using swept sinusoids, it often occurs that the sine sweep room response recording is terminated soon after either the sine sweep ends or the long-lasting low-frequency modes fully decay. In the presence of typical acoustic background noise levels, perceivable artifacts can emerge from the process of converting such a prematurely truncated sweep response into an impulse response. In particular, a low-pass noise process with a time-varying cutoff frequency will appear in the measured room impulse response, a result of the frequency-dependent time shift applied to the sweep response to form the impulse response. Here, we detail the artifact, describe methods for restoring the impulse response measurement, and present a case study using measurements from the Berkeley Art Museum shortly before its demolition. We show that while the difficulty may be avoided using circular convolution, nonlinearities typical of loudspeakers will corrupt the room impulse response. This problem can be alleviated by stitching synthesized noise onto the end of the sweep response before converting it into an impulse response. Two noise synthesis methods are described: the first uses a filter bank to estimate the frequency-dependent measurement noise power and then filter synthesized white Gaussian noise. The second uses a linearphase filter formed by smoothing the recorded noise across perceptual bands to filter Gaussian noise. In both cases, we demonstrate that by time-extending the recording with noise similar to the recorded background noise that we can push the problem out in time such that it no longer interferes with the measured room impulse response.
Download Antialiased State Trajectory Neural Networks for Virtual Analog Modeling
In recent years, virtual analog modeling with neural networks experienced an increase in interest and popularity. Many different modeling approaches have been developed and successfully applied. In this paper we do not propose a novel model architecture, but rather address the problem of aliasing distortion introduced from nonlinearities of the modeled analog circuit. In particular, we propose to apply the general idea of antiderivative antialiasing to a state-trajectory network (STN). Applying antiderivative antialiasing to a stateful system in general leads to an integral of a multivariate function that can only be solved numerically, which is too costly for real-time application. However, an adapted STN can be trained to approximate the solution while being computationally efficient. It is shown that this approach can decrease aliasing distortion in the audioband significantly while only moderately oversampling the network in training and inference.
Download Analysis and resynthesis of quasi-harmonic sounds: an iterative filterbank approach
We employ a hybrid state-space sinusoidal model for general use in analysis-synthesis based audio transformations. This model, which has appeared previously in altered forms (e.g. [5], [8], perhaps others) combines the advantages of a source-filter model with the flexible, time-frequency based transformations of the sinusoidal model. For this paper, we specialize the parameter identification task to a class of “quasi-harmonic” sounds. The latter represent a variety of acoustic sources in which multiple, closely spaced modes cluster about principal harmonics loosely following a harmonic structure (some inharmonicity is allowed.) To estimate the sinusoidal parameters, an iterative filterbank splits the signal into subbands, one per principal harmonic. Each filter is optimally designed by a linear programming approach to be concave in the passband, monotonic in transition regions, and to specifically null out sinusoids in other subband regions. Within each subband, the constant frequencies and exponential decay rates of each mode are estimated by a Steiglitz-McBride approach, then time-varying amplitudes and phases are tracked by a Kalman filter. The instantaneous phase estimate is used to derive an average instantaneous frequency estimate; the latter averaged over all modes in the subband region updates the filter’s center frequency for the next iteration. In this way, the filterbank structure progressively adapts to the specific inharmonicity structure of the source recording. Analysissynthesis applications are demonstrated with standard (time/pitchscaling) transformation protocols, as well as some possibly novel effects facilitated by the “source-filter” aspect.
Download Practical Modeling of Bucket-Brigade Device Circuits
This paper discusses the sonic characteristics of the bucket-brigade device (BBD) and associated circuitry. BBDs are integrated circuits which produce a time-delayed version of an input signal. In order to reduce aliasing, distortion, and noise, BBDs are typically accompanied by low-pass filters and compander circuitry. Through circuit analysis and measurements, each component of the BBD system can be accurately modeled.
Download Visual representations of digital audio effects and of their control
This article gathers some reflections on how the graphical representations of sound and the graphical interfaces can help making and controlling digital audio effects. Time-frequency visual representations can help understand digital audio effects but are also a tool in themselves for these sound transformations. The basic laws for analysistransformations-resynthesis will be recalled. Graphical interfaces help for the control of effects. The choice of a relation between user control parameters and the effective values for the effect show the importance of the mapping and the graphical design. Graphical and sonic editor programs have both in common the use of plug-ins, which use the same kind of interface and deal with the same kind of approach.