Download Gestural exploitation of ecological information in continuous sonic feedback – The case of balancing a rolling ball
Continuous sensory–motor loops form a topic dealt with rather rarely in experiments and applications of ecological auditory perception. Experiments with a tangible audio–visual interface around a physics-based sound synthesis core address this aspect. Initially dealing with the evaluation of a specific work of sound and interaction design, they deliver new arguments and notions for non-speech auditory display and are also to be seen in a wider context of psychoacoustic knowledge and methodology.
Download A Pitch Salience Function Derived from Harmonic Frequency Deviations for Polyphonic Music Analysis
In this paper, a novel approach for the computation of a pitch salience function is presented. The aim of a pitch (considered here as synonym for fundamental frequency) salience function is to estimate the relevance of the most salient musical pitches that are present in a certain audio excerpt. Such a function is used in numerous Music Information Retrieval (MIR) tasks such as pitch, multiple-pitch estimation, melody extraction and audio features computation (such as chroma or Pitch Class Profiles). In order to compute the salience of a pitch candidate f , the classical approach uses the weighted sum of the energy of the short time spectrum at its integer multiples frequencies hf . In the present work, we propose a different approach which does not rely on energy but only on frequency location. For this, we first estimate the peaks of the short time spectrum. From the frequency location of these peaks, we evaluate the likelihood that each peak is an harmonic of a given fundamental frequency. The specificity of our method is to use as likelihood the deviation of the harmonic frequency locations from the pitch locations of the equal tempered scale. This is used to create a theoretical sequence of deviations which is then compared to an observed one. The proposed method is then evaluated for a task of multiple-pitch estimation using the MAPS test-set.
Download Diffusion-Based Music Audio Editing System Using Differentiable Digital Signal Processing Mixture Model
This paper proposes a music audio editing system that enables source-wise editing of harmonic instrument mixtures without explicit source separation. It builds on our previously proposed score-informed method for estimating source-wise synthesis parameters, i.e., time-varying controls used to synthesize each source, such as fundamental frequency and loudness. The method directly estimates these parameters from a mixture signal and the corresponding musical score in an analysis-by-synthesis framework. Using the estimated parameters, the proposed system allows users to edit individual sources by modifying note sequences and instrument types, and then re-synthesizes the edited mixture. Through demonstrations on two-instrument mixtures, we show that the system supports note-level phrasing modification and instrument conversion of selected sources.
Download Adaptive Modeling of Synthetic Nonstationary Sinusoids
Nonstationary oscillations are ubiquitous in music and speech, ranging from the fast transients in the attack of musical instruments and consonants to amplitude and frequency modulations in expressive variations present in vibrato and prosodic contours. Modeling nonstationary oscillations with sinusoids remains one of the most challenging problems in signal processing because the fit also depends on the nature of the underlying sinusoidal model. For example, frequency modulated sinusoids are more appropriate to model vibrato than fast transitions. In this paper, we propose to model nonstationary oscillations with adaptive sinusoids from the extended adaptive quasi-harmonic model (eaQHM). We generated synthetic nonstationary sinusoids with different amplitude and frequency modulations and compared the modeling performance of adaptive sinusoids estimated with eaQHM, exponentially damped sinusoids estimated with ESPRIT, and log-linear-amplitude quadratic-phase sinusoids estimated with frequency reassignment. The adaptive sinusoids from eaQHM outperformed frequency reassignment for all nonstationary sinusoids tested and presented performance comparable to exponentially damped sinusoids.
Download Taming the Red Llama—Modeling a CMOS-Based Overdrive Circuit
The Red Llama guitar overdrive effect pedal differs from most other overdrive effects because it utilizes CMOS inverters, formed by two metal-oxide-semiconductor field-effect transistors (MOSFETs), instead of a combination of operational amplifiers and diodes to obtain nonlinear distortion. This makes it an interesting subject for virtual analog modeling, obviously requiring a suitable model for the CMOS inverters. Therefore, in this paper, we extend a well-known model for MOSFETs by a straight-forward heuristic approach to achieve a good match between the model and measurement data obtained for the individual MOSFETs. This allows a faithful digital simulation of the Red Llama.
Download A Minimal Passive Model of the Operational Amplifier: Application to Sallen-Key Analog Filters
This papers stems from the fact that, whereas there are passive models of transistors and tubes, a minimal passive model of the operational amplifier does not seem to exist. A new behavioural model is presented that is memoryless, fully described by its interaction ports, with a minimal number of equations, for which a passive power balance can be defined. The proposed model handles saturation, asymmetric power supply, and can be used with nonideal voltage references. To illustrate the model in audio applications, the non-inverting voltage amplifier and a saturating Sallen-Key lowpass filter are considered.
Download Continuous and discrete Fourier spectra of aperiodic sequences for sound modeling
The Fourier analysis of aperiodic ordered time structures related with number eight is considered. Recursion relations for the Fourier amplitudes are obtained for a sequence with discrete spectrum. The continuous spectrum of a different type of sequence is also studied . By increasing the number of points in the time axis dynamic spectra can be obtained and used for sound synthesis.
Download Improving intelligibility prediction under informational masking using an auditory saliency model
The reduction of speech intelligibility in noise is usually dominated by energetic masking (EM) and informational masking (IM). Most state-of-the-art objective intelligibility measures (OIM) estimate intelligibility by quantifying EM. Few measures model the effect of IM in detail. In this study, an auditory saliency model, which intends to measure the probability of the sources obtaining auditory attention in a bottom-up process, was integrated into an OIM for improving the performance of intelligibility prediction under IM. While EM is accounted for by the original OIM, IM is assumed to arise from the listener’s attention switching between the target and competing sounds existing in the auditory scene. The performance of the proposed method was evaluated along with three reference OIMs by comparing the model predictions to the listener word recognition rates, for different noise maskers, some of which introduce IM. The results shows that the predictive accuracy of the proposed method is as good as the best reported in the literature. The proposed method, however, provides a physiologically-plausible possibility for both IM and EM modelling.
Download Constrained Pole Optimization for Modal Reverberation
The problem of designing a modal reverberator to match a measured room impulse response is considered. The modal reverberator architecture expresses a room impulse response as a parallel combination of resonant filters, with the pole locations determined by the room resonances and decay rates, and the zeros by the source and listener positions. Our method first estimates the pole positions in a frequency-domain process involving a series of constrained pole position optimizations in overlapping frequency bands. With the pole locations in hand, the zeros are fit to the measured impulse response using least squares. Example optimizations for a mediumsized room show a good match between the measured and modeled room responses.
Download AudioBIFS: The MPEG-4 Standard for Effects Processing
We present a tutorial overview of the AudioBIFS system, part of the Binary Format for Scene Description in the MPEG-4 International Standard. AudioBIFS allows the flexible construction of sound scenes using streaming audio, interactive presentation, 3-D spatialization and environmental auralization, and dynamic download of custom signal-processing routines. MPEG-4 sound scenes are based on a model that is a superset of the model in VRML 2.0, and a comparison between the two models is presented. We discuss the use of SAOL, the MPEG-4 Structured Audio Orchestra Language, for writing downloadable effects. The current status of the standard is described.