Download Physics-Inspired Feature Fusion for Plate Parameter Estimation from Acoustic Impulse Responses ★ Estimating physical plate parameters from impulse responses is a challenging inverse problem. Task A of the first Digital Audio Effects Parameter Estimation Challenge requires the recovery of six identifiable parameters from displacement impulse responses. In this work, we propose a physics-inspired feature fusion network (PIFFN) that combines a pretrained convolutional backbone with a 15-dimensional physics-inspired feature vector computed from the impulse response. These physics-inspired features describe amplitude scale, temporal decay, and spectral structure without relying on modal-distribution priors. The proposed model is evaluated on the official validation set, achieving an overall normalized mean squared error of 0.00362. Compared with the official particle swarm optimization baseline and backbone-only model, PIFFN shows a clear performance improvement, demonstrating its effectiveness for plate parameter estimation.
Download Melody Line Detection and Source Separation in classical Saxophone Recordings We propose a system which separates saxophone melodies from composite recordings of saxophone, piano, and/or orchestra. The system is intended to produce an accompaniment sans saxophone suitable for rehearsal and practice purposes. A Melody Line Detection (MLD) algorithm is proposed as the starting point for a source separation implementation which incorporates known information about typical saxophone melody lines, acoustic characteristics and range of the saxophone in order to prevent and correct detection errors. By extracting reliable information about the soloist melody line, the system separates piano or orchestra accompaniments from the solo part. The system was tested with commercial recordings and a performance of 79.7% of accurate detections was achieved. The accompaniment tracks obtained after source separation successfully remove most of the saxophone sound while preserving the original nature of the accompaniment track.
Download B-Format Acoustic Impulse Response Measurement and Analysis In the Forest at Koli National Park, Finland Acoustic impulse responses are used for convolution based auralisation and reverberation techniques for a range of applications, such as music production, sound design and virtual reality systems. These impulse responses can be measured in real world environments to provide realistic and natural sounding reverberation effects. Analysis of this data can also provide useful information about the acoustic characteristics of a particular space. Currently, impulse responses recorded in outdoor conditions are not widely available for surround sound auralisation and research purposes. This work presents results from a recent acoustic survey of measurements at three locations in the snow covered forest of Koli National Park in Finland during early spring. Acoustic impulse responses were measured using a B-format Soundfield microphone and a single loudspeaker. The results are analysed in terms of reverberation and spatial characteristics. The work is part of a larger study to collect and investigate acoustic impulse responses from a variety of outdoor locations under different climatic conditions.
Download High Accuracy Frame-by-Frame Non-Stationary Sinusoidal Modelling This paper describes techniques for obtaining high accuracy estimates, including those of non-stationarity, of parameters for sinusoidal modelling using a single frame of analysis data. In this case the data used is generated from the time and frequency reassigned short-time Fourier transform (STFT). Such a system offers the potential for quasi real-time (frame-by-frame) spectral modelling of audio signals.
Download Adjusting the Spectral Envelope Evolution of Transposed Sounds with Gabor Mask Prototypes Audio-samplers often require to modify the pitch of recorded sounds in order to generate scales or chords. This article tackles the use of Gabor masks and their capacity to improve the perceptual realism of transposed notes obtained through the classical phasevocoder algorithm. Gabor masks can be seen as operators that allows the modification of time-dependent spectral content of sounds by modifying their time-frequency representation. The goal here is to restore a distribution of energy that is more in line with the physics of the structure that generated the original sound. The Gabor mask is elaborated using an estimation of the spectral envelope evolution in the time-frequency plane, and then applied to the modified Gabor transform. This operation turns the modified Gabor transform into another one which respects the estimated spectral envelope evolution, and therefore leads to a note that is more perceptually convincing.
Download Audio Transport: A Generalized Portamento via Optimal Transport This paper proposes a new method to interpolate between two audio signals. As an interpolation parameter is changed, the pitches in one signal slide to the pitches in the other, producing a portamento, or musical glide. The assignment of pitches in one sound to pitches in the other is accomplished by solving a 1-dimensional optimal transport problem. In addition, we introduce several techniques that preserve the audio fidelity over this highly nonlinear transformation. A portamento is a natural way for a musician to transition between notes, but traditionally it has only been possible for instruments with a continuously variable pitch like the human voice or the violin. Audio transport extends the portamento to any instrument, even polyphonic ones. Moreover, the effect can be used to transition between different instruments, groups of instruments, or any other pair of audio signals. The audio transport effect operates in real-time; we provide an open-source implementation. In experiments with sinusoidal inputs, the interpolating effect is indistinguishable from ideal sine sweeps. More generally, the effect produces clear, musical results for a wide variety of inputs.
Download Audio Content Transmission Content description has become a topic of interest for many researchers in the audiovisual field [1][2]. While manual annotation has been used for many years in different applications, the focus now is on finding automatic contentextraction and content-navigation tools. An increasing number of projects, in some of which we are actively involved, focus on the extraction of meaningful features from an audio signal. Meanwhile, standards like MPEG7 [3] are trying to find a convenient way of describing audiovisual content. Nevertheless, content description is usually thought of as an additional information stream attached to the ‘actual content’ and the only envisioned scenario is that of a search and retrieval framework. However, in this article it will be argued that if there is a suitable content description, the actual content itself may no longer be needed and we can concentrate on transmitting only its description. Thus, the receiver should be able to interpret the information that, in the form of metadata, is available at its inputs, and synthesize new content relying only on this description. It is possibly in the music field where this last step has been further developed, and that fact allows us to think of such a transmission scheme being available on the near future.
Download An Extension for Source Separation Techniques Avoiding Beats The problem of separating individual sound sources from a mixture of these, known as Source Separation or Computational Auditory Scene Analysis (CASA), has become popular in the recent decades. A number of methods have emerged from the study of this problem, some of which perform very well for certain types of audio sources, e.g. speech. For separation of instruments in music, there are several shortcomings. In general when instruments play together they are not independent of each other. More specifically the time-frequency distributions of the different sources will overlap. Harmonic instruments in particular have high probability of overlapping partials. If these overlapping partials are not separated properly, the separated signals will have a different sensation of roughness, and the separation quality degrades. In this paper we present a method to separate overlapping partials in stereo signals. This method looks at the shapes of partial envelopes, and uses minimization of the difference between such shapes in order to demix overlapping partials. The method can be applied to enhance existing methods for source separation, e.g. blind source separation techniques, model based techniques, and spatial separation techniques. We also discuss other simpler methods that can work with mono signals.
Download Synthesis by Mathematical Models Sound synthesis methods can be interpreted, from a mathe matical point of view, as a collection of techniques of selecting and conceptually organizing elements of a Hilbert space. In this sense, mathematics, being a highly structured and sophisticated system of classification, modeling and categorization, seems to be the natural tool to describe existing synthesis methods and to pro pose new ones. Because, from this perspective, one can think of any available (or theoretically predictable, or imaginable) synthe sis method as a collection of procedures to deal with meaningful parameters, with the term ”synthesis by mathematical models” we mean an extensive use of the modeling and categorization power of mathematics applied to the world of sounds. In this paper we give a few examples of sound synthesis tech niques, based on mathematical models. After reviewing shortly FM synthesis and synthesis by nonlinear distortion, and suggest ing some, to our advice, interesting open problems, we propose two different new methods: synthesis by means of elliptic func tions and synthesis by means of nowhere (or almostnowhere) dif ferentiable functions and lacunary series. The resulting waveforms have been produced using CSound as an audio engine, driven by Python scripts.
Download Audio-Tactile Glove This paper introduces the Audio-Tactile Glove, an experimental tool for the analysis of vibrotactile feedback in instrument design. Vibrotactile feedback provides essential information in the operation of acoustic instruments. The Audio-Tactile Glove is designed as a research tool for the investigation of the various techniques used to apply this theory to digital interfaces. The user receives vibrations via actuators distributed throughout the glove, located so as not to interrupt the physical contact required between user and interface. Using this actuator array, researchers will be able to independently apply vibrotactile information to six stimulation points across each hand exploiting the broad frequency range of the device, with specific sensitivity within the haptic frequency range of the hand. It is proposed that researchers considering the inclusion of vibrotactile feedback in existing devices can utilize this device without altering their initial designs.