Download Accurate Plate Reverb Parameter Estimation Using Two-Stage Evolutionary Search We describe our submission to Task A of the 1st DAFx parameter estimation challenge. The task is to recover the six physical parameters of a simulated metal-plate reverberator—its dimensions and material properties—from a single impulse response (IR). We treat this as a black-box optimization: candidate parameter sets are fed to the simulator and scored by a loss against the target IR. The method has two stages. The first uses CMA-ES, an evolutionary optimizer, to recover five of the six parameters, comparing IRs under an amplitude-normalized loss. Amplitude normalization makes the search robust but discards the cue to the sixth parameter, the plate’s surface density; a second stage therefore estimates it alone, with a ternary search on the un-normalized loss. As the choice of loss strongly affects the search, we select it beforehand, and analyze why compression in the common multi-scale spectral loss degrades recovery. Finally, we test our method on a validation set of 50 IRs, discuss a pathological failure mode, and ablate to justify having two different stages instead of a unified CMA-ES search.
Download Declaratively Programmable Ultra Low-Latency Audio Effects Processing on FPGA WaveCore is a coarse-grained reconfigurable processor architecture, based on data-flow principles. The processor architecture consists of a scalable and interconnected cluster of Processing Units (PU), where each PU embodies a small floating-point RISC processor. The processor has been designed in technology-independent VHDL and mapped on a commercially available FPGA development platform. The programming methodology is declarative, and optimized to the application domain of audio and acoustical modeling. A benchmark demonstrator algorithm (guitar-model, comprehensive effects-gear box, and distortion/cabinet model) has been developed and applied to the WaveCore development platform. The demonstrator algorithm proved that WaveCore is very well suited for efficient modeling of complex audio/acoustical algorithms with negligible latency and virtually zero jitter. An experimental Faust-to-WaveCore compiler has shown the feasibility of automated compilation of Faust code to the WaveCore processor target. Keywords: ultra-low latency, zero-jitter, coarse-grained reconfigurable computing, declarative programming, automated manycore compilation, Faust-compatible, massively-parallel
Download Real-Time Implementation of an Elasto-Plastic Friction Model using Finite-Difference Schemes The simulation of a bowed string is challenging due to the strongly non-linear relationship between the bow and the string. This relationship can be described through a model of friction. Several friction models in the literature have been proposed, from simple velocity dependent to more accurate ones. Similarly, a highly accurate technique to simulate a stiff string is the use of finitedifference time-domain (FDTD) methods. As these models are generally computationally heavy, implementation in real-time is challenging. This paper presents a real-time implementation of the combination of a complex friction model, namely the elastoplastic friction model, and a stiff string simulated using FDTD methods. We show that it is possible to keep the CPU usage of a single bowed string below 6 percent. For real-time control of the bowed string, the Sensel Morph is used.
Download Navigating in a Space of Synthesized Interaction-Sounds: Rubbing, Scratching and Rolling Sounds In this paper, we investigate a control strategy of synthesized interaction-sounds. The framework of our research is based on the action/object paradigm that considers that sounds result from an action on an object. This paradigm presumes that there exists some sound invariants, i.e. perceptually relevant signal morphologies that carry information about the action or the object. Some of these auditory cues are considered for rubbing, scratching and rolling interactions. A generic sound synthesis model, allowing the production of these three types of interaction together with a control strategy of this model are detailed. The proposed control strategy allows the users to navigate continuously in an ”action space”, and to morph between interactions, e.g. from rubbing to rolling.
Download Approximating non-linear inductors using time-variant linear filters In this paper we present an approach to modeling the non-linearities of analog electronic components using time-variant digital linear filters. The filter coefficients are computed at every sample depending on the current state of the system. With this technique we are able to accurately model an analog filter including a nonlinear inductor with a saturating core. The value of the magnetic permeability of a magnetic core changes according to its magnetic flux and this, in turn, affects the inductance value. The cutoff frequency of the filter can thus be seen as if it is being modulated by the magnetic flux of the core. In comparison to a reference nonlinear model, the proposed approach has a lower computational cost while providing a reasonably small error.
Download Implementing a Low Latency Parallel Graphic Equalizer with Heterogeneous Computing This paper describes the implementation of a recently introduced parallel graphic equalizer (PGE) in a heterogeneous way. The control and audio signal processing parts of the PGE are distributed to a PC and to a signal processor, of WaveCore architecture, respectively. This arrangement is particularly suited to the algorithm in question, benefiting from the low-latency characteristics of the audio signal processor as well as general purpose computing power for the more demanding filter coefficient computation. The design is achieved cleanly in a high-level language called Kronos, which we have adapted for the purposes of heterogeneous code generation from a uniform program source.
Download A Perceptually Inspired Single Parameter Auditory Distance Renderer for Music Production Conveying auditory distance in a digital audio workstation requires balancing several uncoupled tools (reverb, gain, equalization, pre-delay) by hand, a workflow that is cognitively demanding and easily produces spatially incoherent results. We demonstrate a real-time VST3 plugin that derives five correlated distance cues from a single normalized control and keeps them mutually coherent by construction, grounded in the psychoacoustics of auditory distance perception. A headphone listening test with 18 participants showed that the coupled renderer roughly halves distance placement error relative to an uncoupled manual mix and was unanimously preferred on composite spatial quality. Users sweep one knob and hear sources move convincingly from near to far on multitrack material, compare the result against a manual uncoupled mix, and toggle an optional binaural externalization stage. Source code is available on GitHub.
Download Auditory Perception of Spatial Extent in the Horizontal and Vertical Plane This article investigates the accuracy with which listeners can identify the spatial extent of distributed sound sources. Either the complementary frequency bands comprising a source signal or the individual grains of a granular synthesis-based stimulus were distributed directly on discrete loudspeakers. Loudspeakers were arranged either on the horizontal or the vertical axis. The algorithms were applied on white noise, an impulse train, and a rain drops stimulus. Absolute judgments of spatial extent were obtained separately for each orientation, algorithm, and stimulus using three different magnitudes of horizontal or vertical extent. Horizontal spatial extent judgments varied systematically with physical extent for all conditions in the experiment. The correspondence between perceived and actual vertical extent was poor. The time-based synthesis algorithm resulted in significantly larger judgments of spatial extent irrespective of orientation and stimulus compared to the frequency-based algorithm.
Download A Corpus-Driven Parametric Modal Reverberator A parametric modal reverberator is presented in which synthesis parameters are derived from a large, curated corpus of room impulse responses (IRs). The collected responses are subjected to modal decomposition, yielding per-mode frequencies, damping coefficients, and residue amplitudes, together with a short early-reflection finite impulse response (FIR) filter. From the decomposed data, a feature table is constructed per IR comprising standard acoustic indices, per-band damping and density statistics, amplitude distributions, and FIR descriptors—50 variables in total. Six acoustically meaningful user controls are selected; since these exhibit substantial pairwise correlations across the corpus, they are orthogonalised via principal component analysis (PCA) prior to regression.
Download Hard real-time onset detection of percussive instruments To date, the most successful onset detectors are those based on frequency representation of the signal. However, for such methods the time between the physical onset and the reported one is unpredictable and may largely vary according to the type of sound being analyzed. Such variability and unpredictability of spectrum-based onset detectors may not be convenient in some real-time applications. This paper proposes a real-time method to improve the temporal accuracy of state-of-the-art onset detectors. The method is grounded on the theory of hard real-time operating systems where the result of a task must be reported at a certain deadline. It consists of the combination of a time-base technique (which has a high degree of accuracy in detecting the physical onset time but is more prone to false positives and false negatives) with a spectrum-based technique (which has a high detection accuracy but a low temporal accuracy). The developed hard real-time onset detector was tested on a dataset of single non-pitched percussive sounds using the high frequency content detector as spectral technique. Experimental validation showed that the proposed approach was effective in better retrieving the physical onset time of about 50% of the hits detected by the spectral technique, with an average improvement of about 3 ms and maximum one of about 12 ms. The results also revealed that the use of a longer deadline may capture better the variability of the spectral technique, but at the cost of a bigger latency.