Download Real-Time Black-Box Modelling With Recurrent Neural Networks This paper proposes to use a recurrent neural network for black-box modelling of nonlinear audio systems, such as tube amplifiers and distortion pedals. As a recurrent unit structure, we test both Long Short-Term Memory and a Gated Recurrent Unit. We compare the proposed neural network with a WaveNet-style deep neural network, which has been suggested previously for tube amplifier modelling. The neural networks are trained with several minutes of guitar and bass recordings, which have been passed through the devices to be modelled. A real-time audio plugin implementing the proposed networks has been developed in the JUCE framework. It is shown that the recurrent neural networks achieve similar accuracy to the WaveNet model, while requiring significantly less processing power to run. The Long Short-Term Memory recurrent unit is also found to outperform the Gated Recurrent Unit overall. The proposed neural network is an important step forward in computationally efficient yet accurate emulation of tube amplifiers and distortion pedals.
Download Sample Rate Independent Recurrent Neural Networks for Audio Effects Processing In recent years, machine learning approaches to modelling guitar amplifiers and effects pedals have been widely investigated and have become standard practice in some consumer products. In particular, recurrent neural networks (RNNs) are a popular choice for modelling non-linear devices such as vacuum tube amplifiers and distortion circuitry. One limitation of such models is that they are trained on audio at a specific sample rate and therefore give unreliable results when operating at another rate. Here, we investigate several methods of modifying RNN structures to make them approximately sample rate independent, with a focus on oversampling. In the case of integer oversampling, we demonstrate that a previously proposed delay-based approach provides high fidelity sample rate conversion whilst additionally reducing aliasing. For non-integer sample rate adjustment, we propose two novel methods and show that one of these, based on cubic Lagrange interpolation of a delay-line, provides a significant improvement over existing methods. To our knowledge, this work provides the first in-depth study into this problem.
Download On the numerical solution of the 2D wave equation with compact FDTD schemes This paper discusses compact-stencil nite difference time domain (FDTD) schemes for approximating the 2D wave equation in the context of digital audio. Stability, accuracy, and efciency are investigated and new ways of viewing and interpreting the results are discussed. It is shown that if a tight accuracy constraint is applied, implicit schemes outperform explicit schemes. The paper also discusses the relevance to digital waveguide mesh modelling, and highlights the optimally efcient explicit scheme.
Download GPGPU Patterns for Serial and Parallel Audio Effects Modern commodity GPUs offer high numerical throughput per
unit of cost, but often sit idle during audio workstation tasks. Various researches in the field have shown that GPUs excel at tasks
such as Finite-Difference Time-Domain simulation and wavefield
synthesis. Concrete implementations of several such projects are
available for use.
Benchmarks and use cases generally concentrate on running
one project on a GPU. Running multiple such projects simultaneously is less common, and reduces throughput. In this work
we list some concerns when running multiple heterogeneous tasks
on the GPU. We apply optimization strategies detailed in developer documentation and commercial CUDA literature, and show
results through the lens of real-time audio tasks. We benchmark
the cases of (i) a homogeneous effect chain made of previously
separate effects, and (ii) a synthesizer with distinct, parallelizable
sound generators.
Download Sound Spatialisation Towards the end of the nineteenth century two inventions, the telephone and the phonograph, appeared which were to change the way music was dealt with. Prior to the developments they brought about, every musical performance was indivisible from its place in time and space. Their appearance meant that music could be presented remotely in both time and space from its origin. This has inevitably resulted in various forms of distortion - nonlinear, spectral, temporal, spatial, of the original. Whilst it has proved relatively easy to deal satisfactorily with the first three so that we can now present “remoted” music which is excellent in all of those three aspects, removing the distortions in the spatial presentation has proved far more intractable. Even the best systems in use today for sound “spatialisation” are relatively crude, allowing for little more than the creation of an illusion, sometimes very good, more often poor. However, despite this, the creative use of sound spatialisation is becoming more and more important, whether for serious avant-garde composers, computer game designers, in cinema, television and multimedia productions or in audio recording. It is anticipated that the demand will escalate even more with the appearance of DVD with its multiple audio channel capability. This tutorial paper briefly covers the basic directional hearing mechanisms of the human brain before examining in more detail the various different ways of dealing with sound spatialisation, starting with headphone related technologies such as binaural and transaural. Loudspeaker-based systems will then be covered, starting with conventional stereo followed by cinema style surround sound systems. Finally a true 3-d system, Ambisonics, will be examined. The advantages and limitations of all the systems, both aurally and in terms of difficulty of implementation or control, will be covered. It is hoped to give a number of demonstrations.
Download Realistic Gramophone Noise Synthesis Using a Diffusion Model This paper introduces a novel data-driven strategy for synthesizing gramophone noise audio textures. A diffusion probabilistic model is applied to generate highly realistic quasiperiodic noises. The proposed model is designed to generate samples of length equal to one disk revolution, but a method to generate plausible periodic variations between revolutions is also proposed. A guided approach is also applied as a conditioning method, where an audio signal generated with manually-tuned signal processing is refined via reverse diffusion to improve realism. The method has been evaluated in a subjective listening test, in which the participants were often unable to recognize the synthesized signals from the real ones. The synthetic noises produced with the best proposed unconditional method are statistically indistinguishable from real noise recordings. This work shows the potential of diffusion models for highly realistic audio synthesis tasks.
Download Development of an outdoor auralisation prototype with 3D sound reproduction Auralisation of outdoor sound has a strong potential for demonstrating the impact of different community noise scenarios. We describe here the development of an auralisation tool for outdoor noise such as traffic or industry. The tool calculates the sound propagation from source to listener using the Nord2000 model, and represents the sound field at the listener’s position using spherical harmonics. Because of this spherical harmonics approach, the sound may be reproduced in various formats, such as headphones, stereo, or surround. Dynamic reproduction in headphones according to the listener’s head orientation is also possible through the use of head tracking.
Download Harmonic Mixing Based on Roughness and Pitch Commonality The practice of harmonic mixing is a technique used by DJs for the beat-synchronous and harmonic alignment of two or more pieces of music. In this paper, we present a new harmonic mixing method based on psychoacoustic principles. Unlike existing commercial DJ-mixing software which determine compatible matches between songs via key estimation and harmonic relationships in the circle of fifths, our approach is built around the measurement of musical consonance at the signal level. Given two tracks, we first extract a set of partials using a sinusoidal model and average this information over sixteenth note temporal frames. Then within each frame, we measure the consonance between all combinations of dyads according to psychoacoustic models of roughness and pitch commonality. By scaling the partials of one track over ± 6 semitones (in 1/8th semitone steps), we can determine the optimal pitch-shift which maximises the consonance of the resulting mix. Results of a listening test show that the most consonant alignments generated by our method were preferred to those suggested by an existing commercial DJ-mixing system.
Download Numerical Simulation of Spring Reverberation Virtual analog modeling of spring reverberation presents a challenging problem to the algorithm designer, regardless of the particular strategy employed. The difficulties lie in the behaviour of the helical spring, which, due to its inherent curvature, shows characteristics of both coherent and dispersive wave propagation. Though it is possible to emulate such effects in an efficient manner using audio signal processing constructs such as delay lines (for coherent wave propagation) and chains of allpass filters (for dispersive wave propagation), another approach is to make use of direct numerical simulation techniques, such as the finite difference time domain method (FDTD) in order to solve the equations of motion directly. Such an approach, though more computationally intensive, allows a closer link with the underlying model system— and yet, there are severe numerical difficulties associated with such designs, and in particular anomalous numerical dispersion, requiring some care at the design stage. In this paper, a complete model of helical spring vibration is presented; dispersion analysis from an audio perspective allows for model simplification. A detailed description of novel FDTD designs follows, with special attention is paid to issues such as numerical stability, loss modeling, numerical boundary conditions, and computational complexity. Simulation results are presented.
Download Neural-Driven Multi-Band Processing for Automatic Equalization and Style Transfer We present a Neural-Driven Multi-Band Processor (NDMP), a differentiable audio processing framework that augments a static sixband Parametric Equalizer (PEQ) with per-band dynamic range
compression. We optimize this processor using neural inference
for two tasks: Automatic Equalization (AutoEQ), which estimates
tonal and dynamic corrections without a reference, and Production
Style Transfer (NDMP-ST), which adapts the processing of an input signal to match the tonal and dynamic characteristics of a reference. We train NDMP using a self-supervised strategy, where the
model learns to recover a clean signal from inputs degraded with
randomly sampled NDMP parameters and gain adjustments. This
setup eliminates the need for paired input–target data and enables
end-to-end training with audio-domain loss functions. In the inference, AutoEQ enhances previously unseen inputs in a blind setting, while NDMP-ST performs style transfer by predicting taskspecific processing parameters. We evaluate our approach on the
MUSDB18 dataset using both objective metrics (e.g., SI-SDR,
PESQ, STFT loss) and a listening test.
Our results show that
NDMP consistently outperforms traditional PEQ and a PEQ+DRC
(single-band) baseline, offering a robust neural framework for audio enhancement that combines learned spectral and dynamic control.