Download Physics-Inspired Feature Fusion for Plate Parameter Estimation from Acoustic Impulse Responses ★
Estimating physical plate parameters from impulse responses is a challenging inverse problem. Task A of the first Digital Audio Effects Parameter Estimation Challenge requires the recovery of six identifiable parameters from displacement impulse responses. In this work, we propose a physics-inspired feature fusion network (PIFFN) that combines a pretrained convolutional backbone with a 15-dimensional physics-inspired feature vector computed from the impulse response. These physics-inspired features describe amplitude scale, temporal decay, and spectral structure without relying on modal-distribution priors. The proposed model is evaluated on the official validation set, achieving an overall normalized mean squared error of 0.00362. Compared with the official particle swarm optimization baseline and backbone-only model, PIFFN shows a clear performance improvement, demonstrating its effectiveness for plate parameter estimation.
Download Amp-Space: A Large-Scale Dataset for Fine-Grained Timbre Transformation
We release Amp-Space, a large-scale dataset of paired audio samples: a source audio signal, and an output signal, the result of a timbre transformation. The types of transformations we study are from blackbox musical tools (amplifiers, stompboxes, studio effects) traditionally used to shape the sound of guitar, bass, or synthesizer sounds. For each sample of transformed audio, the set of parameters used to create it are given. Samples are from both real and simulated devices, the latter allowing for orders of magnitude greater data than found in comparable datasets. We demonstrate potential use cases of this data by (a) pre-training a conditional WaveNet model on synthetic data and show that it reduces the number of samples necessary to digitally reproduce a real musical device, and (b) training a variational autoencoder to shape a continuous space of timbre transformations for creating new sounds through interpolation.
Download A Generative Model for Raw Audio Using Transformer Architectures
This paper proposes a novel way of doing audio synthesis at the waveform level using Transformer architectures. We propose a deep neural network for generating waveforms, similar to wavenet . This is fully probabilistic, auto-regressive, and causal, i.e. each sample generated depends on only the previously observed samples. Our approach outperforms a widely used wavenet architecture by up to 9% on a similar dataset for predicting the next step. Using the attention mechanism, we enable the architecture to learn which audio samples are important for the prediction of the future sample. We show how causal transformer generative models can be used for raw waveform synthesis. We also show that this performance can be improved by another 2% by conditioning samples over a wider context. The flexibility of the current model to synthesize audio from latent representations suggests a large number of potential applications. The novel approach of using generative transformer architectures for raw audio synthesis is, however, still far away from generating any meaningful music similar to wavenet, without using latent codes/meta-data to aid the generation process.
Download SCAPES: Semantically Conditioned Autoregressive Prior for Environmental Sounds
This paper presents SCAPES, a semantically conditioned autoregressive prior for environmental sound generation. The system models discrete audio representations using an autoregressive architecture conditioned on semantic information, enabling the generation of environmental sounds that follow user-specified concepts. By learning a prior over audio tokens, SCAPES combines high-level semantic control with detailed temporal modeling. Experimental evaluation investigates the quality, diversity, and semantic consistency of generated sounds, demonstrating the potential of autoregressive priors for controllable environmental sound synthesis.
Download Extracting More Detail from the Spectrum with Phase Distortion Analysis
In the sinusoidal analysis of sound, using the Short Time Fourier Transform (STFT), there is the assumption that the signal is locally stationary within each FFT frame. If, as in practice, this assumption is violated, the spectrum becomes distorted. Phase Distortion Analysis (PDA) was introduced in 1995 [1] to enhance the analysis of degraded peaks, by using the distortion itself as a source of information about the signal nonstationarity. It was shown that the first order frequency and amplitude modulation could be measured from the degree of phase shift close to the maximum of the mainlobe peak. This paper presents advances with the PDA technique, in particular a neural network implementation that makes estimation robust to noise. The capability to analyse nonstationarities relaxes the restraint on keeping the FFT analysis window short and therefore effectively improves time-frequency resolution. This, in turn, promises greater analysis-synthesis quality through improved identification and tracking of partials during the analysis phase.
Download SEND: A Spatial Event Neural Detector for Intentional Object Motion in Immersive Music Mixing
Deciding exactly when to move audio objects in immersive mixes is a labor-intensive artistic task. Current tools react strictly to instantaneous frequency overlaps, lacking the macroscopic awareness required for musically intentional spatial transitions. To model these decisions, we propose SEND (Spatial Event Neural Detector). Its dual-stream architecture analyzes the target track against its background context, combining a Spec-TNT backbone and a Temporal Convolutional Network (TCN) to capture hierarchical spectral features and precise rhythmic cues. Their dynamic interplay is modeled via a novel Cross-Track Gating Interaction (CTGI) mechanism.
Download Neural Networks for Physical Parameter Estimation of Plate Reverberation from Impulse Responses ★
This paper presents our Task A submission to the 1st DAFx Parameter Estimation Challenge. We use the official ModalPlate dataset generator to synthesize 1000 one-second plate impulse responses with randomly sampled parameters inside the public ranges. A time-domain CNN-GRU regressor then estimates the six official Task A parameters from each unnormalised waveform. The model combines three one-dimensional convolutional blocks with a bidirectional gated recurrent unit and is trained with mean squared error on min-max normalised targets. The generated data are split into 700/150/150 train/validation/test examples, and the test split is never used during training or model selection. The implementation follows the official Task A format and exports evaluation-compatible prediction files for both development evaluation and blind-set submission.
Download Differentiable White-Box Virtual Analog Modeling
Component-wise circuit modeling, also known as “white-box” modeling, is a well established and much discussed technique in virtual analog modeling. This approach is generally limited in accuracy by lack of access to the exact component values present in a real example of the circuit. In this paper we show how this problem can be addressed by implementing the white-box model in a differentiable form, and allowing approximate component values to be learned from raw input–output audio measured from a real device.
Download Modulation Extraction for LFO-driven Audio Effects
Low frequency oscillator (LFO) driven audio effects such as phaser, flanger, and chorus, modify an input signal using time-varying filters and delays, resulting in characteristic sweeping or widening effects. It has been shown that these effects can be modeled using neural networks when conditioned with the ground truth LFO signal. However, in most cases, the LFO signal is not accessible and measurement from the audio signal is nontrivial, hindering the modeling process. To address this, we propose a framework capable of extracting arbitrary LFO signals from processed audio across multiple digital audio effects, parameter settings, and instrument configurations. Since our system imposes no restrictions on the LFO signal shape, we demonstrate its ability to extract quasiperiodic, combined, and distorted modulation signals that are relevant to effect modeling. Furthermore, we show how coupling the extraction model with a simple processing network enables training of end-to-end black-box models of unseen analog or digital LFO-driven audio effects using only dry and wet audio pairs, overcoming the need to access the audio effect or internal LFO signal. We make our code available and provide the trained audio effect models in a real-time VST plugin1 .
Download RAVE for Speech: Efficient Voice Conversion at High Sampling Rates
Voice conversion has gained increasing popularity within the field of audio manipulation and speech synthesis. Often, the main objective is to transfer the input identity to that of a target speaker without changing its linguistic content. While current work provides high-fidelity solutions they rarely focus on model simplicity, high-sampling rate environments or stream-ability. By incorporating speech representation learning into a generative timbre transfer model, traditionally created for musical purposes, we investigate the realm of voice conversion generated directly in the time domain at high sampling rates. More specifically, we guide the latent space of a baseline model towards linguistically relevant representations and condition it on external speaker information. Through objective and subjective assessments, we demonstrate that the proposed solution can attain levels of naturalness, quality, and intelligibility comparable to those of a state-of-the-art solution for seen speakers, while significantly decreasing inference time. However, despite the presence of target speaker characteristics in the converted output, the actual similarity to unseen speakers remains a challenge.