Download Multi-View Subband Autoregressive Pole Harvesting for Modal Plate Identification We describe a Task B submission for the 1st DAFx Parameter Estimation Challenge. A matching-pursuit anchor stage seeds a multi-view subband autoregressive (AR) pole harvester on the raw IR and its first two finite differences; since linear filtering preserves pole locations while changing residues, the three views expose complementary subsets of the same pole set. Bands in which the AR order saturates are recursively split, and any remaining under-resolved region is completed from an IR-derived saturation indicator. Gains are assigned by a global ridge least-squares (LS) fit in the damped-biquad atom convention. The pipeline uses only the unnormalized IR; so it does not use plate parameters, analytical modal-frequency or decay laws, .wav files, or Task A information.
Download Loopback Frequency Modulation Using a Time-Varying Delay Line This work examines the use of the time-varying delay line (TVDL) to implement loopback frequency modulation (LBFM), an oscillator that loops back to modulate its own frequency. Digital delay lines are used regularly in sound synthesis/processing to model the pure delay associated with one-dimensional acoustic propagation. When the delay is made time varying, the TVDL time warps the input according to a delay function, altering the input's instantaneous frequency and phase. As a result, the TVDL is well suited for delay-based effects and, in particular, those involving frequency/phase modulation for which the TVDL delay function is oscillatory and thus bounded by a maximum and minimum delay. Sustaining a constant change in sounding frequency however, corresponds to a delay function having a term that is linear in time, making it limited only by the length of the input signal. While TVDLs may still be used when there is a pitch shift, limiting the delay by simple wrapping of the delay function and/or cross fading between multiple TVDLs may not be adequate to avoid audible artifacts. In LBFM, the resulting phase has both linear and oscillating terms and the resulting signal undergoes a sustained shift in the fundamental frequency that makes a TVDL implementation more challenging. An alternate closed-form representation of the LBFM oscillator, however, provides the information necessary for accurately wrapping the TVDL delay function and ensuring it is suitably bounded so that the produced sound is free of phase distortion and audible artifacts.
Download Residual-Driven Adaptive Multi-Rate Quadratic Programming Framework for Nonlinear Analog Audio Circuit Emulation This work extends our previously proposed Quadratic Programming (QP) approach for the emulation of nonlinear analog audio circuits by formalizing its main numerical ingredients and introducing a residual-driven adaptive multi-rate scheme. Starting from a state-space Differential Algebraic System of Equations (DAE) formulation, the nonlinear algebraic circuit device relations are replaced inside the QP by a first-order surrogate linear constraint, and the post-step nonlinear residual is shown to act as a valid defect indicator for adaptive step-size control. This yields a single-step simulation procedure that avoids the usual combination of nonlinear iterative solves and separate integration updates. The method is evaluated on a diode clipper, a BJT common-emitter amplifier, and a Colpitts oscillator, using SPICE as a baseline reference. The results show that adaptive step sizing considerably improves agreement with the reference solution, that the pseudo-inverse implementation is essentially equivalent to the full equality-constrained QP in the tested cases, and that the proposed formulation remains effective beyond the baseline clipper example, including for a self-oscillating circuit. These results position the proposed method as a promising bridge between SPICE-like interpretability and the efficiency demands of virtual analog (VA) audio applications.
Download SCAPES: Semantically Conditioned Autoregressive Prior for Environmental Sounds This paper presents SCAPES, a semantically conditioned autoregressive prior for environmental sound generation. The system models discrete audio representations using an autoregressive architecture conditioned on semantic information, enabling the generation of environmental sounds that follow user-specified concepts. By learning a prior over audio tokens, SCAPES combines high-level semantic control with detailed temporal modeling. Experimental evaluation investigates the quality, diversity, and semantic consistency of generated sounds, demonstrating the potential of autoregressive priors for controllable environmental sound synthesis.
Download Explicit Wave Digital Model of the Fulltone OCD Pedal Based on Canonical Piecewise-Linear Functions Virtual Analog (VA) modeling aims at digitally emulating analog audio equipment while preserving its characteristic nonlinear behavior and musical expressiveness. In the context of guitar effects, overdrive pedals represent a cornerstone of many signal chains, as they strongly contribute to the perceived dynamics, articulation, and timbral identity of the instrument. Among these, the Fulltone OCD overdrive is considered a standard in both studio and live environments, being widely adopted across rock and metal genres. In this article, we present an explicit Wave Digital (WD) model of the Fulltone OCD (v2) pedal. By exploiting the circuit topology, the MOSFETs and the germanium diode composing the asymmetric clipping stage are grouped into a single equivalent nonlinear element, enabling an explicit WD realization that avoids costly iterative solvers. The resulting nonlinear characteristic is approximated by means of a Canonical Piecewise-Linear (CPWL) function, yielding a compact and efficient explicit model suitable for real-time implementation. The proposed model is validated against reference simulations and implemented both in MATLAB and as a real-time audio plug-in using the JUCE framework.
Download Diffusion-Based Music Audio Editing System Using Differentiable Digital Signal Processing Mixture Model This paper proposes a music audio editing system that enables source-wise editing of harmonic instrument mixtures without explicit source separation. It builds on our previously proposed score-informed method for estimating source-wise synthesis parameters, i.e., time-varying controls used to synthesize each source, such as fundamental frequency and loudness. The method directly estimates these parameters from a mixture signal and the corresponding musical score in an analysis-by-synthesis framework. Using the estimated parameters, the proposed system allows users to edit individual sources by modifying note sequences and instrument types, and then re-synthesizes the edited mixture. Through demonstrations on two-instrument mixtures, we show that the system supports note-level phrasing modification and instrument conversion of selected sources.
Download Evaluating AI Coding Assistants in Audio DSP Education: A Small Scale Study Recent advances in AI-assisted coding tools raise questions about how programming-intensive subjects such as audio digital signal processing should be taught and how exam projects should be evaluated. This paper presents a small-scale controlled exploratory study conducted in a graduate course on music DSP. As a final project at the end of the course, the students implemented a modular synthesizer plugin in C++. Half of the students had access to AI-assisted coding support, while the other group developed the plugin manually. All students had to follow a protocol and provide data at the end of the project, together with their code, which was discussed with them as part of the course exam. Although the scale of the study is small, the paper shares a qualitative analysis of the results and a few takeaway messages for future reference among lecturers in the field. Overall, AI-assisted coding does provide some advantage to students but only in certain regards. The used AI tools, trained on GitHub repositories, seem to have only partial awareness of the state of the art in digital audio processing (e.g. antialiasing oscillators, virtual analog filters, etc.). Finally, the use of AI seems to not interfere excessively with the ability of the students to learn from their practical experience.
Download Perceptual Optimisation of Loudspeaker-Based Reproduction This paper proposes POLAR, a framework for the optimisation of loudspeaker signals using end-to-end differentiable perceptual loss functions. The framework optimises multiple perceptual attributes across multiple listeners, offering a versatile method for a range of problems. This versatility stems from the ability to customise the number of loudspeakers, listeners, and the weighting applied to different perceptual attributes. Here, we apply the method to four problems: (a) source panning for a single listener in stereo reproduction, (b) single-listener colouration matching in stereo reproduction, (c) extended sweet spot using stereo pairs beamforming, and (d) multi-attribute perceptually driven panning in stereo reproduction. The first three problems are evaluated against solutions traditionally used for these tasks: solutions of (a) are shown to be similar to those obtained with tangent panning law and vector-base amplitude panning (VBAP), solutions of (b) are shown to be similar to those obtained for cross-talk cancellation, and solutions of (c) are shown to be similar to those obtained in earlier work on directivity pattern optimisation for sweet spot widening. Each of these solutions was previously obtained using fundamentally different methodologies, demonstrating the flexibility and broad applicability of the proposed framework.
Download Bunkervik Spatial Reverb Demo This paper accompanies a demonstration of a real-time audio plug-in for a dynamic spatial reverb. The reverb is based on acoustic measurements of the Bunkervik creative arts space in Brescia, Italy. A modal synthesis reverberation engine was created based on measured impulse responses from three locations in the tunnel. Using a common set of modal frequencies the position of the receiver can be dynamically moved through the space by interpolating between data sets of residue weights and FIR filter taps. The audio plug-in also allows real-time manipulation of the high-frequency content, damping, and microphone rotation, all of which can be modulated using two LFOs.
Download Count-Density Networks for Modal Plate Parameter Estimation We describe two submissions to Task B of the 1st DAFx Parameter Estimation Challenge, which estimates an unknown number of modal frequency, decay, and gain triples from a synthetic plate-reverb impulse response. The first method combines pooled spectral features with time-domain and absolute-scale conditioning in a real-valued convolutional count-density network, while the second uses a complex-valued Transformer count-density network. Both methods jointly infer the modal count and per-mode attributes directly from the IR. On an independently generated 100-IR comparison set, the two neural estimators achieve lower overall challenge error than the evaluated classical baselines, with frequency and decay estimation substantially more accurate than gain estimation.