Download Sound Matching with a Differentiable Karplus-Strong Algorithm
We present a self-supervised, event-based sound matching model using a differentiable extended Karplus-Strong algorithm. To avoid relying on external onset and fundamental frequency detectors, we explore training methodologies combining parameter losses on synthetic data with audio losses. We demonstrate that time-domain fractional delay interpolation provides gradient accuracy comparable to frequency-sampling while avoiding time-aliasing in highly resonant time-varying scenarios. Through systematic gradient analysis, we reveal that standard spectral losses provide no meaningful directional gradients for onset times, heavily degrading joint training. Training exclusively with parameter losses on synthetic data effectively learns fundamental frequency, timbral parameters, and onset times, but struggles to generalise to monophonic studio recordings of plucked guitar. External detectors combined with audio losses generalise best, isolating the model to timbre optimisation. While our Karplus-Strong decoder recovers interpretable parameters and naturally captures the transient characteristics of plucked guitar, Harmonics plus Noise baselines yield higher reconstruction fidelity by most metrics.
Download Praat AudioTools: Analysis Objects as Compositional Controllers for Interpretable Sound Transformation
This demonstration presents Praat AudioTools, an open-source hybrid toolkit that repurposes Praat's phonetic-analysis environment for electroacoustic composition, sound design, and offline analysis–resynthesis workflows. Rather than treating analysis data as temporary measurements hidden inside an audio processor, Praat AudioTools exposes pitch contours, formant structures, temporal segmentations, spectral descriptors, phrase boundaries, stochastic trajectories, and host-application exchange files as editable compositional objects. These objects can be inspected, modified, chained, reused, and rendered into new sound transformations. The demonstration focuses on seven offline workflows: Neural Ambient Drone Designer, Praat for Max and Max for Live, Phase-Space Composer, Reich Generator, MCMC Musical Variation, Messagesquisse Opening, and Vector/Full-Chain composition workflows. None of the examples are presented as real-time effects. Instead, they show an "edit-in-the-middle" model in which sound is analyzed, intermediate representations are made visible, compositional decisions are applied to those representations, and the result is rendered as audio. The aim is to demonstrate a transparent alternative to both conventional black-box audio effects and end-to-end generative audio systems: a compositional environment where analysis objects become controllers, traces, scores, and reproducible technical artifacts.
Download A DDSP Framework for Adaptive Room Equalization
Adaptive room equalization remains challenging under time-varying acoustic conditions and complex excitation signals, such as music. In these scenarios, classical filtered-x least mean squares (Fx-LMS) methods falter due to their rigid formulation. We present a modular differentiable digital signal processing (DDSP) framework for closed-loop adaptive room equalization that recovers Fx-LMS as a special case through automatic differentiation. The framework supports interchangeable EQ structures, response estimation methods, loss functions, and optimizers. Experiments with time-varying measured room impulse responses show that frequency-domain objectives provide more stable adaptation than time-domain objectives in the considered scenarios. Relative to the non-equalized response, system distance is reduced by 70% and mel-spectral distance by 13% (worst-case scenario). We further examine how online room response estimation accuracy and frame length affect the trade-off between responsiveness and convergence stability. Overall, the framework provides a unified open-source basis for exploring synergies between classical adaptive filtering and DDSP-based optimization.
Download A Comparative Study of Kolmogorov-Arnold Networks and Multi-Layer Perceptrons for Virtual Analog Modeling in Wave Digital Filters
The design of Virtual Analog (VA) algorithms has traditionally been divided between white-box (physics-based) and black-box (data-driven) approaches. Recent work has shown that hybrid methods, combining physical modeling with neural networks, can effectively leverage the strengths of both paradigms. In particular, Wave Digital Filters (WDFs) can be coupled with Multi-Layer Perceptrons (MLPs) to model circuits with multiple nonlinearities in a fully explicit manner. In this paper, we present a comparative study investigating the use of Kolmogorov-Arnold Networks (KANs) for VA modeling within the WDF framework. Unlike MLPs, KANs shift the learning paradigm by parameterizing activation functions instead of relying exclusively on learned weight matrices, potentially enabling more compact representations. Results show that, for our case study, KANs achieve accuracy comparable to MLPs while requiring approximately 70% fewer parameters at the cost of increased computational complexity. These findings suggest that KANs may represent a promising alternative in scenarios where memory footprint is a primary constraint, such as embedded audio applications, or when target models feature numerous nonlinear elements.
Download Neural Morphing: Sequence-Optimized Token-Level Morphing in Neural Audio Codecs
Neural audio codecs were originally developed for high-fidelity compression; however, their latent token representations and expressive decoders also constitute a powerful substrate for controllable audio transformation. This work introduces Neural Morphing, a training-free token-domain audio effect that selects residual-vector-quantized (RVQ) token grains from a user palette and decodes the edited stream through a pretrained codec. The method combines an RVQ-group transfer policy that separates coarse, middle, and fine codebook groups with a continuity-constrained sequence matcher that replaces independent greedy selection with bounded beam search. The intended output is a controlled hybrid: the source preserves rhythmic organization while the palette contributes timbral color and residual detail. We focus on the implementation and real-time behavior of a deployable VST3/AU system, including chunked rendering, palette-size scaling, and backend health checks.
Download Audio-to-Audio via Diffusion Warm Initialization
In this paper, we propose diffusion warm initialization as a simple yet effective approach for a range of audio-to-audio transformation tasks. To illustrate the generality of the approach, we demonstrate its use in timbre transfer, MIDI-to-Real synthesis, and multiple audio enhancement tasks. We conduct a detailed empirical analysis on timbre transfer to investigate the role of the initialization time t_init. The effect of t_init is evaluated using pitch-based Jaccard Distance and Fréchet Audio Distance to quantify faithfulness to the input signal and alignment with the target distribution. Our results provide practical guidance for selecting t_init and show that, once properly chosen, a single pretrained diffusion model combined with warm initialization can support multiple transformation objectives without task-specific training or conditioning. Despite its simplicity, this approach already achieves competitive results when compared with more complex pipelines designed specifically for these tasks.
Download WildFX: A DAW-Powered Pipeline for In-the-Wild Audio FX Graph Modeling
We present WildFX, a digital-audio-workstation-powered pipeline for modeling audio-effects graphs from in-the-wild audio. The system uses a DAW environment to construct, render, and evaluate effect-processing graphs, enabling research on realistic effect chains beyond isolated processors or synthetic training settings. WildFX supports the analysis and reconstruction of complex audio transformations by combining flexible plugin routing with data-driven modeling. The pipeline is designed to facilitate scalable dataset creation and experimentation with effect graph inference, parameter estimation, and audio transformation in practical production contexts.
Download Perceptually Motivated Alignment and Interpolation of Pitch-Aligned Time-Frequency Representations
This paper proposes methods for alignment and interpolation of pitch-aligned time-frequency representations, building on the tonal interval vector. Extensions reformulate it as an invertible operator, enabling alignment via permutation search under perceptually weighted distances and interpolation via optimal transport with a circular formulation respecting harmonic structure. The work also develops a geometric scale representation factorizing scale structure into root, density, and color.
Download A Unified Framework for Real-Time Concatenation-Driven Convolution
This work introduces a novel framework for Concatenation-Driven Convolution (CDC), unifying concatenative synthesis and real-time convolution into a single integrated audio processing paradigm. While concatenative synthesis has traditionally been used for corpus-based sound generation and convolution has served as a largely static filtering technique, the proposed approach reconceptualizes impulse responses (IRs) as dynamic, navigable sonic material. In the CDC framework, a corpus of audio segments is analyzed using perceptual features and organized via a self-organizing map (SOM), enabling intuitive, gesture-based traversal of a structured timbral space; the resulting concatenative output is treated as a continuously evolving impulse response and injected directly into a partitioned convolution engine. Its central technical contribution is single-engine frequency-domain kernel interpolation: rather than crossfading the outputs of two convolution engines, the FFT-domain kernels of the current and target IRs are interpolated within a single engine, preserving the internal convolution state across IR transitions and avoiding the warm-up energy loss inherent to dual-engine crossfading.
Download Evaluating Dynamic Range Compressor Models Using Control-Voltage Measurements: An Approach and Dataset
The quantity that defines the behavior of a dynamic range compressor is the time-varying gain applied to the signal as a function of the input level. However, models of these devices are typically evaluated using proxy metrics because isolating the gain reduction signal from the audio input–output data included in existing datasets creates an ill-conditioned inverse problem. It is unclear how accurately these metrics describe the behavior the model is tasked with emulating, particularly as waveform-based metrics can be influenced by secondary effects introduced by analog processing and capture, even when those effects are inaudible. We investigate a method of evaluation in which the gain-reduction signal produced by a model is measured directly against a gain-reduction control voltage signal produced by the hardware. To evaluate the efficacy of this metric as a learning objective, a gray-box model is trained using loss computed directly over the gain control signals alongside two models trained using common proxy losses. The models trained using proxy losses did not achieve parity with models trained directly on the gain control signal when evaluated with respect to the underlying control trajectory, and the waveform-domain metrics assigned similar errors to models that were clearly separated by the direct metric. To facilitate further exploration of this method of evaluation, we present a Solid State Logic bus compressor dataset that includes the gain control voltage signal captured alongside the audio output.