Download SEND: A Spatial Event Neural Detector for Intentional Object Motion in Immersive Music Mixing Deciding exactly when to move audio objects in immersive mixes is a labor-intensive artistic task. Current tools react strictly to instantaneous frequency overlaps, lacking the macroscopic awareness required for musically intentional spatial transitions. To model these decisions, we propose SEND (Spatial Event Neural Detector). Its dual-stream architecture analyzes the target track against its background context, combining a Spec-TNT backbone and a Temporal Convolutional Network (TCN) to capture hierarchical spectral features and precise rhythmic cues. Their dynamic interplay is modeled via a novel Cross-Track Gating Interaction (CTGI) mechanism.
Download Band-Count Dense Modal Estimation with Fixed-Frequency Differentiable Resonator Refinement ★ Task B of the 1st DAFx Parameter Estimation Challenge requires estimating the frequencies, decay rates, gains, and number of modes in a dense plate-reverb impulse response. Weak and overlapping modes make sparse peak detection prone to severe undercounting. We train an ExtraTrees regressor on simulator-generated data to predict mode counts in four frequency bands. These counts define dense frequency grids, after which a differentiable all-pole resonator model refines decay and gain while keeping frequency fixed. On two separate synthetic validation sets, the system reduces a local challenge-style error by about 66% relative to the official default peak-picking baseline. The improvement is mainly associated with lower mode-count mismatch, while decay and gain remain the largest error sources. These findings support separating modal-density estimation from continuous parameter fitting.
Download L-BOW: Gesture-Driven Digital Audio Effects for Augmented Violin in a Unified Csound Environment Live performance leaves little room for a sensor pipeline that misfires; when a gesture fails to map correctly to an intended effect, the error is immediately audible. This paper presents L-Bow, a wrist-worn six-degree-of-freedom (6-DoF) inertial measurement unit (IMU) controller for augmented violin performance. By removing intermediate software layers, L-Bow integrates gesture sensing, six performance modes, and a shared digital effects chain within a single, self-contained Csound file. This is achieved using Csound's native arduinoRead opcode for direct serial communication rather than an external Python–OSC bridge. The paper discusses the architecture of this system and its implications for designing dependable, low-maintenance interactive digital audio effects.
Download A Frequency-Domain Reverberator Plug-In We present FDverb, a frequency-domain artificial reverberator, based on the idea of a vocoder with a noise carrier signal. Using a short-time Fourier transform (STFT) for analysis and synthesis, FDverb generates late reverberation by weighting spectral noise components with envelopes. We extend FDverb with early reflections, nonlinear decay, and pitch shifting. These extensions enable creative sound-design applications. We provide FDverb as an open-source DAW plug-in, using the JUCE framework.
Download Fourier Neural Operators for Sample-Rate-Independent Virtual Analog Modeling Neural networks that operate directly on time-domain signals are widely used for virtual analog (VA) modeling. A key limitation of these models is their dependence on the sampling rate used during training, which becomes implicitly encoded in the learned parameters, so that changing it generally alters the realized dynamics. Although architectural modifications to recurrent neural networks have been proposed to enable sample-rate independent operation, these approaches are inherently tailored to upsampling and do not accommodate downsampling scenarios. In this manuscript, we present a VA modeling framework based on Fourier Neural Operators (FNOs) adapted to process fixed-duration audio frames. The proposed formulation defines the learned mapping over a fixed temporal support and evaluates it on uniform grids of different densities, so that a model trained at a single sampling rate can be applied at unseen sampling resolutions. Numerical results on a nonlinear transistor circuit show that the proposed model achieves competitive accuracy in upsampling scenarios while remaining directly applicable to downsampling, unlike a sample-rate independent baseline recurrent architecture.
Download A Unified Framework for Real-Time Concatenation-Driven Convolution This work introduces a novel framework for Concatenation-Driven Convolution (CDC), unifying concatenative synthesis and real-time convolution into a single integrated audio processing paradigm. While concatenative synthesis has traditionally been used for corpus-based sound generation and convolution has served as a largely static filtering technique, the proposed approach reconceptualizes impulse responses (IRs) as dynamic, navigable sonic material. In the CDC framework, a corpus of audio segments is analyzed using perceptual features and organized via a self-organizing map (SOM), enabling intuitive, gesture-based traversal of a structured timbral space; the resulting concatenative output is treated as a continuously evolving impulse response and injected directly into a partitioned convolution engine. Its central technical contribution is single-engine frequency-domain kernel interpolation: rather than crossfading the outputs of two convolution engines, the FFT-domain kernels of the current and target IRs are interpolated within a single engine, preserving the internal convolution state across IR transitions and avoiding the warm-up energy loss inherent to dual-engine crossfading.
Download Measurement-Informed Nonlinear Modal Synthesis of 65 Classical Guitars When a classical guitar string is plucked, vibration energy flows through the bridge into the body and is radiated as sound. Synthesising this process for a large collection of instruments requires both an efficient nonlinear string model and a robust method for extracting instrument-specific parameters from measurements. This paper addresses both issues. Starting from the publicly available dataset of Mores, which provides impulse-response measurements on 65 classical guitars, modal parameters of the bridge compliance and of the bridge-to-air radiation path are extracted for each instrument. These feed a nonlinear string model in which transverse vibration is governed by a geometrically exact elastic potential coupled at an interior bridge point to the measured body data. The nonlinear potential is quadratised via the Scalar Auxiliary Variable (SAV) method, so that the equations of motion become linear in a scalar variable and a known gradient vector, even at the continuous level. After time discretisation, the coupled system is inverted through two sequential Sherman–Morrison rank-one updates (one for the bridge coupling, one for the SAV nonlinearity), yielding an O(N) algorithm per time step. Two regularisation techniques prevent long-term drift of the auxiliary variable. The complete pipeline is demonstrated by synthesising plucked notes across all frets and strings for each of the 65 guitars.
Download Robust Recovery of Deterministic Timecode Signals Under Analog Degradation This paper presents a controlled evaluation framework for recovering deterministic stereo timecode waveforms used in digital vinyl systems. Three lightweight decoding policies are compared: Fixed-Threshold Symbol Decoder (FTSD), Adaptive-Threshold Symbol Decoder (ATSD), and State-Constrained Temporal Decoder (SCTD). The work provides matched comparison and explicit separation of availability, reference-relative correctness, and temporal behavior across noise, dropout, and clipping conditions.
Download Winding Numbers and Monodromy of Vector Bundles over a Circular Buffer The Möbius strip is perhaps the most recognizable topological object of general knowledge. It can be described mathematically in various ways including the formalism of line bundles. In this paper we discuss the bundle idea in the context of digital processing over a circular buffer and show how the idea leads to a more general notion known as monodromy, which describes the effect of the space on traversing a circle once. In this formulation, the monodromy of the Möbius strip is an orientation inversion characterized by a change in sign. This in turns leads to the concept of the winding number, which describes how many windings it takes to return to the original state. We discuss variable monodromy and illustrate that the winding number is robust under this variation. This will allow us to interpret previous disparate results in audio signal processing from Möbius waveguides to chaotic oscillators in delay loops in one unified framework. We close by showing how extending from line to vector bundles opens up the notion of braids to describe monodromy.
Download Quantifying Nonlinear Behavior in Digital Moog Ladder Filters: Cross-Implementation Comparison and Common-Core Ablation Digital Moog ladder filters are often compared by their linear frequency response, even though musically important differences emerge under nonlinear operation. This paper presents an open, reproducible SPICE-referenced evaluation framework that combines two components: a cross-implementation comparison of five digital ladder-filter implementations against the same SPICE reference, and a controlled common-core ablation. The cross-implementation comparison shows that close linear agreement can mask substantial nonlinear divergence, while the ablation shows that ladder nonlinear behavior depends not only on saturator choice and nonlinearity placement, but also on the non-uniform contribution of different ladder stages: in partial ablations of a SPICE-referenced TPT/ZDF core, retaining earlier-stage nonlinearities preserves harmonic behavior better than retaining later-stage nonlinearities alone. Beyond the specific comparisons reported here, the framework provides a reproducible basis for evaluating fidelity–cost trade-offs when nonlinear structure must be simplified.