Download A Perceptually Inspired Single Parameter Auditory Distance Renderer for Music Production Conveying auditory distance in a digital audio workstation requires balancing several uncoupled tools (reverb, gain, equalization, pre-delay) by hand, a workflow that is cognitively demanding and easily produces spatially incoherent results. We demonstrate a real-time VST3 plugin that derives five correlated distance cues from a single normalized control and keeps them mutually coherent by construction, grounded in the psychoacoustics of auditory distance perception. A headphone listening test with 18 participants showed that the coupled renderer roughly halves distance placement error relative to an uncoupled manual mix and was unanimously preferred on composite spatial quality. Users sweep one knob and hear sources move convincingly from near to far on multitrack material, compare the result against a manual uncoupled mix, and toggle an optional binaural externalization stage. Source code is available on GitHub.
Download A Corpus-Driven Parametric Modal Reverberator A parametric modal reverberator is presented in which synthesis parameters are derived from a large, curated corpus of room impulse responses (IRs). The collected responses are subjected to modal decomposition, yielding per-mode frequencies, damping coefficients, and residue amplitudes, together with a short early-reflection finite impulse response (FIR) filter. From the decomposed data, a feature table is constructed per IR comprising standard acoustic indices, per-band damping and density statistics, amplitude distributions, and FIR descriptors—50 variables in total. Six acoustically meaningful user controls are selected; since these exhibit substantial pairwise correlations across the corpus, they are orthogonalised via principal component analysis (PCA) prior to regression.
Download Sound Matching with a Differentiable Karplus-Strong Algorithm We present a self-supervised, event-based sound matching model using a differentiable extended Karplus-Strong algorithm. To avoid relying on external onset and fundamental frequency detectors, we explore training methodologies combining parameter losses on synthetic data with audio losses. We demonstrate that time-domain fractional delay interpolation provides gradient accuracy comparable to frequency-sampling while avoiding time-aliasing in highly resonant time-varying scenarios. Through systematic gradient analysis, we reveal that standard spectral losses provide no meaningful directional gradients for onset times, heavily degrading joint training. Training exclusively with parameter losses on synthetic data effectively learns fundamental frequency, timbral parameters, and onset times, but struggles to generalise to monophonic studio recordings of plucked guitar. External detectors combined with audio losses generalise best, isolating the model to timbre optimisation. While our Karplus-Strong decoder recovers interpretable parameters and naturally captures the transient characteristics of plucked guitar, Harmonics plus Noise baselines yield higher reconstruction fidelity by most metrics.
Download Quality Audio Prototyping: A Prototype System for Unified Sound Retrieval and Procedural Generation This paper presents Quality Audio Prototyping (QAP), a unified prototype system for sound retrieval and procedural generation. The system is designed to support rapid exploration of sound effects through a common interface that combines retrieval from existing audio collections with controllable procedural synthesis. By bringing these two paradigms together, QAP allows users to search for recorded sounds, generate new material, and iteratively refine results within a single workflow. The prototype emphasizes usability, extensibility, and practical sound-design applications, providing a foundation for future work on integrated retrieval and generation systems.
Download Pulsetable Synthesis of Wind Instrument Tones We revisit pulsetable synthesis, an efficient technique for generating plausible and expressive wind instrument tones. Based on the principles of pulse forming theory, this method models sound production as the periodic repetition of shaped pulses characterizing the target instruments' spectral envelope. In this approach, single-cycle waveforms, referred to as pulses, are stored in pulsetables indexed by their corresponding fundamental frequency. During synthesis, the pulses are read from these tables to form a periodic waveform, which is further shaped by time-varying low-pass filtering, amplification, and reverberation. These processes are guided by control signal contours that describe how fundamental frequency, brightness, and loudness evolve over time. Through case studies with real-world wind instrument recordings, we show how the interplay between these control signals gives rise to articulations such as attack transients, vibrato, and growl. Finally, we discuss the potential of this framework for integration into Differentiable Digital Signal Processing (DDSP) models, where neural networks could learn synthesis parameters directly from training data.
Download Shimmer Reverberation with Nonlinear Feedback Delay Networks Shimmer reverberation is an effect used in music production to deliver ethereal, pitch-shifted textures and evolving ambient soundscapes. This paper explores the synthesis of shimmer effects using the feedback delay network architecture, a popular real-time reverberator. We propose five distinct approaches for integrating nonlinear and time-varying operations into the feedback loop, focusing on expanding the harmonic content while adhering to energy-preservation and stability criteria. Our approach can generate a wide range of sonic characteristics, from harmonically rich distortions to musically coherent pitch-shifted reverberation, while maintaining stability and controllable decay behavior.
Download Eigensystem Realization of Violin Bridge Admittances Modeling violin bridge admittance is a long-standing problem in musical acoustics, with applications in sound analysis, synthesis, and virtual instrument design. In this work, we investigate the use of the Eigensystem Realization Algorithm (ERA) for deriving reduced-order state-space models directly from measured impulse responses. The proposed approach allows us to extract dominant system dynamics and obtain compact realizations without requiring explicit modal parameterization. We evaluate ERA on a dataset of modern and historical violins and compare it against established modal and state-space identification methods. Experimental results demonstrate that ERA outperforms existing approaches by achieving lower reconstruction errors in both the time and frequency domains while preserving perceptually relevant characteristics of the bridge response. Furthermore, we show that the state-space realizations obtained using ERA reproduce the target frequency-dependent energy decay more accurately than models obtained using the baseline methods. These findings support the use of ERA as an efficient and flexible alternative for modeling violin bridge admittances, with applications that span from audio synthesis and processing to instrument virtualization.
Download Peak-Residual Modal Estimation with Learned Calibration and High-Band Density Correction ★ This paper describes two related submissions to Task B of the 1st DAFx Parameter Estimation Challenge. Both estimate modal frequency, decay, and gain directly from an unnormalised plate impulse response without using plate parameters, the excluded analytical modal-frequency law, or official-test ground truth. The primary system constructs a large candidate pool through prominence-graded spectral peak picking, iterative residual analysis, multi-view consensus, band-wise budgeting, and a learned file-level mode-count target. Raw decay and gain estimates are then corrected by a small mode-wise neural network that is not allowed to move frequencies or change the number of rows. A secondary variant addresses suspected high-frequency under-counting with a separately gated, non-oracle density-fill stage in the 6–10 kHz band. The paper reports development diagnostics, reproducibility information, and descriptive statistics for the 16 official outputs. The two variants expose a deliberate precision–recall trade-off: one preserves a visible spectral justification for every row, while the other tests bounded hidden-multiplicity augmentation in densely overlapped regions.
Download Parametric Resynthesis of Measured Spatial Room Impulse Responses Spatial Room Impulse Responses (SRIRs) are fundamental to immersive audio rendering and have become a key focus of recent machine learning research in acoustics and auralization. Due to the high computational cost of direct convolution, spatial audio systems commonly employ artificial reverberation algorithms. However, these approaches often fail to accurately reproduce the spatial, temporal, and spectral characteristics of early reflections, leading to notable deviations from measured SRIRs. This paper presents a comprehensive framework for the analysis and efficient resynthesis of SRIRs captured with Spherical Microphone Arrays (SMAs). The proposed method accounts for hardware-induced artifacts, including scattering and spatial aliasing. Early reflections are reconstructed using a parametric approach based on the Herglotz analysis method, while late reverberation is synthesized using a Directional Feedback Delay Network (DFDN) with optimized filter-attenuation and correlation-matching. The proposed framework produces signals whose spatial correlation and Energy Decay Relief (EDR) closely match those of measured SRIRs, demonstrating its effectiveness for both real-time spatial audio rendering and realistic dataset generation for machine learning applications.
Download IRIS: Continuous Spatial Navigation of Measured Acoustic Fields via Impulse Response Interpolation Impulse response (IR) collections are useful in virtual acoustics, sound design, and field-based acoustic research, but they remain difficult to explore as continuous resources in lightweight real-time plugin workflows. This demo paper presents IRIS, a VST3 plugin for arranging, navigating, and auditioning measured or user-defined IR collections in a two-dimensional navigation plane. Each IR is represented as a node whose position can be imported from metadata or assigned manually. During navigation, nearby responses are combined using Gaussian distance-based weighting, while a bounded active set limits the number of simultaneous convolutions. The system also includes smoothing, hysteresis, optional preprocessing, boundary attenuation, OSC control, and coupled multichannel handling. The demo focuses on workflow and audible behavior rather than perceptual validation. A short timing characterization reports practical real-time limits as a function of IR length, buffer size, and active-set size. IRIS is presented as a practical tool for exploratory, analytical, and creative navigation of IR collections rather than as a physically optimal interpolation method.