Download Multi-Source Extension and Hyperparameter Optimization of the DiffRIR Framework for Room Impulse Response Synthesis Efficient prediction of Room Impulse Responses (RIRs) is a cornerstone for immersive virtual acoustics and scalable room acoustic modeling. This study extends the DiffRIR framework – proposed by Wang et al. in Hearing Anything Anywhere – by introducing a multi-source training logic and systematically optimizing its convergence behavior to overcome the inherent limitations of the original framework. Our results reveal that multi-source training acts as implicit data augmentation, where the resulting increase in spatial entropy enhances the model's spectral accuracy. Furthermore, we demonstrate that the model exhibits remarkable robustness against geometric inaccuracies, maintaining numerical stability even with source positional offsets of up to 4 m in single-source baseline evaluations. By identifying a learning rate of 3×10⁻², we were able to reduce the training duration to 23% of the original baseline without compromising prediction accuracy. While the increased complexity of multi-source fields necessitates a trade-off in temporal precision – quantified via our newly integrated Energy Decay Convergence (EDC) metric – this research provides an efficient and resilient solution for acoustic simulations in complex environments.
Download Transformer-Based Plate Parameter Estimation with Differentiable and Particle-Swarm Refinement We present two Transformer-based methods for Task A of the 1st DAFx Parameter Estimation Challenge, which requires estimating six effective physical parameters of a synthetic plate-reverb model from its impulse response (IR). Method A1 combines an Audio Spectrogram Transformer encoder and Transformer regressor with differentiable IR refinement. Method A2 uses the same encoder to condition a continuous normalizing flow and refines sampled candidates using particle swarm optimisation (PSO) and gradient polishing. Both methods preserve the absolute IR scale to recover surface density. On a synthetic holdout set of 100 IRs, both refinement procedures reduce waveform and parameter errors by more than three orders of magnitude relative to the unrefined neural outputs. The PSO-based pipeline achieves the lowest errors, indicating near-perfect recovery in this matched synthetic setting.
Download Pulsetable Synthesis of Wind Instrument Tones We revisit pulsetable synthesis, an efficient technique for generating plausible and expressive wind instrument tones. Based on the principles of pulse forming theory, this method models sound production as the periodic repetition of shaped pulses characterizing the target instruments' spectral envelope. In this approach, single-cycle waveforms, referred to as pulses, are stored in pulsetables indexed by their corresponding fundamental frequency. During synthesis, the pulses are read from these tables to form a periodic waveform, which is further shaped by time-varying low-pass filtering, amplification, and reverberation. These processes are guided by control signal contours that describe how fundamental frequency, brightness, and loudness evolve over time. Through case studies with real-world wind instrument recordings, we show how the interplay between these control signals gives rise to articulations such as attack transients, vibrato, and growl. Finally, we discuss the potential of this framework for integration into Differentiable Digital Signal Processing (DDSP) models, where neural networks could learn synthesis parameters directly from training data.
Download Quality Audio Prototyping: A Prototype System for Unified Sound Retrieval and Procedural Generation This paper presents Quality Audio Prototyping (QAP), a unified prototype system for sound retrieval and procedural generation. The system is designed to support rapid exploration of sound effects through a common interface that combines retrieval from existing audio collections with controllable procedural synthesis. By bringing these two paradigms together, QAP allows users to search for recorded sounds, generate new material, and iteratively refine results within a single workflow. The prototype emphasizes usability, extensibility, and practical sound-design applications, providing a foundation for future work on integrated retrieval and generation systems.
Download A DDSP Framework for Adaptive Room Equalization Adaptive room equalization remains challenging under time-varying acoustic conditions and complex excitation signals, such as music. In these scenarios, classical filtered-x least mean squares (Fx-LMS) methods falter due to their rigid formulation. We present a modular differentiable digital signal processing (DDSP) framework for closed-loop adaptive room equalization that recovers Fx-LMS as a special case through automatic differentiation. The framework supports interchangeable EQ structures, response estimation methods, loss functions, and optimizers. Experiments with time-varying measured room impulse responses show that frequency-domain objectives provide more stable adaptation than time-domain objectives in the considered scenarios. Relative to the non-equalized response, system distance is reduced by 70% and mel-spectral distance by 13% (worst-case scenario). We further examine how online room response estimation accuracy and frame length affect the trade-off between responsiveness and convergence stability. Overall, the framework provides a unified open-source basis for exploring synergies between classical adaptive filtering and DDSP-based optimization.
Download From Arbitrary Audio to EDM: Audio-Conditioned Retrieval of Discrete Rhythm Archetypes We present a system for transforming arbitrary audio into Electronic Dance Music (EDM) drum patterns while preserving the timbral identity of the source material. A Vector Quantized Variational Autoencoder (VQ-VAE) trained on 7,999 EDM drum loops learns a discrete codebook of 256 rhythm archetypes, validated through UMAP and hierarchical clustering to exhibit semantically meaningful structure. At inference, spectral features extracted from arbitrary input audio select the nearest archetype via nearest-neighbor retrieval in a shared audio feature space. A training sample from the selected archetype is reconstructed through the VQ-VAE, and a second decoder predicts per-hit velocity dynamics. The user's sounds are then placed at the reconstructed hit positions, scaled by predicted velocity. Applied to 2,000 files from the ESC-50 environmental sound dataset, the system activates 128 of 256 codebook entries (50% coverage), demonstrating broad responsiveness to diverse non-EDM audio.
Download Parameter Estimation via Differentiable Modal Plate Synthesis We present our submission to Task A of the 1st DAFx Parameter Estimation Challenge, which concerns the estimation of the physical parameters of a vibrating plate from a synthetic impulse response. Our approach introduces a differentiable modal plate synthesizer and estimates the plate parameters through inference-time gradient-based optimization of the synthesizer parameters. The six target parameters are recovered by minimizing a multi-scale spectral loss via backpropagation through the differentiable plate model. To handle the non-convexity of the loss landscape, we adopt a two-phase training strategy consisting of multiple short-term probe optimizations, followed by full-scale refinement initialized from the best candidate. We evaluate the approach on eight impulse responses synthesized with the official challenge dataset generator. Compared with a constant-value predictor and the particle swarm optimization baseline provided by the challenge, the proposed method reduces the prediction error by approximately one order of magnitude.
Download Enhancing Automatic Chord Recognition via Pseudo-Labeling and Knowledge Distillation Automatic Chord Recognition (ACR) is constrained by the scarcity of aligned chord annotations, which are costly to acquire. At the same time, open-weight pre-trained models are more accessible than their proprietary training data. In this work, we present a two-stage training pipeline that leverages pre-trained models together with unlabeled audio. The proposed method decouples training into two stages. In the first stage, we use the pre-trained BTC model as a teacher to generate pseudo-labels for over 1,000 hours of diverse unlabeled audio and train a student model solely on these pseudo-labels. In the second stage, the student is continually trained on ground-truth labels as they become available. To prevent catastrophic forgetting of the representations learned in the first stage, we apply selective knowledge distillation (KD) from the teacher as a regularizer. In our experiments, two models (BTC, 2E1D) were used as students. In Stage 1, using only pseudo-labels, the BTC student achieves about 99% of the teacher's performance, while the 2E1D model achieves about 97% of the teacher's performance across seven standard mir_eval metrics. After continual training with labeled data in Stage 2, the resulting BTC student model consistently surpasses both the traditional supervised learning baseline and the original pre-trained teacher model across all metrics. The resulting 2E1D student model also outperforms the supervised baseline and approaches teacher-level performance, with both models demonstrating substantial gains on rare chord qualities.
Download Explicit Wave Digital Model of the Fulltone OCD Pedal Based on Canonical Piecewise-Linear Functions Virtual Analog (VA) modeling aims at digitally emulating analog audio equipment while preserving its characteristic nonlinear behavior and musical expressiveness. In the context of guitar effects, overdrive pedals represent a cornerstone of many signal chains, as they strongly contribute to the perceived dynamics, articulation, and timbral identity of the instrument. Among these, the Fulltone OCD overdrive is considered a standard in both studio and live environments, being widely adopted across rock and metal genres. In this article, we present an explicit Wave Digital (WD) model of the Fulltone OCD (v2) pedal. By exploiting the circuit topology, the MOSFETs and the germanium diode composing the asymmetric clipping stage are grouped into a single equivalent nonlinear element, enabling an explicit WD realization that avoids costly iterative solvers. The resulting nonlinear characteristic is approximated by means of a Canonical Piecewise-Linear (CPWL) function, yielding a compact and efficient explicit model suitable for real-time implementation. The proposed model is validated against reference simulations and implemented both in MATLAB and as a real-time audio plug-in using the JUCE framework.
Download Sound Matching with a Differentiable Karplus-Strong Algorithm We present a self-supervised, event-based sound matching model using a differentiable extended Karplus-Strong algorithm. To avoid relying on external onset and fundamental frequency detectors, we explore training methodologies combining parameter losses on synthetic data with audio losses. We demonstrate that time-domain fractional delay interpolation provides gradient accuracy comparable to frequency-sampling while avoiding time-aliasing in highly resonant time-varying scenarios. Through systematic gradient analysis, we reveal that standard spectral losses provide no meaningful directional gradients for onset times, heavily degrading joint training. Training exclusively with parameter losses on synthetic data effectively learns fundamental frequency, timbral parameters, and onset times, but struggles to generalise to monophonic studio recordings of plucked guitar. External detectors combined with audio losses generalise best, isolating the model to timbre optimisation. While our Karplus-Strong decoder recovers interpretable parameters and naturally captures the transient characteristics of plucked guitar, Harmonics plus Noise baselines yield higher reconstruction fidelity by most metrics.