Download Deep Regularized RNNs for Virtual Analog Virtual analog (VA) modeling methods seek to emulate analog audio hardware using digital signal processing (DSP). Modeling approaches fall into three broad categories: white-box methods, which use detailed device knowledge for accurate simulation; gray-box methods that use generic DSP blocks to model the system; and black-box methods, which rely solely on opaque models learned from input–output data. A category of architectures used widely in black-box modeling are recurrent neural networks (RNNs). To model device controls, the control values can be provided as conditioning input to the network. However, when the conditioning is time-varied, the models are susceptible to producing noise artifacts. Regularization of the RNN dynamics significantly reduces these artifacts, though at a loss in modeling accuracy. This paper closes the dynamics regularization quality gap by introducing deep control-conditioned LSTMs and a gammatone filterbank (GFB) loss. Experiments indicate that the proposed method achieves comparable modeling performance as unregularized baselines while avoiding the noise artifacts caused by time-varying control inputs.
Download PAEDB: A Synthetic Primary-Ambient Dataset Generation Pipeline for Automatic Upmixing Using Deep Neural Networks Automatic blind upmixing aims to convert audio from a smaller channel format (e.g. mono or stereo) into a multichannel format using estimates of direct and diffuse spatial statistics within the signal. Current approaches rely on primary-ambient extraction (PAE) algorithms, which lack real-world context through limited processing windows. Deep learning music source separation (MSS) models have been applied in voice-primary-ambient extraction (VPA) upmixing systems for handling direct components, but still rely on DSP methods of surround channel generation. This work further investigates utilizing source separation within VPA upmixing, focusing specifically on the task of stereo decorrelation and ambience extraction for 5.1 surround. We also release PAEDB (Primary–Ambient Extraction Dataset), a high-quality music dataset derived from MUSDB18-HQ and MoisesDB, comprising 1,809 primary–ambient stem pairs totaling over 550 hours of audio. The performance of selected DNNs trained on PAEDB is then evaluated using signal metrics and a listening study. Our findings indicate that DNNs can effectively model the behavior of PAE algorithms, establishing PAEDB as a strong foundation for ML upmixing systems and underscoring the need for higher-quality multichannel data to advance beyond conventional methods.
Download Diagonal Complex-Valued State Space Models for System Identification and Modeling of Metal Plate Reverbs Accurate and interpretable modeling of plate reverbs remains an important challenge in virtual analog modeling of audio effects. While existing neural network-based black-box approaches already achieve high-quality synthesis and strong perceptual quality, they often lack the possibility to identify the underlying physically meaningful complex, long-memory modal behavior. In this work, we address this limitation by proposing a restricted complex-valued diagonal State Space Model (SSM), showing its equivalence to a parallel second-order all-pole filter, also utilizing efficient training via parallel state computation using the parallel scan algorithm. Additionally, we propose a Matrix Pencil (MP) guided eigenvalue initialization, improving synthesis quality and system identification performance.
Download WildFX: A DAW-Powered Pipeline for In-the-Wild Audio FX Graph Modeling We present WildFX, a digital-audio-workstation-powered pipeline for modeling audio-effects graphs from in-the-wild audio. The system uses a DAW environment to construct, render, and evaluate effect-processing graphs, enabling research on realistic effect chains beyond isolated processors or synthetic training settings. WildFX supports the analysis and reconstruction of complex audio transformations by combining flexible plugin routing with data-driven modeling. The pipeline is designed to facilitate scalable dataset creation and experimentation with effect graph inference, parameter estimation, and audio transformation in practical production contexts.
Download CLEAN2FX: Label-Conditioned Modeling for Clean-to-Effect Guitar Audio Transformations We present Clean2FX, a study and demo of label-conditioned clean-to-effect transformation for electric guitar audio. Given a clean guitar input and a target effect label, the task is to synthesize the corresponding effected signal while preserving the musical content. Training and evaluation pairs are constructed from EGFxSet real, single-tone recordings by assembling matched clean/effected chords, melodies, and mixed timelines. This allows for controlled comparison across effects. We evaluate four neural approaches under a common spectrogram-based transformation setting: two variational autoencoders and two U-Net models that differ in whether they operate on linear or log-magnitude representations. Performance is measured using linear-magnitude spectrogram MSE and Fréchet Audio Distance. The U-Net models outperform the variational autoencoder variants. Per-effect results show that distortion effects are most readily improved, whereas delay and reverb effects exhibit weaker FAD gains despite substantial spectral-error reductions. A conditioning-sensitivity diagnostic provides evidence that the best model responds to target labels rather than collapsing to a single transformation. Our demo website compares two models applied on real-world guitar performances outside training and validation data, providing audio and spectrogram examples of the practical clean-to-effect behavior.
Download WaveNet-Style Guitar Amplifier Model Pruning for Real-Time iOS Deployment WaveNet-style convolutional networks emulate tube amplifiers and distortion pedals with high fidelity, but their computational cost has confined them to desktops or dedicated DSP hardware. We present a sparse-enabled WaveNet inference engine for iOS that runs heavily pruned neural guitar amplifier models in real time on iPhones. Aggressive iterative magnitude pruning removes 90% of the network weights with no perceptible loss in quality. A custom sparse C++ engine turns this sparsity directly into compute savings, sustaining low-latency real-time operation on a CPU-only iPhone implementation where the dense model cannot. On-device output matches the trained model to within int16 quantization error. At the demonstration, visitors will play a guitar through the app on iPhone hardware and A/B the on-device pruned model against the physical pedal it emulates. Source code and audio examples are available online.
Download Multi-Source Extension and Hyperparameter Optimization of the DiffRIR Framework for Room Impulse Response Synthesis Efficient prediction of Room Impulse Responses (RIRs) is a cornerstone for immersive virtual acoustics and scalable room acoustic modeling. This study extends the DiffRIR framework – proposed by Wang et al. in Hearing Anything Anywhere – by introducing a multi-source training logic and systematically optimizing its convergence behavior to overcome the inherent limitations of the original framework. Our results reveal that multi-source training acts as implicit data augmentation, where the resulting increase in spatial entropy enhances the model's spectral accuracy. Furthermore, we demonstrate that the model exhibits remarkable robustness against geometric inaccuracies, maintaining numerical stability even with source positional offsets of up to 4 m in single-source baseline evaluations. By identifying a learning rate of 3×10⁻², we were able to reduce the training duration to 23% of the original baseline without compromising prediction accuracy. While the increased complexity of multi-source fields necessitates a trade-off in temporal precision – quantified via our newly integrated Energy Decay Convergence (EDC) metric – this research provides an efficient and resilient solution for acoustic simulations in complex environments.
Download Neural Networks for Physical Parameter Estimation of Plate Reverberation from Impulse Responses ★ This paper presents our Task A submission to the 1st DAFx Parameter Estimation Challenge. We use the official ModalPlate dataset generator to synthesize 1000 one-second plate impulse responses with randomly sampled parameters inside the public ranges. A time-domain CNN-GRU regressor then estimates the six official Task A parameters from each unnormalised waveform. The model combines three one-dimensional convolutional blocks with a bidirectional gated recurrent unit and is trained with mean squared error on min-max normalised targets. The generated data are split into 700/150/150 train/validation/test examples, and the test split is never used during training or model selection. The implementation follows the official Task A format and exports evaluation-compatible prediction files for both development evaluation and blind-set submission.
Download Real-Time Neural Audio on Apple Silicon: Benchmarking Inference Frameworks Under Realistic DAW Contention Neural network models are increasingly deployed in audio plugins across a wide range of applications, including amplifier emulation, effects modeling, and synthesis. This paper evaluates widely used inference options including BNNSGraph, RTNeural, LibTorch, ONNX Runtime, and anira on model architectures commonly used in neural audio plugins. The key contribution is moving beyond isolated benchmarks to evaluate performance under realistic DAW contention, constructing mix sessions with configurable plugin loads. Results show that isolated benchmarks can be misleading, and BNNSGraph proves most robust for convolutional models on Apple Silicon.
Download FoleySet: A Multi-Level Human-Annotated Foley Sound Dataset We introduce FoleySet, a human-annotated Foley sound dataset designed to support research on sound-event understanding and Foley sound generation. The dataset provides annotations at multiple levels of granularity, capturing both broad event categories and more detailed semantic or production-related attributes. This multi-level structure supports tasks such as classification, retrieval, captioning, and controllable generation. FoleySet is intended to address the limited availability of systematically annotated Foley material and to provide a common resource for evaluating models across different levels of semantic detail.