Download Diffusion-Based Music Audio Editing System Using Differentiable Digital Signal Processing Mixture Model This paper proposes a music audio editing system that enables source-wise editing of harmonic instrument mixtures without explicit source separation. It builds on our previously proposed score-informed method for estimating source-wise synthesis parameters, i.e., time-varying controls used to synthesize each source, such as fundamental frequency and loudness. The method directly estimates these parameters from a mixture signal and the corresponding musical score in an analysis-by-synthesis framework. Using the estimated parameters, the proposed system allows users to edit individual sources by modifying note sequences and instrument types, and then re-synthesizes the edited mixture. Through demonstrations on two-instrument mixtures, we show that the system supports note-level phrasing modification and instrument conversion of selected sources.
Download SEND: A Spatial Event Neural Detector for Intentional Object Motion in Immersive Music Mixing Deciding exactly when to move audio objects in immersive mixes is a labor-intensive artistic task. Current tools react strictly to instantaneous frequency overlaps, lacking the macroscopic awareness required for musically intentional spatial transitions. To model these decisions, we propose SEND (Spatial Event Neural Detector). Its dual-stream architecture analyzes the target track against its background context, combining a Spec-TNT backbone and a Temporal Convolutional Network (TCN) to capture hierarchical spectral features and precise rhythmic cues. Their dynamic interplay is modeled via a novel Cross-Track Gating Interaction (CTGI) mechanism.
Download Fourier Neural Operators for Sample-Rate-Independent Virtual Analog Modeling Neural networks that operate directly on time-domain signals are widely used for virtual analog (VA) modeling. A key limitation of these models is their dependence on the sampling rate used during training, which becomes implicitly encoded in the learned parameters, so that changing it generally alters the realized dynamics. Although architectural modifications to recurrent neural networks have been proposed to enable sample-rate independent operation, these approaches are inherently tailored to upsampling and do not accommodate downsampling scenarios. In this manuscript, we present a VA modeling framework based on Fourier Neural Operators (FNOs) adapted to process fixed-duration audio frames. The proposed formulation defines the learned mapping over a fixed temporal support and evaluates it on uniform grids of different densities, so that a model trained at a single sampling rate can be applied at unseen sampling resolutions. Numerical results on a nonlinear transistor circuit show that the proposed model achieves competitive accuracy in upsampling scenarios while remaining directly applicable to downsampling, unlike a sample-rate independent baseline recurrent architecture.
Download Enhancing Automatic Chord Recognition via Pseudo-Labeling and Knowledge Distillation Automatic Chord Recognition (ACR) is constrained by the scarcity of aligned chord annotations, which are costly to acquire. At the same time, open-weight pre-trained models are more accessible than their proprietary training data. In this work, we present a two-stage training pipeline that leverages pre-trained models together with unlabeled audio. The proposed method decouples training into two stages. In the first stage, we use the pre-trained BTC model as a teacher to generate pseudo-labels for over 1,000 hours of diverse unlabeled audio and train a student model solely on these pseudo-labels. In the second stage, the student is continually trained on ground-truth labels as they become available. To prevent catastrophic forgetting of the representations learned in the first stage, we apply selective knowledge distillation (KD) from the teacher as a regularizer. In our experiments, two models (BTC, 2E1D) were used as students. In Stage 1, using only pseudo-labels, the BTC student achieves about 99% of the teacher's performance, while the 2E1D model achieves about 97% of the teacher's performance across seven standard mir_eval metrics. After continual training with labeled data in Stage 2, the resulting BTC student model consistently surpasses both the traditional supervised learning baseline and the original pre-trained teacher model across all metrics. The resulting 2E1D student model also outperforms the supervised baseline and approaches teacher-level performance, with both models demonstrating substantial gains on rare chord qualities.
Download FM Synthesizer Audio-Parameter Shared Embeddings Given a target sound, finding the synthesizer preset that best reproduces it remains a core problem in sound design. Existing methods treat synthesis parameters as flat vectors, discarding the signal routing and parameter interactions that produce audio. We make two contributions. First, to learn a representation of parameters including their signal routing, we design a graph neural network whose message passing structure imitates FM signal processing. Second, we adapt the multimodal objective from SLAP to learn joint embeddings of audio and FM synthesizer parameters, enabling preset retrieval from a gallery. We focus on the Yamaha DX7, where six identical sinusoid operators interact according to one of 32 routing topologies. Our graph encoder's message passing weights are shared across all nodes and layers, enabling processing of arbitrary topologies of any size. When every topology is seen during training, the DX7-GNN and two baselines achieve strong audio-to-preset retrieval. When some topologies are held out for testing, the DX7-GNN substantially outperforms both baselines despite having the fewest parameters. Our ablations further support the claim that imitating FM signal flow in a parameter encoder improves generalization to unseen topologies.
Download Audio-to-Audio via Diffusion Warm Initialization In this paper, we propose diffusion warm initialization as a simple yet effective approach for a range of audio-to-audio transformation tasks. To illustrate the generality of the approach, we demonstrate its use in timbre transfer, MIDI-to-Real synthesis, and multiple audio enhancement tasks. We conduct a detailed empirical analysis on timbre transfer to investigate the role of the initialization time t_init. The effect of t_init is evaluated using pitch-based Jaccard Distance and Fréchet Audio Distance to quantify faithfulness to the input signal and alignment with the target distribution. Our results provide practical guidance for selecting t_init and show that, once properly chosen, a single pretrained diffusion model combined with warm initialization can support multiple transformation objectives without task-specific training or conditioning. Despite its simplicity, this approach already achieves competitive results when compared with more complex pipelines designed specifically for these tasks.
Download Parameter Estimation via Differentiable Modal Plate Synthesis We present our submission to Task A of the 1st DAFx Parameter Estimation Challenge, which concerns the estimation of the physical parameters of a vibrating plate from a synthetic impulse response. Our approach introduces a differentiable modal plate synthesizer and estimates the plate parameters through inference-time gradient-based optimization of the synthesizer parameters. The six target parameters are recovered by minimizing a multi-scale spectral loss via backpropagation through the differentiable plate model. To handle the non-convexity of the loss landscape, we adopt a two-phase training strategy consisting of multiple short-term probe optimizations, followed by full-scale refinement initialized from the best candidate. We evaluate the approach on eight impulse responses synthesized with the official challenge dataset generator. Compared with a constant-value predictor and the particle swarm optimization baseline provided by the challenge, the proposed method reduces the prediction error by approximately one order of magnitude.
Download Benchmarking Integrated GPU Acceleration of Real-Time Neural Audio Inference on Snapdragon This paper investigates whether integrated GPUs on Qualcomm Snapdragon SoCs can accelerate streaming inference of neural audio models. Five models spanning three orders of magnitude in parameter count are benchmarked across three inference approaches (best available CPU, QNN CPU, and QNN GPU). Results reveal when GPU acceleration offers meaningful gains, when per-call overhead negates benefits, and how model size and architecture determine GPU suitability.
Download Quality Audio Prototyping: A Prototype System for Unified Sound Retrieval and Procedural Generation This paper presents Quality Audio Prototyping (QAP), a unified prototype system for sound retrieval and procedural generation. The system is designed to support rapid exploration of sound effects through a common interface that combines retrieval from existing audio collections with controllable procedural synthesis. By bringing these two paradigms together, QAP allows users to search for recorded sounds, generate new material, and iteratively refine results within a single workflow. The prototype emphasizes usability, extensibility, and practical sound-design applications, providing a foundation for future work on integrated retrieval and generation systems.
Download A DDSP Framework for Adaptive Room Equalization Adaptive room equalization remains challenging under time-varying acoustic conditions and complex excitation signals, such as music. In these scenarios, classical filtered-x least mean squares (Fx-LMS) methods falter due to their rigid formulation. We present a modular differentiable digital signal processing (DDSP) framework for closed-loop adaptive room equalization that recovers Fx-LMS as a special case through automatic differentiation. The framework supports interchangeable EQ structures, response estimation methods, loss functions, and optimizers. Experiments with time-varying measured room impulse responses show that frequency-domain objectives provide more stable adaptation than time-domain objectives in the considered scenarios. Relative to the non-equalized response, system distance is reduced by 70% and mel-spectral distance by 13% (worst-case scenario). We further examine how online room response estimation accuracy and frame length affect the trade-off between responsiveness and convergence stability. Overall, the framework provides a unified open-source basis for exploring synergies between classical adaptive filtering and DDSP-based optimization.