Download Perceptual Optimisation of Loudspeaker-Based Reproduction
This paper proposes POLAR, a framework for the optimisation of loudspeaker signals using end-to-end differentiable perceptual loss functions. The framework optimises multiple perceptual attributes across multiple listeners, offering a versatile method for a range of problems. This versatility stems from the ability to customise the number of loudspeakers, listeners, and the weighting applied to different perceptual attributes. Here, we apply the method to four problems: (a) source panning for a single listener in stereo reproduction, (b) single-listener colouration matching in stereo reproduction, (c) extended sweet spot using stereo pairs beamforming, and (d) multi-attribute perceptually driven panning in stereo reproduction. The first three problems are evaluated against solutions traditionally used for these tasks: solutions of (a) are shown to be similar to those obtained with tangent panning law and vector-base amplitude panning (VBAP), solutions of (b) are shown to be similar to those obtained for cross-talk cancellation, and solutions of (c) are shown to be similar to those obtained in earlier work on directivity pattern optimisation for sweet spot widening. Each of these solutions was previously obtained using fundamentally different methodologies, demonstrating the flexibility and broad applicability of the proposed framework.
Download FM Parameter Estimation with Low-Order Rational Constraints on Wasserstein Loss Landscape
Frequency modulation (FM) synthesis has been widely used in music production and sound design due to its ability to generate rich timbres with few control parameters. However, estimating the frequency parameters from a target sound remains challenging because different parameter configurations can yield similar spectra, creating numerous local minima in the loss landscape. In this paper, we analyze the Wasserstein distance loss landscape for two-operator FM synthesis under practical FFT-based spectral representations and show that it exhibits non-differentiable ridges at rational frequency ratios, arising from negative-frequency folding and spectral ordering transitions. Exploiting this structure, we propose a constrained gradient-based optimization strategy that constrains the frequency ratio in each optimization run to an interval bounded by consecutive low-order rational ratios and retains the lowest-loss candidate across intervals. Experimental results from controlled ablations show that maintaining the constraint throughout optimization improves reliability over random initialization and initialization-only constraints, particularly for more complex spectra at higher modulation indices.
Download Differentiable Articulatory Copy-Synthesis of Biphonic Singing
Sygyt is a Tuvan style of biphonic singing in which a low vocal drone is sustained while a high harmonic is selectively amplified in the 1–3 kHz region. Copy-synthesizing this effect remains challenging for articulatory models, since it requires fine control of narrowly focused resonances that standard low-dimensional tract parameterizations cannot easily reproduce. We address this problem with a differentiable Kelly–Lochbaum waveguide augmented with a sublingual second source, cubic B-spline tract parameterization, and spatially varying learnable damping, optimized end-to-end by gradient descent from audio. On 20 segments from two independent sygyt datasets (5 singers, 10 pitches), the proposed model reduces log-spectral distance by 30–38% relative to an articulatory baseline, with the largest gains concentrated in the overtone region. Cepstral-envelope analysis further shows more accurate recovery of the merged formant structure characteristic of sygyt production. The model also outperforms a DDSP harmonic-plus-noise baseline with direct per-harmonic spectral control, suggesting that explicit acoustic structure is a useful inductive bias for overtone-singing copy-synthesis.
Download Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion
Recent diffusion-based generative models have achieved strong results in domain-specific audio generation tasks such as speech, singing, and instrumental music synthesis. However, these models are typically specialized and do not generalize well to mixed or intermediate audio types. In this work, we adapt a diffusion-based model originally designed for multi-instrument music synthesis to voice conversion, covering both speech and singing within a unified framework. Specifically, we extend musical note-based conditioning to include phonetic posteriorgrams (PPGs) and pitch contours, and reinterpret timbre conditioning as speaker or singer identity via feature-wise linear modulation. Experiments show that the adapted model matches or surpasses a dedicated voice conversion system in terms of naturalness and performer similarity, while maintaining accurate pitch control across speech and singing. At the same time, we observe limitations in phonetic fidelity and a degradation in vocal quality when incorporating instrumental training data. Furthermore, we demonstrate that off-the-shelf feature extractors provide effective conditioning signals, enabling large-scale self-supervised training without manual annotations. These results highlight the potential of cross-domain model transfer towards unified audio generation systems capable of handling speech, singing, and music.
Download Probing Low-Level Acoustic Attribute Encoding in CLAP Audio Embeddings
This work analyzes CLAP audio embeddings through a probing framework, studying the encoding of reverberation (RT60), loudness (LUFS), spectral content (SC), and relative pitch (RP). Results show that all attributes are reliably recoverable from CLAP embeddings, with RT60, LUFS, and RP approximately linearly encoded, while SC requires non-linear probes. The identified patterns generalize across eight additional audio foundation models.
Download DAFx Challenge Introduction & Results
The 1st DAFx Parameter Estimation Challenge is an open initiative to advance the state of the art in parameter estimation for acoustic modeling. Stated as a system identification problem, this first edition focuses on plate reverberation—an archetypal dense, modal and weakly damped acoustic system. Participants tackled two tasks: (A) estimating the physical parameters of a vibrating plate from its impulse response, and (B) recovering the modal parameters of the same system. Both rest on a simulation framework based on the damped Kirchhoff–Love plate equation, and both are posed and scored entirely on synthetic data produced by that framework: no measurement of a real plate is involved. Two participants solved Task A down to machine precision by different strategies: one a neural network trained on a very large dataset, and one gradient-free optimization with many inexpensive evaluations. Task B proved considerably harder: the best submission attains a relative error of 0.33 on a [0, 2] scale, and every method recovers modal frequencies and decay rates far more accurately than modal gains. A complementary frequency-domain evaluation reorders the ranking and exposes a systematic gain bias to which the per-mode metric is blind.