Download Ambisonic Decoder Equalization in Reverberant Environments via Closed-Hull Crosstalk Inversion In this work, the authors develop a higher order Ambisonic (HOA) decoder that compensates for listening room reverberation by cascading a conventional decode matrix with a crosstalk matrix derived from room impulse responses (RIR) constrained over the listener area, namely by sampling over a spherical boundary according to HOA convention. A decoder is generated for simulated RIRs of mixed specular and diffuse reverberation, and spectral and spatiotemporal energy distributions are shown for directional and diffuse HOA signals rendered through the reverberant room. Results demonstrate that the decoder renders temporally-variant directional sources with higher directivity as compared to a conventional decoder.ing Ambisonic decoders in reverberant environments using closed-hull crosstalk convolution. By modeling the acoustic path from each loudspeaker to the listener's ears via measured room impulse responses, we formulate equalization filters that compensate for room-induced distortions in the decoded signals. The proposed closed-hull approach constrains the solution to preserve perceptually relevant spatial cues while minimizing spectral coloration. Listening tests demonstrate improved externalization, timbral transparency, and localization accuracy compared to standard free-field decoding in reverberant conditions.
Download Perceptual Optimisation of Loudspeaker-Based Reproduction This paper proposes POLAR, a framework for the optimisation of loudspeaker signals using end-to-end differentiable perceptual loss functions. The framework optimises multiple perceptual attributes across multiple listeners, offering a versatile method for a range of problems. This versatility stems from the ability to customise the number of loudspeakers, listeners, and the weighting applied to different perceptual attributes. Here, we apply the method to four problems: (a) source panning for a single listener in stereo reproduction, (b) single-listener colouration matching in stereo reproduction, (c) extended sweet spot using stereo pairs beamforming, and (d) multi-attribute perceptually driven panning in stereo reproduction. The first three problems are evaluated against solutions traditionally used for these tasks: solutions of (a) are shown to be similar to those obtained with tangent panning law and vector-base amplitude panning (VBAP), solutions of (b) are shown to be similar to those obtained for cross-talk cancellation, and solutions of (c) are shown to be similar to those obtained in earlier work on directivity pattern optimisation for sweet spot widening. Each of these solutions was previously obtained using fundamentally different methodologies, demonstrating the flexibility and broad applicability of the proposed framework.
Download Multi-Source Extension and Hyperparameter Optimization of the DiffRIR Framework for Room Impulse Response Synthesis Efficient prediction of Room Impulse Responses (RIRs) is a cornerstone for immersive virtual acoustics and scalable room acoustic modeling. This study extends the DiffRIR framework – proposed by Wang et al. in Hearing Anything Anywhere – by introducing a multi-source training logic and systematically optimizing its convergence behavior to overcome the inherent limitations of the original framework. Our results reveal that multi-source training acts as implicit data augmentation, where the resulting increase in spatial entropy enhances the model's spectral accuracy. Furthermore, we demonstrate that the model exhibits remarkable robustness against geometric inaccuracies, maintaining numerical stability even with source positional offsets of up to 4 m in single-source baseline evaluations. By identifying a learning rate of 3×10⁻², we were able to reduce the training duration to 23% of the original baseline without compromising prediction accuracy. While the increased complexity of multi-source fields necessitates a trade-off in temporal precision – quantified via our newly integrated Energy Decay Convergence (EDC) metric – this research provides an efficient and resilient solution for acoustic simulations in complex environments.
Download Parametric Resynthesis of Measured Spatial Room Impulse Responses Spatial Room Impulse Responses (SRIRs) are fundamental to immersive audio rendering and have become a key focus of recent machine learning research in acoustics and auralization. Due to the high computational cost of direct convolution, spatial audio systems commonly employ artificial reverberation algorithms. However, these approaches often fail to accurately reproduce the spatial, temporal, and spectral characteristics of early reflections, leading to notable deviations from measured SRIRs. This paper presents a comprehensive framework for the analysis and efficient resynthesis of SRIRs captured with Spherical Microphone Arrays (SMAs). The proposed method accounts for hardware-induced artifacts, including scattering and spatial aliasing. Early reflections are reconstructed using a parametric approach based on the Herglotz analysis method, while late reverberation is synthesized using a Directional Feedback Delay Network (DFDN) with optimized filter-attenuation and correlation-matching. The proposed framework produces signals whose spatial correlation and Energy Decay Relief (EDR) closely match those of measured SRIRs, demonstrating its effectiveness for both real-time spatial audio rendering and realistic dataset generation for machine learning applications.
Download Physical Model of the Chinese Yehu for Sound Synthesis The yehu is a Chinese bowed string instrument featuring a resonator carved from a coconut shell, a seashell-based bridge, and two silk strings. This paper proposes a physical model of the yehu and reports on simulations using a finite-difference scheme with measurement-based physical characterization. The proposed model consists of two stiff strings coupled at the bridge, a bow with elastic bow hairs, a stopping finger, and a modal model of the bridge. A non-iterative solver based on energy quadratization is used to model the finger–string contact force, while an iterative solver is used for elasto-plastic bow-string friction force. The bridge-body model is based on a modal characterization obtained from the measured bridge admittance. The measured radiation transfer function is represented as a bank of parallel second-order filters and is applied to the simulated bridge force to incorporate body radiation characteristics. Finally, computational performance tests are conducted, showing that the proposed model is capable of real-time computation.
Download Eigensystem Realization of Violin Bridge Admittances Modeling violin bridge admittance is a long-standing problem in musical acoustics, with applications in sound analysis, synthesis, and virtual instrument design. In this work, we investigate the use of the Eigensystem Realization Algorithm (ERA) for deriving reduced-order state-space models directly from measured impulse responses. The proposed approach allows us to extract dominant system dynamics and obtain compact realizations without requiring explicit modal parameterization. We evaluate ERA on a dataset of modern and historical violins and compare it against established modal and state-space identification methods. Experimental results demonstrate that ERA outperforms existing approaches by achieving lower reconstruction errors in both the time and frequency domains while preserving perceptually relevant characteristics of the bridge response. Furthermore, we show that the state-space realizations obtained using ERA reproduce the target frequency-dependent energy decay more accurately than models obtained using the baseline methods. These findings support the use of ERA as an efficient and flexible alternative for modeling violin bridge admittances, with applications that span from audio synthesis and processing to instrument virtualization.
Download Measurement-Informed Nonlinear Modal Synthesis of 65 Classical Guitars When a classical guitar string is plucked, vibration energy flows through the bridge into the body and is radiated as sound. Synthesising this process for a large collection of instruments requires both an efficient nonlinear string model and a robust method for extracting instrument-specific parameters from measurements. This paper addresses both issues. Starting from the publicly available dataset of Mores, which provides impulse-response measurements on 65 classical guitars, modal parameters of the bridge compliance and of the bridge-to-air radiation path are extracted for each instrument. These feed a nonlinear string model in which transverse vibration is governed by a geometrically exact elastic potential coupled at an interior bridge point to the measured body data. The nonlinear potential is quadratised via the Scalar Auxiliary Variable (SAV) method, so that the equations of motion become linear in a scalar variable and a known gradient vector, even at the continuous level. After time discretisation, the coupled system is inverted through two sequential Sherman–Morrison rank-one updates (one for the bridge coupling, one for the SAV nonlinearity), yielding an O(N) algorithm per time step. Two regularisation techniques prevent long-term drift of the auxiliary variable. The complete pipeline is demonstrated by synthesising plucked notes across all frets and strings for each of the 65 guitars.
Download Differentiable Articulatory Copy-Synthesis of Biphonic Singing Sygyt is a Tuvan style of biphonic singing in which a low vocal drone is sustained while a high harmonic is selectively amplified in the 1–3 kHz region. Copy-synthesizing this effect remains challenging for articulatory models, since it requires fine control of narrowly focused resonances that standard low-dimensional tract parameterizations cannot easily reproduce. We address this problem with a differentiable Kelly–Lochbaum waveguide augmented with a sublingual second source, cubic B-spline tract parameterization, and spatially varying learnable damping, optimized end-to-end by gradient descent from audio. On 20 segments from two independent sygyt datasets (5 singers, 10 pitches), the proposed model reduces log-spectral distance by 30–38% relative to an articulatory baseline, with the largest gains concentrated in the overtone region. Cepstral-envelope analysis further shows more accurate recovery of the merged formant structure characteristic of sygyt production. The model also outperforms a DDSP harmonic-plus-noise baseline with direct per-harmonic spectral control, suggesting that explicit acoustic structure is a useful inductive bias for overtone-singing copy-synthesis.
Download Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Recent diffusion-based generative models have achieved strong results in domain-specific audio generation tasks such as speech, singing, and instrumental music synthesis. However, these models are typically specialized and do not generalize well to mixed or intermediate audio types. In this work, we adapt a diffusion-based model originally designed for multi-instrument music synthesis to voice conversion, covering both speech and singing within a unified framework. Specifically, we extend musical note-based conditioning to include phonetic posteriorgrams (PPGs) and pitch contours, and reinterpret timbre conditioning as speaker or singer identity via feature-wise linear modulation. Experiments show that the adapted model matches or surpasses a dedicated voice conversion system in terms of naturalness and performer similarity, while maintaining accurate pitch control across speech and singing. At the same time, we observe limitations in phonetic fidelity and a degradation in vocal quality when incorporating instrumental training data. Furthermore, we demonstrate that off-the-shelf feature extractors provide effective conditioning signals, enabling large-scale self-supervised training without manual annotations. These results highlight the potential of cross-domain model transfer towards unified audio generation systems capable of handling speech, singing, and music.
Download Explicit Wave Digital Model of the Fulltone OCD Pedal Based on Canonical Piecewise-Linear Functions Virtual Analog (VA) modeling aims at digitally emulating analog audio equipment while preserving its characteristic nonlinear behavior and musical expressiveness. In the context of guitar effects, overdrive pedals represent a cornerstone of many signal chains, as they strongly contribute to the perceived dynamics, articulation, and timbral identity of the instrument. Among these, the Fulltone OCD overdrive is considered a standard in both studio and live environments, being widely adopted across rock and metal genres. In this article, we present an explicit Wave Digital (WD) model of the Fulltone OCD (v2) pedal. By exploiting the circuit topology, the MOSFETs and the germanium diode composing the asymmetric clipping stage are grouped into a single equivalent nonlinear element, enabling an explicit WD realization that avoids costly iterative solvers. The resulting nonlinear characteristic is approximated by means of a Canonical Piecewise-Linear (CPWL) function, yielding a compact and efficient explicit model suitable for real-time implementation. The proposed model is validated against reference simulations and implemented both in MATLAB and as a real-time audio plug-in using the JUCE framework.