Download A Dual-Stream Framework Combining Audio Spectrogram Transformer and Dynamic Mode Decomposition for Plate Modal Parameter Estimation ★ Plate reverberation is characterized by a dense distribution of resonant modes, which makes the estimation of modal parameters from observed responses a challenging inverse problem. To address this problem, we propose a physics guided dual stream framework that integrates an Audio Spectrogram Transformer (AST) with Dynamic Mode Decomposition (DMD). The AST branch models the global temporal and spectral structure of the impulse response, whereas the DMD branch extracts local descriptors associated with modal dynamics. The resulting representations are fused and processed by convolutional prediction heads to jointly estimate mode presence and the corresponding modal parameters. Experiments on the official validation set of Task B in the DAFx Challenge show that the proposed method reduces the overall relative error from 1.976 for the official baseline to 0.867. These results demonstrate that integrating local dynamic information derived from physical modeling with global transformer based representations substantially improves plate modal parameter estimation.
Download Sound Synthesis for Nonlinear Plates In this paper, a simple finite difference scheme for a rectangular dynamic nonlinear plate, under free boundary conditions is presented. The algorithm is straightforward to program, and is capable of reproducing, to a first approximation, the behaviour of various percussion instruments whose timbre depends crucially on nonlinear effects (due to high-speed strikes), including transient pitch glides and the buildup of high-frequency energy. Though computationally intensive, algorithms such as that presented here promise more faithful sound synthesis and, as with all physical model inspired synthesis algorithms, require the specification of only a few, physically meaningful parameters. Full details of the algorithm, including the setting of boundary conditions and computational demands are provided. Numerical simulation results are presented.
Download A Real-Time Synthesis Oriented Tanpura Model Physics-based synthesis of tanpura drones requires accurate simulation of stiff, lossy string vibrations while incorporating sustained contact with the bridge and a cotton thread. Several challenges arise from this when seeking efficient and stable algorithms for real-time sound synthesis. The approach proposed here to address these combines modal expansion of the string dynamics with strategic simplifications regarding the string-bridge and stringthread contact, resulting in an efficient and provably stable timestepping scheme with exact modal parameters. Attention is given also to the physical characterisation of the system, including string damping behaviour, body radiation characteristics, and determination of appropriate contact parameters. Simulation results are presented exemplifying the key features of the model.
Download Spatial Sound Synthesis for Circular Membranes Physical models of real or virtual instruments are usually only exploited for the generation of wave forms. However, models of twoand three-dimensional vibrating structures contain also information about the sound radiation into the free field. This contribution presents a model for a membrane from which the required driving functions for a multichannel loudspeaker array are derived. The resulting sound field reproduces not only the musical timbre of the sounding body but also its spatial radiation characteristics. It is suitable for real-time synthesis without pre-recorded or presynthesized source tracks.
Download Modeling Bowl Resonators Using Circular Waveguide Networks We propose efficient implementations of a glass harmonica and a Tibetan bowl using circular digital waveguide networks. Circular networks provide a physically meaningful representation of bowl resonators. Just like the real instruments, both models can be either struck or rubbed using a hard mallet, a violin bow, or a wet finger.
Download Higher-Order Scattering Delay Networksfor Artificial Reverberation Computer simulations of room acoustics suffer from an efficiency vs accuracy trade-off, with highly accurate wave-based models being highly computationally expensive, and delay-network-based models lacking in physical accuracy. The Scattering Delay Network (SDN) is a highly efficient recursive structure that renders first order reflections exactly while approximating higher order ones. With the purpose of improving the accuracy of SDNs, in this paper, several variations on SDNs are investigated, including appropriate node placement for exact modeling of higher order reflections, redesigned scattering matrices for physically-motivated scattering, and pruned network connections for reduced computational complexity. The results of these variations are compared to state-of-the-art geometric acoustic models for different shoebox room simulations. Objective measures (Normalized Echo Densities (NEDs) and Energy Decay Curves (EDCs)) showed a close match between the proposed methods and the references. A formal listening test was carried out to evaluate differences in perceived naturalness of the synthesized Room Impulse Responses. Results show that increasing SDNs’ order and adding directional scattering in a fully-connected network improves perceived naturalness, and higher-order pruned networks give similar performance at a much lower computational cost.
Download Interaction-optimized Sound Database Representation Interactive navigation within geometric, feature-based database representations allows expressive musical performances and installations. Once mapped to the feature space, the user’s position in a physical interaction setup (e.g. a multitouch tablet) can be used to select elements or trigger audio events. Hence physical displacements are directly connected to the evolution of sonic characteristics — a property we call analytic sound–control correspondence. However, automatically computed representations have a complex geometry which is unlikely to fit the interaction setup optimally. After a review of related work, we present a physical model-based algorithm that redistributes the representation within a user-defined region according to a user-defined density. The algorithm is designed to preserve the analytic sound-control correspondence property as much as possible, and uses a physical analogy between the triangulated database representation and a truss structure. After preliminary pre-uniformisation steps, internal repulsive forces help to spread points across the whole region until a target density is reached. We measure the algorithm performance relative to its ability to produce representations corresponding to user-specified features and to preserve analytic sound–control correspondence during a standard density-uniformisation task. Quantitative measures and visual evaluation outline the excellent performances of the algorithm, as well as the interest of the pre-uniformisation steps.
Download Simulation of piano sustain-pedal effect by parallel second-order filters This paper presents a sustain-pedal effect simulation algorithm for piano synthesis, by using parallel second-order filters. A robust two-step filter design procedure, based on frequency-zooming ARMA modeling and least squares fit, is applied to calibrate the algorithm from impulse responses of the soundboard and the string register. The model takes into account the differences in coupling between the various strings. The algorithm can be applied to both sample-based and physics-based piano synthesizers.
Download A perceptually inspired generative model of rigid-body contact sounds Contact between rigid-body objects produces a diversity of impact and friction sounds. These sounds can be synthesized with detailed simulations of the motion, vibration and sound radiation of the objects, but such synthesis is computationally expensive and prohibitively slow for many applications. Moreover, detailed physical simulations may not be necessary for perceptually compelling synthesis; humans infer ecologically relevant causes of sound, such as material categories, but not with arbitrary precision. We present a generative model of impact sounds which summarizes the effect of physical variables on acoustic features via statistical distributions fit to empirical measurements of object acoustics. Perceptual experiments show that sampling from these distributions allows efficient synthesis of realistic impact and scraping sounds that convey material, mass, and motion.
Download Estimation and Modeling of Pinna-Related Transfer Functions This paper considers the problem of modeling pinna-related transfer functions (PRTFs) for 3-D sound rendering. Following a structural modus operandi, we present an algorithm for the decomposition of PRTFs into ear resonances and frequency notches due to reflections over pinna cavities. Such an approach allows to control the evolution of each physical phenomenon separately through the design of two distinct filter blocks during PRTF synthesis. The resulting model is suitable for future integration into a structural head-related transfer function model, and for parametrization over anthropometrical measurements of a wide range of subjects.