Download Digital Synthesis Models of Clarinet-Like Instruments Including Nonlinear Losses in the Resonator
This paper presents a real-time algorithm for the synthesis of reed instruments, taking into account nonlinear losses at the first open tonehole. The physical model on which the synthesis model relies on is based on the experimental works of Dalmont et al. who have shown that for high pressure levels within the bore, an air jet obeying the Bernoulli flow model, hence acting as a nonlinear resistance, is created at the open end of the bore. We study the effect of these additional losses on the response of the bore to an acoustic flow impulse at different levels and on the self oscillations. We show that at low frequencies, these nonlinear losses are of the same order of magnitude than the viscothermal linear losses and modifie the functioning of the whole instrument. For real-time synthesis purposes, a simplified algorithm is proposed and compared to the more accurate model.
Download Decomposition of steady state instrument data into excitation system and formant filter components
This paper describes a method for decomposing steady-state instrument data into excitation and formant filter components. The input data, taken from several series of recordings of acoustical instruments is analyzed in the frequency domain, and for each series a model is built, which most accurately represents the data as a source-filter system. The source part is taken to be a harmonic excitation system with frequency-invariant magnitudes, and the filter part is considered to be responsible for all spectral inhomogenieties. This method has been applied to the SHARC database of steady state instrument data to create source-filter models for a large number of acoustical instruments. Subsequent use of such models can have a wide variety of applications, including wavetable and physical modeling synthesis, high quality pitch shifting, and creation of “hybrid” instrument timbres.
Download Sparse Decomposition, Clustering and Noise for Fire Texture Sound Re-Synthesis
In this paper we introduce a framework that represents environmental texture sounds as a linear superposition of independent foreground and background layers that roughly correspond to entities in the physical production of the sound. Sound samples are decomposed into a sparse representation with the matching pursuit algorithm and a dictionary of Daubechies wavelet atoms. An agglomerative clustering procedure groups atoms into short transient molecules. A foreground layer is generated by sampling these sound molecules from a distribution, whose parameters are estimated from the input sample. The residual signal is modelled by an LPC-based source-filter model, synthesizing the background sound layer. The capability of the system is demonstrated with a set of fire sounds.
Download Modelling of Brass Instrument Valves
Finite difference time domain (FDTD) approaches to physical modeling sound synthesis, though more computationally intensive than other techniques (such as, e.g., digital waveguides), offer a great deal of flexibility in approaching some of the more interesting real-world features of musical instruments. One such case, that of brass instruments, including a set of time-varying valve components, will be approached here using such methods. After a full description of the model, including the resonator, and incorporating viscothermal loss, bell radiation, a simple lip model, and time varying valves, FDTD methods are introduced. Simulations of various characteristic features of valve instruments, including half-valve impedances, note transitions, and characteristic multiphonic timbres are presented, as are illustrative sound examples.
Download A Multi-Resolution Spectrogram Approach for Estimating the Physical Parameters of a Plate Reverb
The ResNet-18 image classification model is employed to determine the physical parameters of a plate reverb from a recording of the impulse response. The model is adapted to derive parameters using normalized and down-sampled multi-resolution spectrograms computed from the provided impulse responses (IRs). To refine the prediction of the output location, the spectral phase response is also included as an additional input channel to the network since multiple output locations can give the same magnitude response for high-order resonant modes. On a 5000 IR validation set, our model achieves an average normalized mean squared error (NMSE) of 0.02920 across all parameters, with the lowest average NMSE occurring for parameters yo (0.00228), Ly (0.00347), and xo (0.00574).
Download Experimental Study of Guitar Pickup Nonlinearity
In this paper, we focus on studying nonlinear behavior of the pickup of an electric guitar and on its modeling. The approach is purely experimental, based on physical assumptions and attempts to find a nonlinear model that, with few parameters, would be able to predict the nonlinear behavior of the pickup. In our experimental setup a piece of string is attached to a shaker and vibrates perpendicularly to the pickup in frequency range between 60 Hz and 400 Hz. The oscillations are controlled by a linearizion feedback to create a purely sinusoidal steady state movement of the string. In the first step, harmonic distortions of three different magnetic pickups (a single-coil, a humbucker, and a rail-pickup) are compared to check if they provide different distortions. In the second step, a static nonlinearity of Paiva’s model is estimated from experimental signals. In the last step, the pickup nonlinearities are compared and an empirical model that fits well all three pickups is proposed.
Download WaveNet-Style Guitar Amplifier Model Pruning for Real-Time iOS Deployment
WaveNet-style convolutional networks emulate tube amplifiers and distortion pedals with high fidelity, but their computational cost has confined them to desktops or dedicated DSP hardware. We present a sparse-enabled WaveNet inference engine for iOS that runs heavily pruned neural guitar amplifier models in real time on iPhones. Aggressive iterative magnitude pruning removes 90% of the network weights with no perceptible loss in quality. A custom sparse C++ engine turns this sparsity directly into compute savings, sustaining low-latency real-time operation on a CPU-only iPhone implementation where the dense model cannot. On-device output matches the trained model to within int16 quantization error. At the demonstration, visitors will play a guitar through the app on iPhone hardware and A/B the on-device pruned model against the physical pedal it emulates. Source code and audio examples are available online.
Download Non-Iterative Solvers For Nonlinear Problems: The Case of Collisions
Nonlinearity is a key feature in musical instruments and electronic circuits alike, and thus in simulation, for the purposes of physics-based modeling and virtual analog emulation, the numerical solution of nonlinear differential equations is unavoidable. Ensuring numerical stability is thus a major consideration. In general, one may construct implicit schemes using well-known discretisation methods such as the trapezoid rule, requiring computationally-costly iterative solvers at each time step. Here, a novel family of provably numerically stable time-stepping schemes is presented, avoiding the need for iterative solvers, and thus of greatly reduced computational cost. An application to the case of the collision interaction in musical instrument modeling is detailed.
Download Time-domain model of the singing voice
A combined physical model for the human vocal folds and vocal tract is presented. The vocal fold model is based on a symmetrical 16 mass model by Titze. Each vocal fold is modeled with 8 masses that represent the mucosal membrane coupled by non-linear springs to another 8 masses for the vocalis muscle together with the ligament. Iteratively, the value of the glottal flow is calculated and taken as input for calculation of the aerodynamic forces. Together with the spring forces and damping forces they yield the new positions of the masses that are then used for the calculation of a new glottal flow value. The vocal tract model consists of a number of uniform cylinders of fixed length. At each discontinuity incident, reflected and transmitted waves are calculated including damping. Assuming a linear system, the pressure signal generated by the vocal fold model is either convoluted with the Green’s function calculated by the vocal tract model or calculated interactively assuming variable reflection coefficients for the glottis and the vocal tract during phonation. The algorithms aim at real-time performance and are implemented in MATLAB.
Download A Source Localization/Separation/Respatialization System Based on Unsupervised Classification of Interaural Cues
In this paper we propose a complete computational system for Auditory Scene Analysis. This time-frequency system localizes, separates, and spatializes an arbitrary number of audio sources given only binaural signals. The localization is based on recent research frameworks, where interaural level and time differences are combined to derive a confident direction of arrival (azimuth) at each frequency bin. Here, the power-weighted histogram constructed in the azimuth space is modeled as a Gaussian Mixture Model, whose parameter structure is revealed through a weighted Expectation Maximization. Afterwards, a bank of Gaussian spatial filters is configured automatically to extract the sources with significant energy accordingly to a posterior probability. In this frequency-domain framework, we also inverse a geometrical and physical head model to derive an algorithm that simulates a source as originating from any azimuth angle.