Download Computer-generating emotional music: The design of an affective music algorithm
Download Source-Filter based Clustering for Monaural Blind Source Separation
In monaural blind audio source separation scenarios, a signal mixture is usually separated into more signals than active sources. Therefore it is necessary to group the separated signals to the final source estimations. Traditionally grouping methods are supervised and thus need a learning step on appropriate training data. In contrast, we discuss unsupervised clustering of the separated channels by Mel frequency cepstrum coefficients (MFCC). We show that replacing the decorrelation step of the MFCC by the non-negative matrix factorization improves the separation quality significantly. The algorithms have been evaluated on a large test set consisting of melodies played with different instruments, vocals, speech, and noise.
Download Ambisonic Decoder Equalization in Reverberant Environments via Closed-Hull Crosstalk Inversion
In this work, the authors develop a higher order Ambisonic (HOA) decoder that compensates for listening room reverberation by cascading a conventional decode matrix with a crosstalk matrix derived from room impulse responses (RIR) constrained over the listener area, namely by sampling over a spherical boundary according to HOA convention. A decoder is generated for simulated RIRs of mixed specular and diffuse reverberation, and spectral and spatiotemporal energy distributions are shown for directional and diffuse HOA signals rendered through the reverberant room. Results demonstrate that the decoder renders temporally-variant directional sources with higher directivity as compared to a conventional decoder.ing Ambisonic decoders in reverberant environments using closed-hull crosstalk convolution. By modeling the acoustic path from each loudspeaker to the listener's ears via measured room impulse responses, we formulate equalization filters that compensate for room-induced distortions in the decoded signals. The proposed closed-hull approach constrains the solution to preserve perceptually relevant spatial cues while minimizing spectral coloration. Listening tests demonstrate improved externalization, timbral transparency, and localization accuracy compared to standard free-field decoding in reverberant conditions.
Download A Quadric Surface Model of Vacuum Tubes for Virtual Analog Applications
Despite the prevalence of modern audio technology, vacuum tube amplifiers continue to play a vital role in the music industry. For this reason, over the years, many different digital techniques have been introduced for accomplishing their emulation. In this paper, we propose a novel quadric surface model for tube simulations able to overcome the Cardarilli model in terms of efficiency whilst retaining comparable accuracy when grid current is negligible. After showing the model capability to well outline tubes starting from measurement data, we perform an efficiency comparison by implementing the considered tube models as nonlinear 3-port elements in the Wave Digital domain. We do this by taking into account the typical common-cathode gain stage employed in vacuum tube guitar amplifiers. The proposed model turns out to be characterized by a speedup of 4.6× with respect to the Cardarilli model, proving thus to be promising for real-time Virtual Analog applications.
Download Discretization of Parametric Analog Circuits for Real-Time Simulations
The real-time simulation of analog circuits by digital systems becomes problematic when parametric components like potentiometers are involved. In this case the coefficients defining the digital system will change and have to be adapted. One common solution is to recalculate the coefficients in real-time, a possibly computationally expensive operation. With a view to the simulation using state-space representations, two parametric subcircuits found in typical guitar amplifiers are analyzed, namely the tone stack, a linear passive network used as simple equalizer and a distorting preamplifier, limiting the signal amplitude with LEDs. Solutions using trapezoidal rule discretization are presented and discussed. It is shown, that the computational costs in case of recalculation of the coefficients are reduced compared to the related DK-method, due to minimized matrix formulations. The simulation results are compared to reference data and show good match.
Download Distortion Recovery: A Two-Stage Method for Guitar Effect Removal
Removing audio effects from electric guitar recordings makes it easier for post-production and sound editing. An audio distortion recovery model not only improves the clarity of the guitar sounds but also opens up new opportunities for creative adjustments in mixing and mastering. While progress have been made in creating such models, previous efforts have largely focused on synthetic distortions that may be too simplistic to accurately capture the complexities seen in real-world recordings. In this paper, we tackle the task by using a dataset of guitar recordings rendered with commercial-grade audio effect VST plugins. Moreover, we introduce a novel two-stage methodology for audio distortion recovery. The idea is to firstly process the audio signal in the Mel-spectrogram domain in the first stage, and then use a neural vocoder to generate the pristine original guitar sound from the processed Mel-spectrogram in the second stage. We report a set of experiments demonstrating the effectiveness of our approach over existing methods, through both subjective and objective evaluation metrics.
Download Blind Arbitrary Reverb Matching
Reverb provides psychoacoustic cues that convey information concerning relative locations within an acoustical space. The need arises often in audio production to impart an acoustic context on an audio track that resembles a reference track. One tool for making audio tracks appear to be recorded in the same space is by applying reverb to a dry track that is similar to the reverb in a wet one. This paper presents a model for the task of “reverb matching,” where we attempt to automatically add artificial reverb to a track, making it sound like it was recorded in the same space as a reference track. We propose a model architecture for performing reverb matching and provide subjective experimental results suggesting that the reverb matching model can perform as well as a human. We also provide open source software for generating training data using an arbitrary Virtual Studio Technology plug-in.
Download Fast Sinusoid Synthesis for MPEG-4 HILN Parametric Audio Decoding
Additive sinusoidal synthesis is a popular technique for applications like sound synthesis or very low bit rate parametric audio decoding. In this paper, different algorithms for the efficient synthesis of sinusoids on general purpose CPUs as found in today’s PCs are investigated. Fast algorithms for time domain synthesis of constant and linearly changing frequencies are presented and compared to frequency domain synthesis approaches. Execution time and accuracy (SNR) of the algorithms are reported for different CPU types. Finally, the algorithms are implemented in a fast MPEG-4 HILN parametric audio decoder in order to evaluate their performance in a real world application.
Download The caterpillar system for data-driven concatenative sound synthesis
Concatenative data-driven synthesis methods are gaining more interest for musical sound synthesis and effects. They are based on a large database of sounds and a unit selection algorithm which finds the units that match best a given sequence of target units. We describe related work and our C ATERPILLAR synthesis system, focusing on recent new developments: the advantages of the addition of a relational SQL database, work on segmentation by alignment, the reformulation and extension of the unit selection algorithm using a constraint resolution approach, and new applications for musical and speech synthesis.
Download Transaural Stereo in a Beamforming Approach
This paper presents a study on algorithms for headphone-free binaural synthesis using a dedicated loudspeaker configuration. Both algorithms under investigation improve the properties of the binaural synthesis performance of the array. Firstly, beam-forming provides sound radiation localized at two freely adjustable, narrow target spots. Adjusting both spots to the locations of the listener’s ears achieves a good basis. Secondly, an additional interaural crosstalk canceler improves the overall result.