Download Morphing techniques for enhanced scat singing In jazz, scat singing is a phonetic improvisation that imitates instrumental sounds. In this paper, we propose a system that aims to transform singing voice into real instrument sounds, extending the possibilities for scat singers. Analysis algorithms in the spectral domain extract voice parameters, which drive the resulting instrument sound. A small database contains real instrument samples that have been spectrally analyzed offline. Two different prototypes are introduced, producing sounds of a trumpet and a bass guitar respectively.
Download The Influence of Small Variations in a Simplified Guitar Amplifier Model A strongly simplified guitar amplifier model, consisting of four stages, is presented. The exponential sweep technique is used to measure the frequency dependent harmonic spectra. The influence of small variations of the system parameters on the harmonic components is analyzed. The differences of the spectra are explained and visualized.
Download High-Performance Real-Time FIR-Filtering Using Fast Convolution on Graphics Hardware In this paper we examine how graphic hardware can be used for real-time FIR filtering. We implement uniformly-partitioned fast convolution in the frequency-domain and evaluate its performance on a NVIDIA GTX 285 graphics card. Motivated by audio rendering for virtual reality, our focus lies on large-scale realtime filtering with a multitude of channels, long impulse responses and low latencies. Graphics hardware has already been used for audio signal processing — including FIR and IIR filtering with respect to offline and real-time processing. However, the combination of GPU computing and real-time conditions leads to a number of challenges that have not been reviewed in detail. The new contribution of this paper is an implementation and detailled analysis of a frequency-domain fast convolution method on GPUs. We discuss specific problems that emerge under real-time conditions. Our method allows to achieve an outstanding real-time filtering performance. In this work, we do not only regard a timeinvariant filtering, but also time-varying filtering, where filters are exchanged during runtime. Furthermore, we examine the opportunities of distributed computation — using CPU and GPU — in order to maximize the performance. Finally, we identify bottlenecks and explain their impact on filter exchange latencies and update rates.
Download What you hear is what you see: Audio quality from Image Quality Metrics In this study, we investigate the feasibility of utilizing stateof-the-art perceptual image metrics for evaluating audio signals by representing them as spectrograms. The encouraging outcome of the proposed approach is based on the similarity between the neural mechanisms in the auditory and visual pathways. Furthermore, we customise one of the metrics which has a psychoacoustically plausible architecture to account for the peculiarities of sound signals. We evaluate the effectiveness of our proposed metric and several baseline metrics using a music dataset, with promising results in terms of the correlation between the metrics and the perceived quality of audio as rated by human evaluators.
Download Trans-synthesis System for Polyphonic Musical Recordings of Bowed-String Instruments A system that tries to analyze polyphonic musical recordings of bowed-string instruments, extract synthesis parameters of individual instrument and then re-synthesize is proposed. In the analysis part, multiple F0s estimation and partials tracking are performed based on modified WGCDV (weighted greatest common divisor and vote) method and high-order HMM. Then, dynamic time warping algorithm is employed to align the above results with the score to improve the accuracy of the extracted parameters. In the re-synthesis part, simple additive synthesis is employed. Here, one can experiment on changing timbres, pitches and so on or adding vibrato or other effects on the same piece of music.
Download Separation Of Speech Signal From Complex Auditory Scenes The hearing system, even in front of complex auditory scenes and in unfavourable conditions, is able to separate and recognize auditory events accurately. A great deal of effort has gone into the understanding of how, after having captured the acoustical data, the human auditory system processes them. The aim of this work is the digital implementation of the decomposition of a complex sound in separate parts as it would appear to a listener. This operation is called signal separation. In this work, the separation of speech signal from complex auditory scenes has been studied and an experimentation of the techniques that address this problem has been done.
Download Audio Signal Processing and Object-oriented Systems Object-oriented programming (OOP) has been for many years now one of the most important programming paradigms used in a variety of applications. Digital audio signal processing can benefit largely from this approach for systems development. In this paper a number of approaches to using object-orientation in audio processing systems are reviewed. Existing systems of audio processing are introduced and discussed in detail. The paper also draws attention to the different OOP techniques enabled and supported by these systems. Comparative code and tutorial examples are included, providing an insight into the development of signal processing applications using objects.
Download Transforming Singing Voice Expression - The Sweetness Effect We propose a real-time system which is targeted to music production in the context of vocal recordings. The aim is to transform the singer’s voice characteristics in order to achieve a sweet sounding voice. It combines three different transformations namely SubHarmonic Component Reduction (reduction of sub-harmonics, which are found in voices with vocal disorders), Vocal Tract Excitation Modification (to achieve a change in loudness) and the Intonation Modification (to achieve smoother transitions in pitch). The transformations are done in the frequency domain based on an enhanced phase-locked vocoder. The Expression Adaptive Control estimates the amount of present vocal disorder in the singer’s voice. This estimate automatically controls the amount of SubHarmonic Component reduction to assure a natural sounding transformation.
Download Generalised Prior Subspace Analysis for Polyphonic Pitch Transcription A reformulation of Prior Subspace Analysis (PSA) is presented, which restates the problem as that of fitting an undercomplete signal dictionary to a spectrogram. Further, a generalization of PSA is derived which allows the transcription of polyphonic pitched instruments. This involves the translation of a single frequency prior subspace of a note to approximate other notes, overcoming the problem of needing a separate basis function for each note played by an instrument. Examples are then demonstrated which show the utility of the generalised PSA algorithm for the purposes of polyphonic pitch transcription.
Download Modal analysis of impact sounds with ESPRIT in Gabor transforms Identifying the acoustical modes of a resonant object can be achieved by expanding a recorded impact sound in a sum of damped sinusoids. High-resolution methods, e.g. the ESPRIT algorithm, can be used, but the time-length of the signal often requires a sub-band decomposition. This ensures, thanks to sub-sampling, that the signal is analysed over a significant duration so that the damping coefficient of each mode is estimated properly, and that no frequency band is neglected. In this article, we show that the ESPRIT algorithm can be efficiently applied in a Gabor transform (similar to a sub-sampled short-time Fourier transform). The combined use of a time-frequency transform and a high-resolution analysis allows selective and sharp analysis over selected areas of the time-frequency plane. Finally, we show that this method produces high-quality resynthesized impact sounds which are perceptually very close to the original sounds.