Download FoleySet: A Multi-Level Human-Annotated Foley Sound Dataset
We introduce FoleySet, a human-annotated Foley sound dataset designed to support research on sound-event understanding and Foley sound generation. The dataset provides annotations at multiple levels of granularity, capturing both broad event categories and more detailed semantic or production-related attributes. This multi-level structure supports tasks such as classification, retrieval, captioning, and controllable generation. FoleySet is intended to address the limited availability of systematically annotated Foley material and to provide a common resource for evaluating models across different levels of semantic detail.
Download Real-Time Neural Audio on Apple Silicon: Benchmarking Inference Frameworks Under Realistic DAW Contention
Neural network models are increasingly deployed in audio plugins across a wide range of applications, including amplifier emulation, effects modeling, and synthesis. This paper evaluates widely used inference options including BNNSGraph, RTNeural, LibTorch, ONNX Runtime, and anira on model architectures commonly used in neural audio plugins. The key contribution is moving beyond isolated benchmarks to evaluate performance under realistic DAW contention, constructing mix sessions with configurable plugin loads. Results show that isolated benchmarks can be misleading, and BNNSGraph proves most robust for convolutional models on Apple Silicon.
Download Probing Low-Level Acoustic Attribute Encoding in CLAP Audio Embeddings
This work analyzes CLAP audio embeddings through a probing framework, studying the encoding of reverberation (RT60), loudness (LUFS), spectral content (SC), and relative pitch (RP). Results show that all attributes are reliably recoverable from CLAP embeddings, with RT60, LUFS, and RP approximately linearly encoded, while SC requires non-linear probes. The identified patterns generalize across eight additional audio foundation models.
Download Fast Parametric Matrices for Lossless Feedback Delay Networks
This paper presents a framework for designing creative reverbs using parametric orthogonal feedback matrices on Feedback Delay Networks (FDNs) through recursive Kronecker products of 2D rotation and reflection matrices. By parameterizing each 2×2 kernel with a single angle, we construct a family of 2M×2M orthogonal matrices that maintain losslessness while enabling continuous control over network topology. We then exploit their recursive definition to compute the feedback operation with an O(N log₂ N) divide-and-conquer algorithm that matches the Fast Walsh-Hadamard Transform time complexity while offering parametric flexibility. Strategic manipulation of individual kernel angles enables creative sound design applications, such as stereo cross-coupling, selective freeze, and time-varying modulation for resonance breaking.
Download Praat AudioTools: Analysis Objects as Compositional Controllers for Interpretable Sound Transformation
This demonstration presents Praat AudioTools, an open-source hybrid toolkit that repurposes Praat's phonetic-analysis environment for electroacoustic composition, sound design, and offline analysis–resynthesis workflows. Rather than treating analysis data as temporary measurements hidden inside an audio processor, Praat AudioTools exposes pitch contours, formant structures, temporal segmentations, spectral descriptors, phrase boundaries, stochastic trajectories, and host-application exchange files as editable compositional objects. These objects can be inspected, modified, chained, reused, and rendered into new sound transformations. The demonstration focuses on seven offline workflows: Neural Ambient Drone Designer, Praat for Max and Max for Live, Phase-Space Composer, Reich Generator, MCMC Musical Variation, Messagesquisse Opening, and Vector/Full-Chain composition workflows. None of the examples are presented as real-time effects. Instead, they show an "edit-in-the-middle" model in which sound is analyzed, intermediate representations are made visible, compositional decisions are applied to those representations, and the result is rendered as audio. The aim is to demonstrate a transparent alternative to both conventional black-box audio effects and end-to-end generative audio systems: a compositional environment where analysis objects become controllers, traces, scores, and reproducible technical artifacts.
Download Stability Analysis of Time-Varying Virtual Analog Filters
Time-varying virtual analog filters used in digital audio effects and synthesizers are often implemented by discretizing continuous-time state-space systems using trapezoidal integration. When filter parameters such as cutoff frequency or resonance vary over time, as is the norm in musical applications, proving BIBO stability of the resulting time-varying discrete-time system becomes nontrivial. In this paper, we review the technique of common quadratic Lyapunov functions (CQLFs) from the control systems literature and show how a continuous-time CQLF is preserved through discretization. This allows us to prove stability of some time-varying virtual analog filters by working in the often simpler continuous-time domain. We apply this framework to several filters of musical interest, providing new proofs of stability for the state variable filter and the Sallen-Key filter, and new bounds on the stable time-varying parameter range for the Moog ladder and diode ladder filters.
Download PolyMap: A 64-Channel Polyphonic Guitar Pickup System
In electric guitars, the vibrations of the strings are typically sensed by coils of wire combined with a magnet, called pickups. The pickups and their position along the strings contribute strongly to the instrument's sound. Most guitars feature one to three pickups, each spanning across all strings with fixed positions and generating a single mono output. The work of this Master's Thesis at ETH Zürich introduces a new pickup system called PolyMap, which senses each string individually and at multiple locations. The system is demonstrated with a custom-made eight-string guitar that contains eight pickups per string for a total of 64 pickups. The signals from these 64 pickups are individually digitized inside the guitar and transmitted over a multichannel audio digital interface (MADI), a low-latency digital audio interface, to a computer for further processing. PolyMap enables high-resolution sensing of an electric guitar's strings and enables extensive post-processing capabilities for musicians, audio engineers, and researchers. To the best of our knowledge, this is the first polyphonic guitar pickup system with such a complete feature set.
Download Ambisonic Decoder Equalization in Reverberant Environments via Closed-Hull Crosstalk Inversion
In this work, the authors develop a higher order Ambisonic (HOA) decoder that compensates for listening room reverberation by cascading a conventional decode matrix with a crosstalk matrix derived from room impulse responses (RIR) constrained over the listener area, namely by sampling over a spherical boundary according to HOA convention. A decoder is generated for simulated RIRs of mixed specular and diffuse reverberation, and spectral and spatiotemporal energy distributions are shown for directional and diffuse HOA signals rendered through the reverberant room. Results demonstrate that the decoder renders temporally-variant directional sources with higher directivity as compared to a conventional decoder.ing Ambisonic decoders in reverberant environments using closed-hull crosstalk convolution. By modeling the acoustic path from each loudspeaker to the listener's ears via measured room impulse responses, we formulate equalization filters that compensate for room-induced distortions in the decoded signals. The proposed closed-hull approach constrains the solution to preserve perceptually relevant spatial cues while minimizing spectral coloration. Listening tests demonstrate improved externalization, timbral transparency, and localization accuracy compared to standard free-field decoding in reverberant conditions.
Download PAEDB: A Synthetic Primary-Ambient Dataset Generation Pipeline for Automatic Upmixing Using Deep Neural Networks
Automatic blind upmixing aims to convert audio from a smaller channel format (e.g. mono or stereo) into a multichannel format using estimates of direct and diffuse spatial statistics within the signal. Current approaches rely on primary-ambient extraction (PAE) algorithms, which lack real-world context through limited processing windows. Deep learning music source separation (MSS) models have been applied in voice-primary-ambient extraction (VPA) upmixing systems for handling direct components, but still rely on DSP methods of surround channel generation. This work further investigates utilizing source separation within VPA upmixing, focusing specifically on the task of stereo decorrelation and ambience extraction for 5.1 surround. We also release PAEDB (Primary–Ambient Extraction Dataset), a high-quality music dataset derived from MUSDB18-HQ and MoisesDB, comprising 1,809 primary–ambient stem pairs totaling over 550 hours of audio. The performance of selected DNNs trained on PAEDB is then evaluated using signal metrics and a listening study. Our findings indicate that DNNs can effectively model the behavior of PAE algorithms, establishing PAEDB as a strong foundation for ML upmixing systems and underscoring the need for higher-quality multichannel data to advance beyond conventional methods.
Download PolyADAA: Improving Aliasing Reduction in Memoryless Nonlinearities Using Lagrange Interpolation and Polynomial Approximation
Reducing the aliasing of nonlinear functions is an important problem in digital signal processing. The introduction of the Antiderivative Antialiasing (ADAA) method brought many benefits and is an active area of research. The current bottleneck in terms of aliasing reduction is the initial conversion from discrete- to continuous-time, which was previously done by linear interpolation. In this paper we derive PolyADAA, a method for computing the ADAA output when this conversion is done using higher order Lagrange interpolation. To obtain a viable solution, the nonlinear function is approximated using Chebyshev polynomials, which enable the ADAA integral to be computed. The paper provides numerical examples to show the effectiveness of the approach and discusses the advantages of the method.