VoiceFX: CLAP-Based Audio Quality Improvement for Singing and Speech

Elena Georgieva
DAFx-2026 - Cambridge
This project introduces an automatic method for enhancing audio quality in singing and speech. Using recordings from the LibriSpeech and Smule DAMP dataset, I applied a set of degradations and tested a set of audio effect "remedies" designed to reverse them: a high shelf filter, de-esser, noise reduction, and high-pass filter. I used the CLAP (Contrastive Language-Audio Pretraining) model to estimate recording quality and recommend remedies by comparing audio clips to descriptive text prompts in the shared embedding space. To evaluate my method, I conducted a large-scale listener study with 234 participants and 4,600 ratings. While CLAP encoded some relevant information of vocal recording quality, it often favored remedies like noise reduction while listeners preferred the original clips, suggesting that perceptual artifacts introduced by enhancement may not be captured by CLAP. My findings underscore the value of human judgment: embedding models can guide enhancement, but perceptual validation remains valuable. Audio examples are available online on the DAFx demo website.
Download