Benchmarking Integrated GPU Acceleration of Real-Time Neural Audio Inference on Snapdragon

Avery Huang; Gautham Srinivasan; Akito van Troyer; Victor Zappi
DAFx-2026 - Cambridge
This paper investigates whether integrated GPUs on Qualcomm Snapdragon SoCs can accelerate streaming inference of neural audio models. Five models spanning three orders of magnitude in parameter count are benchmarked across three inference approaches (best available CPU, QNN CPU, and QNN GPU). Results reveal when GPU acceleration offers meaningful gains, when per-call overhead negates benefits, and how model size and architecture determine GPU suitability.
Download