Classification of Spatial Audio Location and Content Using Convolutional Neural Networks

Hirvonen, Toni

AES E-Library

Classification of Spatial Audio Location and Content Using Convolutional Neural Networks

This paper investigates the use of Convolutional Neural Networks for spatial audio classification. In contrast to traditional methods that use hand-engineered features and algorithms, we show that a Convolutional Network in combination with generic preprocessing can give good results and allows for specialization to challenging conditions. The method can adapt to e.g. different source distances and microphone arrays, as well as estimate both spatial location and audio content type jointly. For example, with typical single-source material in a simulated reverberant room, we can achieve cross-validation accuracy of 94.3% for 40-ms frames across 16 classes (eight spatial directions, content type speech vs. music).

Author: Hirvonen, Toni
Affiliation: Dolby Laboratories, Stockholm, Sweden
AES Convention: 138 (May 2015) Paper Number: 9294
Publication Date: May 6, 2015 Import into BibTeX
Subject: Sound Localization and Separation
Permalink: https://www.aes.org/e-lib/browse.cfm?elib=17718

Click to purchase paper as a non-member or login as an AES member. If your company or school subscribes to the E-Library then switch to the institutional version. If you are not an AES member and would like to subscribe to the E-Library then Join the AES!

This paper costs $33 for non-members and is free for AES members and E-Library subscribers.

Learn more about the AES E-Library

E-Library Location: (CD 138Papers) /conv/138/9294.pdf

Start a discussion about this paper!

AES E-Library

Classification of Spatial Audio Location and Content Using Convolutional Neural Networks

ABOUT AES

Contact Us