Enhancing LSTM RNN-Based Speech Overlap Detection by Artificially Mixed Data

Hagerer, Gerhard; Pandit, Vedhas; Eyben, Florian; Schuller, Björn

AES E-Library

Enhancing LSTM RNN-Based Speech Overlap Detection by Artificially Mixed Data

This paper presents a new method for Long Short-Term Memory Recurrent Neural Network (LSTM) based speech overlap detection. To this end, speech overlap data is created artificially by mixing large amounts of speech utterances. Our elaborate training strategies and presented network structures demonstrate performance surpassing the considered state-of-the-art overlap detectors. Thereby we target the full ternary task of non-speech, speech, and overlap detection. Furthermore, speakers' gender is recognised, as the first successful combination of this kind within one model.

Authors: Hagerer, Gerhard; Pandit, Vedhas; Eyben, Florian; Schuller, Björn
Affiliations: audEERING GmbH, Gilching, Germany; University of Passau, Passau, Germany(See document for exact affiliation information.)
AES Conference: 2017 AES International Conference on Semantic Audio (June 2017)
Paper Number: P1-1
Publication Date: June 13, 2017 Import into BibTeX
Subject: Semantic Audio
Permalink: https://www.aes.org/e-lib/browse.cfm?elib=18764

Click to purchase paper as a non-member or login as an AES member. If your company or school subscribes to the E-Library then switch to the institutional version. If you are not an AES member and would like to subscribe to the E-Library then Join the AES!

This paper costs $33 for non-members and is free for AES members and E-Library subscribers.

Learn more about the AES E-Library

E-Library Location: /conf/2017/semantic/semantic_audio_2017_paper_13.pdf

Start a discussion about this paper!

AES E-Library

Enhancing LSTM RNN-Based Speech Overlap Detection by Artificially Mixed Data

ABOUT AES

Contact Us