TR2024-118

Sound Event Bounding Boxes

- Ebbers, J., Germain, F.G., Wichern, G., Le Roux, J., "Sound Event Bounding Boxes", Interspeech, DOI: 10.21437/Interspeech.2024-2075, September 2024, pp. 562-566.
  BibTeX TR2024-118 PDF Software
  - @inproceedings{Ebbers2024sep,
  - author = {Ebbers, Janek and Germain, François G and Wichern, Gordon and {Le Roux}, Jonathan},
  - title = {{Sound Event Bounding Boxes}},
  - booktitle = {Interspeech},
  - year = 2024,
  - pages = {562--566},
  - month = sep,
  - doi = {10.21437/Interspeech.2024-2075},
  - issn = {2958-1796},
  - url = {https://www.merl.com/publications/TR2024-118}
  - }
MERL Contacts:
- Gordon
  Wichern
- Jonathan
  Le Roux
Research Areas:

Artificial Intelligence, Speech & Audio

Abstract:

Sound event detection is the task of recognizing sounds and determining their extent (onset/offset times) within an audio clip. Existing systems commonly predict sound presence posteriors in short time frames. Then, thresholding produces binary frame-level presence decisions, with the extent of individual events determined by merging presence in consecutive frames. In this paper, we show that frame-level thresholding deteriorates event extent prediction by coupling it with the system’s sound presence confidence. We propose to decouple the prediction of event extent and confidence by introducing sound event bounding boxes (SEBBs), which format each sound event prediction as a combination of a class type, extent, and overall confidence. We also propose a change-detection-based algorithm to convert frame-level posteriors into SEBBs. We find the algorithm significantly improves the performance of DCASE 2023 Challenge systems, boosting the state of the art from .644 to .686 PSDS1.

Software & Data Downloads

Sound Event Bounding Boxes

Related Publication

Ebbers, J., Germain, F.G., Wichern, G., Le Roux, J., "Sound Event Bounding Boxes", arXiv, June 2024.

BibTeX arXiv

@article{Ebbers2024jun,
author = {Ebbers, Janek and Germain, François G and Wichern, Gordon and {Le Roux}, Jonathan},
title = {{Sound Event Bounding Boxes}},
journal = {arXiv},
year = 2024,
month = jun,
url = {https://arxiv.org/abs/2406.04212}
}

MERL Contacts:

GordonWichern

JonathanLe Roux

Research Areas:

Abstract:

Gordon
Wichern

Jonathan
Le Roux