TR2026-126

Downstream-Task-Aware Unified Source Separation


    •  Mitsui, Y., Aihara, R., Saito, T., Masuyama, Y., Boeddeker, C., Richter, J., Wichern, G., Le Roux, J., "Downstream-Task-Aware Unified Source Separation", International Workshop on Acoustic Signal Enhancement (IWAENC), September 2026.
      BibTeX TR2026-126 PDF
      • @inproceedings{Mitsui2026sep,
      • author = {Mitsui, Yoshiki and Aihara, Ryo and Saito, Tatsuhiko and Masuyama, Yoshiki and Boeddeker, Christoph and Richter, Julius and Wichern, Gordon and {Le Roux}, Jonathan},
      • title = {{Downstream-Task-Aware Unified Source Separation}},
      • booktitle = {International Workshop on Acoustic Signal Enhancement (IWAENC)},
      • year = 2026,
      • month = sep,
      • url = {https://www.merl.com/publications/TR2026-126}
      • }
  • MERL Contacts:
  • Research Areas:

    Artificial Intelligence, Machine Learning, Speech & Audio

Abstract:

Task-aware unified source separation (TUSS) enables a single model to handle diverse separation tasks by conditioning on input prompts. However, conventional TUSS does not account for downstream task requirements, such as whether the enhanced speech will be used for human listening or automatic speech recognition (ASR). In this paper, we propose a prompt extension framework for TUSS that incorporates downstream task information into the input prompts and switches the loss function according to the given prompt during training, enabling outputs with different signal characteristics at inference time. Specifically, we introduce an ASR-dedicated prompt paired with a regularized loss function that reduces speech artifacts to improve ASR robustness, while the standard prompt is paired with the conventional SNR loss function. Experiments on the LibriSpeech and JNAS corpora demonstrate that the proposed joint-training scheme enables a single model to improve ASR performance over noisy input across a wide range of SNR conditions by selecting the ASR-dedicated prompt, while maintaining general speech enhancement quality when the standard prompt is used.