TR2026-137
Speaker Identity as Sole Supervision for Speech Separation
-
- , "Speaker Identity as Sole Supervision for Speech Separation", Interspeech, September 2026.BibTeX TR2026-137 PDF
- @inproceedings{Boeddeker2026sep,
- author = {Boeddeker, Christoph and Masuyama, Yoshiki and Richter, Julius and Edo, Takahiro and Wichern, Gordon and {Le Roux}, Jonathan},
- title = {{Speaker Identity as Sole Supervision for Speech Separation}},
- booktitle = {Interspeech},
- year = 2026,
- month = sep,
- url = {https://www.merl.com/publications/TR2026-137}
- }
- , "Speaker Identity as Sole Supervision for Speech Separation", Interspeech, September 2026.
-
MERL Contacts:
-
Research Areas:
Abstract:
Speech separation systems are typically trained with waveformor spectrogram-level reconstruction losses requiring clean references, spatial cues such as multichannel recordings, or remixing strategies. We investigate whether speech separation can instead be trained using speaker identity as the sole supervision signal. We propose a contrastive objective aligning embeddings of separated outputs with embeddings of auxiliary utterances while repelling competing speakers using in-batch negatives. Unlike target speaker extraction, embeddings are used only to define the training objective and are not required at inference, resulting in a conventional speaker-independent separator. Experiments show that speaker-identity supervision alone can train separation systems from scratch to satisfactory performance and that fine-tuning supervised models on noisy mixtures with this objective further improves separation quality.




