Julius Richter
- Email:
-
Position:
Research / Technical Staff
Visiting Research Scientist -
Education:
Ph.D. in Computer Science, University of Hamburg, Germany, 2025 -
Research Areas:
External Links:
Julius' Quick Links
-
Biography
Julius's research interests include generative models and multimodal learning for audio-visual understanding and restoration. During his Ph.D., he developed novel diffusion-based generative approaches for single-channel speech enhancement. Prior to joining MERL, he was a postdoctoral researcher at Meta Superintelligence Labs.
-
Awards
-
AWARD MERL Team Wins Real-TSE Challenge Track 2 on Offline Target Speaker Extraction Date: July 6, 2026
Awarded to: Dominik Klement, Yoshiki Masuyama, Christoph Boeddeker, Kohei Saijo, Julius Richter, Gordon Wichern, and Jonathan Le Roux
MERL Contacts: Christoph Boeddeker; Jonathan Le Roux; Yoshiki Masuyama; Julius Richter; Gordon Wichern
Research Areas: Artificial Intelligence, Machine Learning, Speech & AudioBriefMERL's Speech & Audio team, led by MERL intern Dominik Klement, ranked 1st out of 11 teams in Track 2, "Offline Target Speaker Extraction," of the Real-TSE Challenge. The challenge focuses on target speaker extraction (TSE) from real-world conversational recordings in either English or Chinese, where the goal is to extract the speech of a target speaker in the presence of interfering speakers, background noise, and reverberation.
While modern TSE systems have achieved strong performance on simulated speech mixtures, their performance can degrade considerably on real-world recordings due to the mismatch between simulated training data and actual conversational environments. The Real-TSE Challenge was designed to advance TSE under these realistic conditions, using real far-field conversational recordings for evaluation.
The MERL team won Track 2 by focusing on training data and curriculum learning rather than introducing a new model architecture. Starting from a strong speech separation model, the team progressively trained the system on fully overlapping synthetic speech, simulated conversations, realistic far-field mixtures, and finally real conversational recordings. This approach reduced the token error rate (TER), measured at either the word (English) or character (Chinese) level, from 70% to 37% on the development set and achieved a final TER of 61.3% on the evaluation set, best among the 11 participating teams. The team also topped the leaderboard in terms of the aggregate ranking across the four measures evaluating intelligibility, target speaker presence rate, speaker similarity, and perceptual quality.
The team also investigated the reliability of the challenge metrics and demonstrated that neural network-based speaker similarity and predicted speech-quality scores could be substantially improved without a corresponding improvement in perceptual quality. Because learned metrics can be susceptible to adversarial attacks or optimization that exploits weaknesses in the metric itself, these findings highlight both the importance of realistic training data for real-world TSE and the need for robust evaluation metrics when developing speech extraction systems.
A paper summarizing the team's findings will be presented at the IEEE Spoken Language Technology (SLT) 2026 workshop, to be held in Palermo, Italy from December 13-16, 2026.
REAL-TSE Challenge: Track 2 rankings — Offline Target Speaker Extraction Rank Team TER ↓ F1 ↑ SIM ↑ P808 ↑ Score ↓ 1 MERL 0.613 (1) 0.861 (2) 0.538 (3) 3.371 (2) 2.00 2 YiJiaHe 0.639 (2) 0.871 (1) 0.565 (1) 3.128 (9) 3.25 3 CARTSE 0.651 (3) 0.857 (4) 0.544 (2) 3.138 (8) 4.25 4 WasedaM 0.675 (5) 0.858 (3) 0.480 (6) 3.232 (6) 5.00 5 SonicAGI 0.680 (6) 0.851 (6) 0.471 (7) 3.258 (5) 6.00 6 WAKA 0.670 (4) 0.847 (8) 0.471 (7) 3.150 (7) 6.50 6 SHNU-TSE 0.731 (9) 0.840 (9) 0.507 (5) 3.362 (3) 6.50 7 ChuEst 0.710 (7) 0.831 (11) 0.532 (4) 3.064 (10) 8.00 8 pyannoteAI 0.728 (8) 0.855 (5) 0.464 (9) 2.904 (12) 8.50 9 AGH-JHU 0.743 (10) 0.837 (10) 0.434 (11) 3.335 (4) 8.75 10 WHU_IASP 0.757 (11) 0.850 (7) 0.465 (8) 2.961 (11) 9.25 11 CUDA_OUT_OF_MEMORY 0.827 (12) 0.819 (13) 0.364 (13) 3.435 (1) 9.75 12 BSRNN_EMB Baseline 0.829 (13) 0.829 (12) 0.417 (12) 2.875 (13) 12.50 12 BSRNN_TFMAP Baseline 0.838 (14) 0.829 (12) 0.443 (10) 2.756 (14) 12.50 ↓ Lower is better; ↑ higher is better. Parentheses show metric ranks. The score is the average of the four dense metric ranks; tied scores share a position. Best metric values are bold. P808 denotes DNSMOS-P808.
Source: Official REAL-TSE Challenge rankings. BSRNN entries are organizer baselines.
-
AWARD MERL Team Wins DCASE 2026 Challenge on Anomalous Sound Detection for Machine Condition Monitoring Date: June 30, 2026
Awarded to: Takuya Fujimura, Gordon Wichern, Yoshiki Masuyama, Christoph Boeddeker, Kohei Saijo, Julius Richter, Takahiro Edo, and Jonathan Le Roux
MERL Contacts: Christoph Boeddeker; Jonathan Le Roux; Yoshiki Masuyama; Julius Richter; Gordon Wichern
Research Areas: Artificial Intelligence, Machine Learning, Signal Processing, Speech & AudioBrief- MERL's Speech & Audio team ranked 1st out of 51 teams in the DCASE 2026 Challenge’s Task 2, “Noise-aware Unsupervised Anomalous Sound Detection for Machine Condition Monitoring.” The team was led by MERL intern Takuya Fujimura, and also included Gordon Wichern, Yoshiki Masuyama, Christoph Boeddeker, Kohei Saijo, Julius Richter, Takahiro Edo, and Jonathan Le Roux.
The IEEE AASP Challenge on Detection and Classification of Acoustic Scenes and Events (DCASE Challenge), started in 2013, has been organized yearly since 2016, and gathers challenges on multiple tasks related to the detection, analysis, and generation of sound events. This year, the DCASE 2026 Challenge received 421 submissions from 135 teams across seven tasks.
The MERL team won Task 2, Noise-aware Unsupervised Anomalous Sound Detection for Machine Condition Monitoring, which aims at building noise-robust systems for automatically detecting machine failure via microphones when only normal machine operating data is available for system development. Task 2 was by far the most popular out of the 7 DCASE 2026 tasks, with 51 teams submitting 168 entries. The MERL team's system was built around MERL’s recently proposed paradigm of noise-aware self-supervised learning, which extracts noise robust features leveraging two-channel recordings, in which one microphone is used to capture noise. Anomaly detection is then performed in the extracted denoised feature space using advanced score normalization. The team's best submission obtained a composite score of 70.24% on five evaluation machines, largely outperforming the 2nd best team's 65.45%.
MERL also participated in Task 4, Spatial Semantic Segmentation of Sound Scenes (S5) and placed 3rd out of 10 teams in separation performance. Our cascaded system consists of universal sound separation with source counting, source classification, and class-aware refinement, where the separation and refinement modules are built upon MERL's TF-Locoformer separation technology. Notably, the team's best submission obtained a label prediction accuracy of 76.92% on the evaluation set, largely outperforming the 2nd best team's 65.54%.
- MERL's Speech & Audio team ranked 1st out of 51 teams in the DCASE 2026 Challenge’s Task 2, “Noise-aware Unsupervised Anomalous Sound Detection for Machine Condition Monitoring.” The team was led by MERL intern Takuya Fujimura, and also included Gordon Wichern, Yoshiki Masuyama, Christoph Boeddeker, Kohei Saijo, Julius Richter, Takahiro Edo, and Jonathan Le Roux.
-
-
MERL Publications
- , "Downstream-Task-Aware Unified Source Separation", International Workshop on Acoustic Signal Enhancement (IWAENC), September 2026.BibTeX TR2026-126 PDF
- @inproceedings{Mitsui2026sep,
- author = {Mitsui, Yoshiki and Aihara, Ryo and Saito, Tatsuhiko and Masuyama, Yoshiki and Boeddeker, Christoph and Richter, Julius and Wichern, Gordon and {Le Roux}, Jonathan},
- title = {{Downstream-Task-Aware Unified Source Separation}},
- booktitle = {International Workshop on Acoustic Signal Enhancement (IWAENC)},
- year = 2026,
- month = sep,
- url = {https://www.merl.com/publications/TR2026-126}
- }
- , "NABEATs: Noise-Aware Audio Representation Learning", International Workshop on Acoustic Signal Enhancement (IWAENC), September 2026.BibTeX TR2026-124 PDF
- @inproceedings{Fujimura2026sep,
- author = {Fujimura, Takuya and Masuyama, Yoshiki and Wichern, Gordon and Boeddeker, Christoph and Richter, Julius and {Le Roux}, Jonathan},
- title = {{NABEATs: Noise-Aware Audio Representation Learning}},
- booktitle = {International Workshop on Acoustic Signal Enhancement (IWAENC)},
- year = 2026,
- month = sep,
- url = {https://www.merl.com/publications/TR2026-124}
- }
- , "Few-Shot Room Impulse Response Interpolation in Latent Domains", International Workshop on Acoustic Signal Enhancement (IWAENC), September 2026.BibTeX TR2026-125 PDF
- @inproceedings{Lin2026sep,
- author = {Lin, Jackie and Masuyama, Yoshiki and Boeddeker, Christoph and Richter, Julius and Wichern, Gordon and Kim, Minje and {Le Roux}, Jonathan},
- title = {{Few-Shot Room Impulse Response Interpolation in Latent Domains}},
- booktitle = {International Workshop on Acoustic Signal Enhancement (IWAENC)},
- year = 2026,
- month = sep,
- url = {https://www.merl.com/publications/TR2026-125}
- }
- , "Technical Report for MERL’s Real-TSE Challenge Submission," Tech. Rep. TR2026-112, Mitsubishi Electric Research Laboratories, July 2026.BibTeX TR2026-112 PDF
- @techreport{Klement2026jul2,
- author = {Klement, Dominik and Masuyama, Yoshiki and Boeddeker, Christoph and Saijo, Kohei and Richter, Julius and Wichern, Gordon and {Le Roux}, Jonathan},
- title = {{Technical Report for MERL’s Real-TSE Challenge Submission}},
- institution = {Real-TSE Challenge},
- year = 2026,
- month = jul,
- url = {https://www.merl.com/publications/TR2026-112}
- }
- , "The MERL Systems for DCASE 2026 Challenge Task 2," Tech. Rep. TR2026-100, IEEE AASP Challenge on Detection and Classification of Acoustic Scenes and Events (DCASE Challenge), June 2026.BibTeX TR2026-100 PDF
- @techreport{Fujimura2026jun,
- author = {{Fujimura, Takuya and Wichern, Gordon and Masuyama, Yoshiki and Boeddeker, Christoph and Saijo, Kohei and Richter, Julius and Edo, Takahiro and Le Roux, Jonathan}},
- title = {{The MERL Systems for DCASE 2026 Challenge Task 2}},
- institution = {IEEE AASP Challenge on Detection and Classification of Acoustic Scenes and Events (DCASE Challenge)},
- year = 2026,
- month = jun,
- url = {https://www.merl.com/publications/TR2026-100}
- }
- , "Downstream-Task-Aware Unified Source Separation", International Workshop on Acoustic Signal Enhancement (IWAENC), September 2026.
-
Other Publications
- , "Do We Need EMA for Diffusion-Based Speech Enhancement? Toward a Magnitude-Preserving Network Architecture", Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2026.BibTeX
- @Inproceedings{Richter2026ICASSPEDM2SE,
- author = {Richter, Julius and de Oliveira, Danilo and Gerkmann, Timo},
- title = {Do We Need EMA for Diffusion-Based Speech Enhancement? Toward a Magnitude-Preserving Network Architecture},
- booktitle = {Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP)},
- year = 2026
- }
- , "Diffusion Models for Audio Restoration", IEEE Signal Processing Magazine, Vol. 41, No. 6, pp. 72-84, 2025.BibTeX
- @Article{Lemercier2025SPMDiffusion,
- author = {Lemercier, Jean-Marie and Richter, Julius and Welker, Simon and Moliner, Eloi and V{\"a}lim{\"a}ki, Vesa and Gerkmann, Timo},
- title = {Diffusion Models for Audio Restoration},
- journal = {IEEE Signal Processing Magazine},
- year = 2025,
- volume = 41,
- number = 6,
- pages = {72--84}
- }
- , "Investigating Training Objectives for Generative Speech Enhancement", Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2025.BibTeX
- @Inproceedings{Richter2025ICASSPObjectives,
- author = {Richter, Julius and de Oliveira, Danilo and Gerkmann, Timo},
- title = {Investigating Training Objectives for Generative Speech Enhancement},
- booktitle = {Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP)},
- year = 2025
- }
- , "ReverbFX: A Dataset of Room Impulse Responses Derived from Reverb Effect Plugins for Singing Voice Dereverberation", Proceedings of the ITG Conference on Speech Communication, 2025.BibTeX
- @Inproceedings{Richter2025ITGReverbFX,
- author = {Richter, Julius and Svajda, Till and Gerkmann, Timo},
- title = {{ReverbFX}: A Dataset of Room Impulse Responses Derived from Reverb Effect Plugins for Singing Voice Dereverberation},
- booktitle = {Proceedings of the ITG Conference on Speech Communication},
- year = 2025
- }
- , "Non-intrusive Speech Quality Assessment with Diffusion Models Trained on Clean Speech", Proceedings of Interspeech, 2025.BibTeX
- @Inproceedings{deOliveira2025InterspeechLikelihood,
- author = {de Oliveira, Danilo and Richter, Julius and Lemercier, Jean-Marie and Welker, Simon and Gerkmann, Timo},
- title = {Non-intrusive Speech Quality Assessment with Diffusion Models Trained on Clean Speech},
- booktitle = {Proceedings of Interspeech},
- year = 2025
- }
- , "Single and Few-step Diffusion for Generative Speech Enhancement", Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2024.BibTeX
- @Inproceedings{Lay2024ICASSPFewStep,
- author = {Lay, Bunlong and Lemercier, Jean-Marie and Richter, Julius and Gerkmann, Timo},
- title = {Single and Few-step Diffusion for Generative Speech Enhancement},
- booktitle = {Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP)},
- year = 2024
- }
- , "EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation", Proceedings of Interspeech, 2024.BibTeX
- @Inproceedings{Richter2024InterspeechEARS,
- author = {Richter, Julius and Wu, Yi-Chiao and Krenn, Steven and Welker, Simon and Lay, Bunlong and Watanabe, Shinji and Richard, Alexander and Gerkmann, Timo},
- title = {{EARS}: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation},
- booktitle = {Proceedings of Interspeech},
- year = 2024
- }
- , "Diffusion-based Speech Enhancement: Demonstration of Performance and Generalization", Audio Imagination Workshop at NeurIPS, 2024.BibTeX
- @Inproceedings{Richter2024NeurIPSAudioImagination,
- author = {Richter, Julius and Gerkmann, Timo},
- title = {Diffusion-based Speech Enhancement: Demonstration of Performance and Generalization},
- booktitle = {Audio Imagination Workshop at NeurIPS},
- year = 2024
- }
- , "Causal Diffusion Models for Generalized Speech Enhancement", IEEE Open Journal of Signal Processing, Vol. 5, pp. 780-789, 2024.BibTeX
- @Article{Richter2024OJSPCausal,
- author = {Richter, Julius and Welker, Simon and Lemercier, Jean-Marie and Lay, Bunlong and Peer, Tal and Gerkmann, Timo},
- title = {Causal Diffusion Models for Generalized Speech Enhancement},
- journal = {IEEE Open Journal of Signal Processing},
- year = 2024,
- volume = 5,
- pages = {780--789}
- }
- , "Reducing the Prior Mismatch of Stochastic Differential Equations for Diffusion-based Speech Enhancement", Proceedings of Interspeech, 2023.BibTeX
- @Inproceedings{Lay2023InterspeechPriorMismatch,
- author = {Lay, Bunlong and Welker, Simon and Richter, Julius and Gerkmann, Timo},
- title = {Reducing the Prior Mismatch of Stochastic Differential Equations for Diffusion-based Speech Enhancement},
- booktitle = {Proceedings of Interspeech},
- year = 2023
- }
- , "StoRM: A Diffusion-based Stochastic Regeneration Model for Speech Enhancement and Dereverberation", IEEE/ACM Transactions on Audio, Speech, and Language Processing, Vol. 31, pp. 2724-2737, 2023.BibTeX
- @Article{Lemercier2023TASLPStoRM,
- author = {Lemercier, Jean-Marie and Richter, Julius and Welker, Simon and Gerkmann, Timo},
- title = {{StoRM}: A Diffusion-based Stochastic Regeneration Model for Speech Enhancement and Dereverberation},
- journal = {IEEE/ACM Transactions on Audio, Speech, and Language Processing},
- year = 2023,
- volume = 31,
- pages = {2724--2737}
- }
- , "Audio-Visual Speech Separation in Noisy Environments with a Lightweight Iterative Model", Proceedings of Interspeech, 2023.BibTeX
- @Inproceedings{Martel2023InterspeechAVSep,
- author = {Martel, Hector and Richter, Julius and Li, Kai and Hu, Xiaolin and Gerkmann, Timo},
- title = {Audio-Visual Speech Separation in Noisy Environments with a Lightweight Iterative Model},
- booktitle = {Proceedings of Interspeech},
- year = 2023
- }
- , "Speech Signal Improvement Using Causal Generative Diffusion Models", Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2023.BibTeX
- @Inproceedings{Richter2023ICASSPSpeechImprovement,
- author = {Richter, Julius and Welker, Simon and Lemercier, Jean-Marie and Lay, Bunlong and Peer, Tal and Gerkmann, Timo},
- title = {Speech Signal Improvement Using Causal Generative Diffusion Models},
- booktitle = {Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP)},
- year = 2023
- }
- , "Audio-Visual Speech Enhancement with Score-Based Generative Models", Proceedings of the ITG Conference on Speech Communication, 2023.BibTeX
- @Inproceedings{Richter2023ITGAVScore,
- author = {Richter, Julius and Frintrop, Simone and Gerkmann, Timo},
- title = {Audio-Visual Speech Enhancement with Score-Based Generative Models},
- booktitle = {Proceedings of the ITG Conference on Speech Communication},
- year = 2023
- }
- , "Speech Enhancement and Dereverberation with Diffusion-Based Generative Models", IEEE/ACM Transactions on Audio, Speech, and Language Processing, Vol. 31, pp. 2351-2364, 2023.BibTeX
- @Article{Richter2023TASLPDiffusion,
- author = {Richter, Julius and Welker, Simon and Lemercier, Jean-Marie and Lay, Bunlong and Gerkmann, Timo},
- title = {Speech Enhancement and Dereverberation with Diffusion-Based Generative Models},
- journal = {IEEE/ACM Transactions on Audio, Speech, and Language Processing},
- year = 2023,
- volume = 31,
- pages = {2351--2364}
- }
- , "On the Behavior of Intrusive and Non-intrusive Speech Enhancement Metrics in Predictive and Generative Settings", Proceedings of the ITG Conference on Speech Communication, 2023.BibTeX
- @Inproceedings{deOliveira2023ITGMetrics,
- author = {de Oliveira, Danilo and Richter, Julius and Lemercier, Jean-Marie and Peer, Tal and Gerkmann, Timo},
- title = {On the Behavior of Intrusive and Non-intrusive Speech Enhancement Metrics in Predictive and Generative Settings},
- booktitle = {Proceedings of the ITG Conference on Speech Communication},
- year = 2023
- }
- , "Continuous Phoneme Recognition based on Audio-Visual Modality Fusion", Proceedings of the IEEE World Congress on Computational Intelligence, 2022.BibTeX
- @Inproceedings{Richter2022WCCIAVPhoneme,
- author = {Richter, Julius and Liebold, Jeanine and Gerkmann, Timo},
- title = {Continuous Phoneme Recognition based on Audio-Visual Modality Fusion},
- booktitle = {Proceedings of the IEEE World Congress on Computational Intelligence},
- year = 2022
- }
- , "Speech Enhancement with Score-Based Generative Models in the Complex STFT Domain", Proceedings of Interspeech, 2022.BibTeX
- @Inproceedings{Welker2022InterspeechComplex,
- author = {Welker, Simon and Richter, Julius and Gerkmann, Timo},
- title = {Speech Enhancement with Score-Based Generative Models in the Complex {STFT} Domain},
- booktitle = {Proceedings of Interspeech},
- year = 2022
- }
- , "Guided Variational Autoencoder for Speech Enhancement with a Supervised Classifier", Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2021.BibTeX
- @Inproceedings{Carbajal2021ICASSPGuidedVAE,
- author = {Carbajal, Guillaume and Richter, Julius and Gerkmann, Timo},
- title = {Guided Variational Autoencoder for Speech Enhancement with a Supervised Classifier},
- booktitle = {Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP)},
- year = 2021
- }
- , "Disentanglement Learning for Variational Autoencoders Applied to Audio-Visual Speech Enhancement", Proceedings of the IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 2021.BibTeX
- @Inproceedings{Carbajal2021WASPAA,
- author = {Carbajal, Guillaume and Richter, Julius and Gerkmann, Timo},
- title = {Disentanglement Learning for Variational Autoencoders Applied to Audio-Visual Speech Enhancement},
- booktitle = {Proceedings of the IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA)},
- year = 2021
- }
- , "Improving Mix-and-Separate Training in Audio-Visual Sound Source Separation with an Object Prior", Proceedings of the International Conference on Pattern Recognition (ICPR), 2020.BibTeX
- @Inproceedings{Nguyen2020ICPRAVPrior,
- author = {Nguyen, Quan and Richter, Julius and Lauri, Mikko and Gerkmann, Timo and Frintrop, Simone},
- title = {Improving Mix-and-Separate Training in Audio-Visual Sound Source Separation with an Object Prior},
- booktitle = {Proceedings of the International Conference on Pattern Recognition (ICPR)},
- year = 2020
- }
- , "Speech Enhancement with Stochastic Temporal Convolutional Networks", Proceedings of Interspeech, 2020.BibTeX
- @Inproceedings{Richter2020InterspeechTCN,
- author = {Richter, Julius and Carbajal, Guillaume and Gerkmann, Timo},
- title = {Speech Enhancement with Stochastic Temporal Convolutional Networks},
- booktitle = {Proceedings of Interspeech},
- year = 2020
- }
- , "Do We Need EMA for Diffusion-Based Speech Enhancement? Toward a Magnitude-Preserving Network Architecture", Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2026.
-
Software & Data Downloads