Paper demo
RegularizedSchrödinger Bridgevia Distortion-Perception Perturbation for High-FidelitySpeech Enhancement
- 1 Jiangsu University
- 2 Wayne State University
† Corresponding Author
LINKS
Abstract
Speech enhancement (SE) requires high-fidelity reconstruction of clean speech that preserves linguistic and paralinguistic cues while maintaining high perceptual quality. Recently, Schrödinger Bridge (SB), a family of diffusion-based generative models, has advanced SE by bridging degraded and clean speech distributions in a principled formulation, enabling higher-quality reconstructions with fewer sampling steps. However, diffusion-based SE methods still face two challenges: (1) the fidelity-realism tradeoff, where they often prioritize perceptual realism encouraged by the learned speech prior, at the expense of fidelity; (2) the exposure bias issue, where iterative multi-step sampling causes early-step prediction errors to accumulate along the sampling trajectory and degrade enhanced speech quality. In this paper, we analyze standard SB training and show that it induces a systematic prediction drift, which biases the multi-step trajectory and amplifies error accumulation. To address this, we propose Regularized Schrödinger Bridge (RSB) for high-fidelity SE, a generative approach that reconciles fidelity and realism while mitigating exposure bias. RSB regularizes training with a Distortion-Perception Perturbation that constructs time-varying targets by interpolating between clean speech and posterior-mean estimates, and trains the network on perturbed intermediate states to correct toward the ground truth progressively. By simulating inference-time prediction errors, this perturbation mitigates the training-inference mismatch and thereby alleviates exposure bias. It also injects posterior-mean estimates as fidelity-preserving guidance, thereby improving reconstruction fidelity. Experiments on the WSJ0 corpus and the VoiceBank+DEMAND dataset demonstrate that RSB improves reconstruction fidelity over strong diffusion-based generative methods, yielding a favorable fidelity-realism tradeoff and reducing exposure bias.
Keywords
Audio Samples
We publicly release 10 samples in total, drawn from two synthetic datasets:WSJ0+WHAM is the denoising dataset and WSJ0+REVERB is the dereverberation dataset.
WSJ0+WHAM
Denoising5 samplesSpeech enhancement under additive environmental noise.
WSJ0+REVERB
Dereverberation5 samplesSpeech enhancement under simulated room reverberation.
Citation
@article{yao2026rsb,
author = {Yao, Qing and Gao, Lijian and Mao, Qirong and Dong, Ming},
title = {Regularized Schr{\"o}dinger Bridge via Distortion-Perception Perturbation for High-Fidelity Speech Enhancement},
journal = {IEEE Transactions on Audio, Speech and Language Processing},
year = {2026},
volume = {34},
pages = {3886-3900},
doi = {10.1109/TASLPRO.2026.3717234}
}