Paper demo

RegularizedSchrödinger Bridgevia Distortion-Perception Perturbation for High-FidelitySpeech Enhancement

Qing Yao 1qyao@stmail.ujs.edu.cnLijian Gao 1ljgao@ujs.edu.cnQirong Mao 1mao_qr@ujs.edu.cnMing Dong 2mdong@wayne.edu

  1. 1 Jiangsu University
  2. 2 Wayne State University

Corresponding Author

Abstract

Speech enhancement (SE) requires high-fidelity reconstruction of clean speech that preserves linguistic and paralinguistic cues while maintaining high perceptual quality. Recently, Schrödinger Bridge (SB), a family of diffusion-based generative models, has advanced SE by bridging degraded and clean speech distributions in a principled formulation, enabling higher-quality reconstructions with fewer sampling steps. However, diffusion-based SE methods still face two challenges: (1) the fidelity-realism tradeoff, where they often prioritize perceptual realism encouraged by the learned speech prior, at the expense of fidelity; (2) the exposure bias issue, where iterative multi-step sampling causes early-step prediction errors to accumulate along the sampling trajectory and degrade enhanced speech quality. In this paper, we analyze standard SB training and show that it induces a systematic prediction drift, which biases the multi-step trajectory and amplifies error accumulation. To address this, we propose Regularized Schrödinger Bridge (RSB) for high-fidelity SE, a generative approach that reconciles fidelity and realism while mitigating exposure bias. RSB regularizes training with a Distortion-Perception Perturbation that constructs time-varying targets by interpolating between clean speech and posterior-mean estimates, and trains the network on perturbed intermediate states to correct toward the ground truth progressively. By simulating inference-time prediction errors, this perturbation mitigates the training-inference mismatch and thereby alleviates exposure bias. It also injects posterior-mean estimates as fidelity-preserving guidance, thereby improving reconstruction fidelity. Experiments on the WSJ0 corpus and the VoiceBank+DEMAND dataset demonstrate that RSB improves reconstruction fidelity over strong diffusion-based generative methods, yielding a favorable fidelity-realism tradeoff and reducing exposure bias.

Keywords

Speech enhancementSchrödinger bridgeDiffusion modelsDistortion-perception tradeoffExposure bias

Audio Samples

We publicly release 10 samples in total, drawn from two synthetic datasets:WSJ0+WHAM is the denoising dataset and WSJ0+REVERB is the dereverberation dataset.

WSJ0+WHAM

Denoising5 samples

Speech enhancement under additive environmental noise.

#1Sample name051o020a_c1_454_snr=5.2.wavSNR5.2 dBSignal-to-noise ratio (SNR) measures speech strength relative to background noise. Higher is cleaner.
Methods
Measurement
Measurement
0:00
-0:00
100
0–200

WSJ0+REVERB

Dereverberation5 samples

Speech enhancement under simulated room reverberation.

#1Sample name050a050v_c1_152_t60=1.57.wavT601.57 sReverberation time (T60) is the time for sound to decay by 60 dB. Lower means less reverberation.
Methods
Measurement
Measurement
0:00
-0:00
100
0–200

Citation

BibTeX
@article{yao2026rsb,
  author  = {Yao, Qing and Gao, Lijian and Mao, Qirong and Dong, Ming},
  title   = {Regularized Schr{\"o}dinger Bridge via Distortion-Perception Perturbation for High-Fidelity Speech Enhancement},
  journal = {IEEE Transactions on Audio, Speech and Language Processing},
  year    = {2026},
  volume  = {34},
  pages   = {3886-3900},
  doi     = {10.1109/TASLPRO.2026.3717234}
}