
Abstract
Diffusion and flow models for speech enhancement face a training–inference mismatch: training uses analytical path states, while inference operates on self-generated trajectories, where prediction and discretization errors accumulate. We introduce Corrective Forcing (CoF), a unified post-training framework that adapts models to these rollout states. CoF corrects clean-speech predictions toward the ground truth under varying sampling schedules and regularizes local transitions against locally corrected counterfactual references. A shared clean-speech prediction parameterization enables the same objective for diffusion and flow models. Experiments with SB-VE and OT-CFM on three benchmarks show gains in perceptual quality and reconstruction fidelity across sampling budgets.
This paper has been submitted to ICASSP 2027 and is under review. This demo is coming soon.
Key ConceptsChallenges
Index Terms
Generative speech enhancementSchrödinger bridgeFlow matchingPost-training methodsExposure bias