DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech
arXiv:2607.04140v1 Announce Type: cross Abstract: Autoregressive (AR) text-to-speech (TTS) models generate discrete speech tokens sequentially, which makes inference slow and can degrade robustness by propagating local errors and hallucinations. This limitation stems from their left-to-right AR commitment: each token must be determined before future speech-token context is available. However, such ordering is not an inherent requirement for TTS, as the full input text is available before synthes