Supported tasks
- Speech Generation: Zero-shot TTS · Instruct TTS
- Content Editing: Lyric Editing · Speech Content Editing
- Acoustic Editing: Pitch Editing · Speed Editing · Volume Editing
- Paralinguistic Editing: Emotion · Timbre · De-accent · Nonverbal Editing · Whisper Conversion
- Enhancement & Separation: Speech Enhancement · Speech Separation · Music Separation
Prompt Enhancer is optional and enabled by default. With it enabled, 0 estimates duration automatically and a value above 0 overrides it. With it disabled, an explicit duration greater than 0 seconds is required, including when reference audio or text is provided.
If the Hugging Face Space is slow or unavailable, try the ModelScope Studio instead.