AuK: An Open-Source Foundational Model for Speech Generation and Editing

Supported tasks

  • Speech Generation: Zero-shot TTS · Instruct TTS
  • Content Editing: Lyric Editing · Speech Content Editing
  • Acoustic Editing: Pitch Editing · Speed Editing · Volume Editing
  • Paralinguistic Editing: Emotion · Timbre · De-accent · Nonverbal Editing · Whisper Conversion
  • Enhancement & Separation: Speech Enhancement · Speech Separation · Music Separation

Prompt Enhancer is optional and enabled by default. With it enabled, 0 estimates duration automatically and a value above 0 overrides it. With it disabled, an explicit duration greater than 0 seconds is required, including when reference audio or text is provided.

Enabled by default: prepares the task and instruction. Keep duration at 0 for automatic estimation, or enter a value to override it. When disabled, a duration greater than 0 is required.
Model
0 30

With Prompt Enhancer: 0 = automatic estimation; values above 0 override the duration. Without Prompt Enhancer: enter a value greater than 0.

4 64
0 5

Base: NFE 32 / CFG 2.0 · Flash: NFE 4 / CFG 0 (locked)

Prompt Enhancer output

Examples

Select an example to load it into the panel, then click Generate. You can also try different seeds.