Tencent AuK-Flash is a 1.5B foundation model that generates and edits audio using simple natural language text instructions in just 4 fast steps.
| Step 1: Input | Step 2: AuK 4-Step Action | Step 3: Result |
|---|---|---|
| Provide 33s WAV audio and a natural text instruction. | 4-step distilled flow transforms speech semantics and acoustics. | Receive enhanced studio WAV audio instantly saved in outputs/. |
- 4-Step Fast Distillation: Speeds up audio sampling from 25 steps down to 4 steps for 4.5x faster throughput.
- Natural Language Control: Edit audio by typing everyday sentences like "replace word X with Y" or "deliver with high energy".
- All-in-One Audio Engine: Single foundation model handles word replacement, tone transfer, and noise cleanup.
- CUDA Acceleration: Automatic GPU device memory detection and PyTorch tensor batch execution.
Run this single PowerShell command to install Python dependencies and download model checkpoints:
pip install torch torchaudio huggingface_hub ; hf download tencent/AuK-Flash --local-dir ./ckpts/AuK-Flash ; hf download Qwen/Qwen2.5-Omni-3B --local-dir ./ckpts/Qwen2.5-Omni-3BExecute the main pipeline to process your input audio file across all target instruction tasks:
python main.py- Targeted Word Swapping: Replace misspoken words in recorded speech without re-recording full paragraphs.
- Creator Tone Conversion: Transform neutral voice narration into energetic, engaging video voiceovers.
- Podcast Noise Cleanup: Eliminate background mic hiss, fan hums, and room reverb for clean clarity.
- Zero-Shot Voice Cloning: Generate new spoken sentences matching a target speaker's vocal timbre.
- Vocal Pitch & Speed Tuning: Adjust vocal pitch by semitones or alter speaking rate for video dubbing.
- Sub-100ms Live Streaming: Real-time streaming audio inference for interactive voice assistants.
- Multi-Speaker Isolation: Advanced audio separation for overlapping multi-person podcast conversations.
- Interactive Web Dashboard: Visual browser waveform editor for word-level speech selection.
- Low-VRAM Quantization: GGUF and INT4 model quants for fast execution on consumer laptops.
- Singing Lyric Editing: Rewrite lyrics in singing recordings while preserving original melody.
Tencent-AuK-Flash/
├── main.py
└── README.md
Tencent AuK-Flash AuK-Flash Audio Generation AI Audio Editing AI Speech Editing Text to Audio Voice Editing AI AI Sound Editing Python Audio AI PyTorch Audio