Skip to content

About

Tencent Just Built a 1.5B AI That Can Generate AND Edit Audio! (AuK-Flash Setup) - Open-source 1.5B audio generation and editing foundation model setup, PyTorch pipeline, and 4-step fast distillation guide.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

⚡ Tencent AuK Flash Audio Suite

License: MIT Python 3.10+ Model: AuK-Flash PyTorch CUDA

Tencent AuK-Flash is a 1.5B foundation model that generates and edits audio using simple natural language text instructions in just 4 fast steps.

🔄 3-Step Audio Workflow

Step 1: Input Step 2: AuK 4-Step Action Step 3: Result
Provide 33s WAV audio and a natural text instruction. 4-step distilled flow transforms speech semantics and acoustics. Receive enhanced studio WAV audio instantly saved in outputs/.

🌟 Key Features & Superpowers

  • 4-Step Fast Distillation: Speeds up audio sampling from 25 steps down to 4 steps for 4.5x faster throughput.
  • Natural Language Control: Edit audio by typing everyday sentences like "replace word X with Y" or "deliver with high energy".
  • All-in-One Audio Engine: Single foundation model handles word replacement, tone transfer, and noise cleanup.
  • CUDA Acceleration: Automatic GPU device memory detection and PyTorch tensor batch execution.

🛠️ Quick Installation & Setup

Run this single PowerShell command to install Python dependencies and download model checkpoints:

pip install torch torchaudio huggingface_hub ; hf download tencent/AuK-Flash --local-dir ./ckpts/AuK-Flash ; hf download Qwen/Qwen2.5-Omni-3B --local-dir ./ckpts/Qwen2.5-Omni-3B

🚀 How to Run

Execute the main pipeline to process your input audio file across all target instruction tasks:

python main.py

🎯 5 Practical Real-World Use Cases

  1. Targeted Word Swapping: Replace misspoken words in recorded speech without re-recording full paragraphs.
  2. Creator Tone Conversion: Transform neutral voice narration into energetic, engaging video voiceovers.
  3. Podcast Noise Cleanup: Eliminate background mic hiss, fan hums, and room reverb for clean clarity.
  4. Zero-Shot Voice Cloning: Generate new spoken sentences matching a target speaker's vocal timbre.
  5. Vocal Pitch & Speed Tuning: Adjust vocal pitch by semitones or alter speaking rate for video dubbing.

🔮 5 Planned Future Features

  1. Sub-100ms Live Streaming: Real-time streaming audio inference for interactive voice assistants.
  2. Multi-Speaker Isolation: Advanced audio separation for overlapping multi-person podcast conversations.
  3. Interactive Web Dashboard: Visual browser waveform editor for word-level speech selection.
  4. Low-VRAM Quantization: GGUF and INT4 model quants for fast execution on consumer laptops.
  5. Singing Lyric Editing: Rewrite lyrics in singing recordings while preserving original melody.

📁 Project File Structure

Tencent-AuK-Flash/
├── main.py
└── README.md

🔑 Keywords & Search Tags

Tencent AuK-Flash AuK-Flash Audio Generation AI Audio Editing AI Speech Editing Text to Audio Voice Editing AI AI Sound Editing Python Audio AI PyTorch Audio

About

Tencent Just Built a 1.5B AI That Can Generate AND Edit Audio! (AuK-Flash Setup) - Open-source 1.5B audio generation and editing foundation model setup, PyTorch pipeline, and 4-step fast distillation guide.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages