A planning, long-form and verification layer for AuK, Tencent Hunyuan's open-source speech generation and editing model.
AuK is excellent at what it does: one natural-language instruction, one clip, one edit. OpenSpeech is everything you need around it to ship real audio.
openspeech run -i interview.wav \
-r "clean up the noise and reverb, keep only the second speaker, calm the delivery, and master to -16 LUFS" \
-o episode.wavThat single command is four operations, ordered into a proper signal chain, split into model-sized pieces across a 40-minute recording, verified after every step, and mastered to a broadcast loudness target. AuK on its own does none of those four things.
These are not hypotheticals. Each one is a specific, cited limitation of AuK 0.1.0.
| What you hit | Where it comes from | What OpenSpeech does |
|---|---|---|
| One task per request. Ask for two things and one is silently discarded. | AuK's Prompt Enhancer instructs its classifier: "when a request contains several unrelated intents, take only the main one" (pe.config.yaml). |
Decomposes a request into an ordered plan of operations, and reports anything it could not match instead of dropping it quietly. |
| A ~30 second ceiling. And it is not 30 seconds of audio — the window holds the source and the generated output. | "source/reference + generated target must fit within 30 [seconds]" (docs/COMFYUI.md). |
Computes the real per-operation budget, cuts at pauses, and stitches back with level-matched crossfades. An equal-length edit gets 13.95s per call; slowing speech to 0.5x gets 9.30s. |
Parameters snapped by an LLM prompt. 1.8x is asked to become a legal 2.0x. |
pe.config.yaml normalisation hints. |
Snaps in code, arithmetically, and tells you: speed_multiplier: 1.8 -> 2.0. |
| No verification. If a generation collapses to silence or ignores the instruction, nothing notices. | No quality signal in the API. | Measures every result against what the operation promised. A +10 dB edit is checked for +10 dB. A pitch shift is checked against the output's actual f0. |
| CLI and Gradio only. Nothing another service can call. | — | A JSON HTTP API with a job queue, progress streaming, and the model loaded once. |
OpenSpeech does not vendor, fork, or modify AuK. It calls AuK's documented Python API and quotes AuK's own instruction wording verbatim, because those strings are what the model was trained on.
git clone https://github.com/zwayth/openspeech
cd openspeech
pip install -e .The core — planning, DSP, chunking, quality control, CLI — depends on numpy alone. It runs and is fully tested with no GPU, no weights, and no torch.
pip install -e ".[service]" # HTTP API
pip install -e ".[llm]" # LLM-backed planner
pip install -e ".[accel]" # faster metering, more audio formats
pip install -e ".[dev]" # tests and lintingOpenSpeech ships a mock backend that approximates AuK with real DSP. It is how the test suite runs in CI, and it lets you build and debug an entire pipeline before downloading 16 GB. It is never confused for the real thing — openspeech run prints a notice, and every report carries engine: mock.
For real output, install AuK and its weights so that ckpts/AuK/auk_base.safetensors and its config.yaml exist, then:
pip install -e ".[auk]"
openspeech doctor # confirms which backend will be used--engine auto (the default) uses AuK when a checkpoint is present and the mock backend otherwise.
No weights needed — the sample generator makes its own audio.
python examples/make_samples.py
openspeech plan "clean up the noise and reverb, calm the delivery, master to -16 LUFS"
openspeech run -i examples/assets/host.wav -r "remove the noise and normalise to -16 LUFS" -o out.wav
openspeech measure out.wavopenspeech plan costs nothing and runs instantly. It is the fastest way to see how a phrasing is understood before spending model time on it:
$ openspeech plan "clean up the noise and reverb, make the host sound calmer, speed it up a bit, and normalise to -16 LUFS"
request : clean up the noise and reverb, make the host sound calmer, speed it up a bit, and normalise to -16 LUFS
language: en planner: rules steps: 4
1. [model] enhance_speech(cleanup_mode='denoise_dereverb')
2. [model] emotion_edit(emotion='calm')
3. [model] speed_edit(speed_multiplier=1.25)
4. [dsp ] loudness_normalize(target_lufs=-16.0, true_peak_db=-1.0)
enhance_speech instruction: "Preserve all speakers, remove noise and reverberation, and output clean speech of the same length."
emotion_edit instruction: "Change the emotion to calm."
speed_edit instruction: "Adjust the speech speed to 1.25x."
Note step 4: loudness normalisation runs locally in DSP, not through the model. A generative model is the wrong tool for arithmetic on gain — local is exact, instant, free of GPU time, and unbound by the context window.
from openspeech import process, read_audio, write_audio
result = process(read_audio("interview.wav"),
"remove the noise, calm the delivery, master to -16 LUFS")
write_audio("episode.wav", result.audio)
print(result.verdict) # pass | warn | fail
for step in result.steps:
print(step.op, step.chunks, step.verdict, step.repairs)Lower level, when you want to hold the parts yourself:
from openspeech import Pipeline, build_op, plan_request, create_engine
from openspeech.plan import Plan
plan = Plan(request="manual", ops=(
build_op("enhance_speech", cleanup_mode="denoise_dereverb"),
build_op("volume_edit", "increase", gain_db=10),
build_op("loudness_normalize", target_lufs=-16.0),
))
result = Pipeline(create_engine("auto")).run(plan, read_audio("in.wav"))A request becomes an ordered pipeline. Ordering is a quality decision, not bookkeeping: restoration runs first because every later model pass does better on clean input, and mastering runs last because any edit before it invalidates a loudness measurement. State the steps backwards and you still get the right order.
Two intents that each ask for half a job are merged into one model pass — "remove the noise, and get rid of the reverb" becomes a single denoise_dereverb call, not two.
The rule-based planner is offline, deterministic, and needs no API key. An LLM planner handles phrasings the rules miss — and every operation it proposes is rebuilt through the same registry, so a hallucinated task name or an illegal parameter is caught locally instead of becoming a bad instruction. See docs/planning.md.
The chunk budget is derived per operation from how much that operation expands the audio, because AuK's window has to hold the source and the target together:
| Operation | Output length | Source budget |
|---|---|---|
enhance_speech |
same | 13.95 s |
emotion_edit (sad) |
×1.22 | 12.57 s |
speed_edit 0.5× |
×2.0 | 9.30 s |
speed_edit 2.0× |
×0.5 | 18.60 s |
Cuts land in pauses, preferring the latest usable pause so the budget is actually spent — fewer calls, fewer seams. Where there is no pause, chunks overlap and are crossfaded. Chunks are level-matched against their own sources, which removes the drift a generative model introduces between independent calls while preserving the recording's real dynamics. Round-tripping unmodified audio through chunking and stitching is sample-exact. See docs/longform.md.
Some operations cannot be chunked, and OpenSpeech says so rather than producing quiet nonsense. A word-level edit can't be split across chunks (you'd edit the wrong one), so it refuses with an actionable message. Speaker separation by talking order is a property of the whole recording, so it warns and suggests identifying the speaker by content instead.
Every result is measured against what the operation promised:
[error] volume_applied: level moved +4.0 dB, asked for +10.0 dB
What happens next depends on the fault. A collapse to silence is a bad draw, so it is retried with a new seed and more sampling steps. A gain that landed 6 dB short is arithmetic — re-rolling the model wastes GPU time and lands somewhere else wrong, so OpenSpeech closes the gap exactly in DSP and says so:
repair: applied the residual +6.0 dB the model left on the table (asked +10.0 dB, got +4.0 dB)
DC offsets, inter-sample peaks and small length overruns are handled the same way. What cannot be repaired is reported honestly rather than papered over. See docs/quality.md.
The metering is not approximate. Loudness is ITU-R BS.1770-4 with proper K-weighting and two-stage gating, reading −22.99 LUFS on EBU Tech 3341's −23.0 test signal; the filter coefficients match the standard's published values to eight decimal places. True-peak detection is band-limited reconstruction, so the classic fs/4-at-45° signal correctly reads 0.0 dBTP despite every sample sitting at −3.01 dBFS. The resampler measures about 60 dB of alias rejection. All of it in numpy, all of it tested.
pip install -e ".[service]"
openspeech serve --port 8000| Endpoint | Purpose |
|---|---|
POST /v1/plan |
Turn a request into a plan. Instant, no model time. |
POST /v1/jobs |
Submit audio + a request. Returns a job id immediately. |
GET /v1/jobs/{id} |
Status, quality verdicts, per-step reports. |
GET /v1/jobs/{id}/events |
Server-sent progress events. |
GET /v1/jobs/{id}/result |
The finished WAV. |
POST /v1/measure |
Objective metrics for a file. |
GET /v1/ops |
The full operation catalogue, generated from the registry. |
GET /healthz |
Health, and which backend is live. |
Full reference in docs/api.md.
Post-production for a real episode — several tracks, each needing its own cleanup, assembled in order, delivered at a specified loudness.
{
"name": "episode-12",
"target_lufs": -16.0,
"gap_seconds": 0.4,
"tracks": [
{"path": "host.wav", "label": "host", "request": "remove the noise and reverb, drop the breaths"},
{"path": "guest.wav", "label": "guest", "request": "clean it up and take the phone-line sound off"}
],
"output": "episode-12.wav"
}openspeech studio episode.json -o build/episode-12: 2 tracks -> build/episode-12.wav
host 42.0s -> 42.0s -28.4 -> -16.0 LUFS [pass]
guest 36.0s -> 36.0s -31.2 -> -16.0 LUFS [pass]
master -16.00 LUFS, true peak -1.00 dBTP, 78.4s
Details in docs/studio.md.
| Architecture | How the layers fit, and why the core has one dependency |
| Planning | Clause splitting, matchers, ordering, the LLM planner |
| Long-form | Budget maths, cut placement, stitching, level drift |
| Quality control | Metrics, gates, retry policy, exact repair |
| Operations | All 20 operations, parameters, instruction templates |
| HTTP API | Endpoints, payloads, job lifecycle |
| Studio | Recipes and production reports |
| Cookbook | Task-oriented recipes |
| Contributing | Setup, tests, adding an operation |
- The mock backend is not AuK. It approximates operations with DSP so the surrounding system can be built and tested. Audio from it is not model output, and every surface says so.
- The rule planner is patterns, not understanding. It covers a wide range of English and Chinese phrasings and reports what it cannot match — but it will miss things. Use
--llmfor open-ended phrasing, or pass an explicit plan. - Chunking cannot rescue every operation. Word-level edits and order-based speaker separation are limited by the context window; OpenSpeech tells you instead of guessing.
- Quality gates are objective proxies. They catch silence, clipping, wrong duration, unapplied gain, wrong pitch and gross timbre substitution. They are not a judgement of whether the result sounds good.
- Not all tasks have bilingual instructions upstream. Where AuK ships wording in one language only, OpenSpeech uses that wording rather than inventing a translation the model never saw.
Bug reports, new intent patterns, and operations are all welcome. pip install -e ".[dev]", then pytest and ruff check src tests. See docs/CONTRIBUTING.md.
OpenSpeech is MIT licensed. It is an independent project and is not affiliated with or endorsed by Tencent.
AuK is MIT licensed, by Tencent Hunyuan and the AuK authors — repository, project page, technical report. See NOTICE for exactly what OpenSpeech derives from it.
@misc{ma2026auktechnicalreportopensource,
title = {AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing},
author = {Ziyang Ma and Zhikang Niu and Wenming Tu and Tianrui Wang and others},
year = {2026},
eprint = {2609.08936},
archivePrefix = {arXiv},
primaryClass = {cs.SD},
url = {https://arxiv.org/abs/2609.08936}
}