Skip to content

About

Orchestration layer for AuK speech editing: multi-intent planning, context-budget-aware chunking with seamless stitching, and BS.1770 quality gates that verify each edit. Pure-numpy core, no GPU needed to run the tests.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

OpenSpeech

A planning, long-form and verification layer for AuK, Tencent Hunyuan's open-source speech generation and editing model.

CI License: MIT Python 3.10+

AuK is excellent at what it does: one natural-language instruction, one clip, one edit. OpenSpeech is everything you need around it to ship real audio.

openspeech run -i interview.wav \
  -r "clean up the noise and reverb, keep only the second speaker, calm the delivery, and master to -16 LUFS" \
  -o episode.wav

That single command is four operations, ordered into a proper signal chain, split into model-sized pieces across a 40-minute recording, verified after every step, and mastered to a broadcast loudness target. AuK on its own does none of those four things.


Why this exists

These are not hypotheticals. Each one is a specific, cited limitation of AuK 0.1.0.

What you hit Where it comes from What OpenSpeech does
One task per request. Ask for two things and one is silently discarded. AuK's Prompt Enhancer instructs its classifier: "when a request contains several unrelated intents, take only the main one" (pe.config.yaml). Decomposes a request into an ordered plan of operations, and reports anything it could not match instead of dropping it quietly.
A ~30 second ceiling. And it is not 30 seconds of audio — the window holds the source and the generated output. "source/reference + generated target must fit within 30 [seconds]" (docs/COMFYUI.md). Computes the real per-operation budget, cuts at pauses, and stitches back with level-matched crossfades. An equal-length edit gets 13.95s per call; slowing speech to 0.5x gets 9.30s.
Parameters snapped by an LLM prompt. 1.8x is asked to become a legal 2.0x. pe.config.yaml normalisation hints. Snaps in code, arithmetically, and tells you: speed_multiplier: 1.8 -> 2.0.
No verification. If a generation collapses to silence or ignores the instruction, nothing notices. No quality signal in the API. Measures every result against what the operation promised. A +10 dB edit is checked for +10 dB. A pitch shift is checked against the output's actual f0.
CLI and Gradio only. Nothing another service can call. — A JSON HTTP API with a job queue, progress streaming, and the model loaded once.

OpenSpeech does not vendor, fork, or modify AuK. It calls AuK's documented Python API and quotes AuK's own instruction wording verbatim, because those strings are what the model was trained on.


Install

git clone https://github.com/zwayth/openspeech
cd openspeech
pip install -e .

The core — planning, DSP, chunking, quality control, CLI — depends on numpy alone. It runs and is fully tested with no GPU, no weights, and no torch.

pip install -e ".[service]"   # HTTP API
pip install -e ".[llm]"       # LLM-backed planner
pip install -e ".[accel]"     # faster metering, more audio formats
pip install -e ".[dev]"       # tests and linting

Connecting the real model

OpenSpeech ships a mock backend that approximates AuK with real DSP. It is how the test suite runs in CI, and it lets you build and debug an entire pipeline before downloading 16 GB. It is never confused for the real thing — openspeech run prints a notice, and every report carries engine: mock.

For real output, install AuK and its weights so that ckpts/AuK/auk_base.safetensors and its config.yaml exist, then:

pip install -e ".[auk]"
openspeech doctor          # confirms which backend will be used

--engine auto (the default) uses AuK when a checkpoint is present and the mock backend otherwise.


Try it in sixty seconds

No weights needed — the sample generator makes its own audio.

python examples/make_samples.py
openspeech plan "clean up the noise and reverb, calm the delivery, master to -16 LUFS"
openspeech run -i examples/assets/host.wav -r "remove the noise and normalise to -16 LUFS" -o out.wav
openspeech measure out.wav

openspeech plan costs nothing and runs instantly. It is the fastest way to see how a phrasing is understood before spending model time on it:

$ openspeech plan "clean up the noise and reverb, make the host sound calmer, speed it up a bit, and normalise to -16 LUFS"

request : clean up the noise and reverb, make the host sound calmer, speed it up a bit, and normalise to -16 LUFS
language: en   planner: rules   steps: 4

  1. [model] enhance_speech(cleanup_mode='denoise_dereverb')
  2. [model] emotion_edit(emotion='calm')
  3. [model] speed_edit(speed_multiplier=1.25)
  4. [dsp  ] loudness_normalize(target_lufs=-16.0, true_peak_db=-1.0)

   enhance_speech instruction: "Preserve all speakers, remove noise and reverberation, and output clean speech of the same length."
   emotion_edit instruction: "Change the emotion to calm."
   speed_edit instruction: "Adjust the speech speed to 1.25x."

Note step 4: loudness normalisation runs locally in DSP, not through the model. A generative model is the wrong tool for arithmetic on gain — local is exact, instant, free of GPU time, and unbound by the context window.


Python

from openspeech import process, read_audio, write_audio

result = process(read_audio("interview.wav"),
                 "remove the noise, calm the delivery, master to -16 LUFS")

write_audio("episode.wav", result.audio)

print(result.verdict)                     # pass | warn | fail
for step in result.steps:
    print(step.op, step.chunks, step.verdict, step.repairs)

Lower level, when you want to hold the parts yourself:

from openspeech import Pipeline, build_op, plan_request, create_engine
from openspeech.plan import Plan

plan = Plan(request="manual", ops=(
    build_op("enhance_speech", cleanup_mode="denoise_dereverb"),
    build_op("volume_edit", "increase", gain_db=10),
    build_op("loudness_normalize", target_lufs=-16.0),
))

result = Pipeline(create_engine("auto")).run(plan, read_audio("in.wav"))

What it actually does

Plans, not classifications

A request becomes an ordered pipeline. Ordering is a quality decision, not bookkeeping: restoration runs first because every later model pass does better on clean input, and mastering runs last because any edit before it invalidates a loudness measurement. State the steps backwards and you still get the right order.

Two intents that each ask for half a job are merged into one model pass — "remove the noise, and get rid of the reverb" becomes a single denoise_dereverb call, not two.

The rule-based planner is offline, deterministic, and needs no API key. An LLM planner handles phrasings the rules miss — and every operation it proposes is rebuilt through the same registry, so a hallucinated task name or an illegal parameter is caught locally instead of becoming a bad instruction. See docs/planning.md.

Long recordings

The chunk budget is derived per operation from how much that operation expands the audio, because AuK's window has to hold the source and the target together:

Operation Output length Source budget
enhance_speech same 13.95 s
emotion_edit (sad) ×1.22 12.57 s
speed_edit 0.5× ×2.0 9.30 s
speed_edit 2.0× ×0.5 18.60 s

Cuts land in pauses, preferring the latest usable pause so the budget is actually spent — fewer calls, fewer seams. Where there is no pause, chunks overlap and are crossfaded. Chunks are level-matched against their own sources, which removes the drift a generative model introduces between independent calls while preserving the recording's real dynamics. Round-tripping unmodified audio through chunking and stitching is sample-exact. See docs/longform.md.

Some operations cannot be chunked, and OpenSpeech says so rather than producing quiet nonsense. A word-level edit can't be split across chunks (you'd edit the wrong one), so it refuses with an actionable message. Speaker separation by talking order is a property of the whole recording, so it warns and suggests identifying the speaker by content instead.

Verification, then exact repair

Every result is measured against what the operation promised:

[error]   volume_applied: level moved +4.0 dB, asked for +10.0 dB

What happens next depends on the fault. A collapse to silence is a bad draw, so it is retried with a new seed and more sampling steps. A gain that landed 6 dB short is arithmetic — re-rolling the model wastes GPU time and lands somewhere else wrong, so OpenSpeech closes the gap exactly in DSP and says so:

repair: applied the residual +6.0 dB the model left on the table (asked +10.0 dB, got +4.0 dB)

DC offsets, inter-sample peaks and small length overruns are handled the same way. What cannot be repaired is reported honestly rather than papered over. See docs/quality.md.

Real measurement underneath

The metering is not approximate. Loudness is ITU-R BS.1770-4 with proper K-weighting and two-stage gating, reading −22.99 LUFS on EBU Tech 3341's −23.0 test signal; the filter coefficients match the standard's published values to eight decimal places. True-peak detection is band-limited reconstruction, so the classic fs/4-at-45° signal correctly reads 0.0 dBTP despite every sample sitting at −3.01 dBFS. The resampler measures about 60 dB of alias rejection. All of it in numpy, all of it tested.


The HTTP service

pip install -e ".[service]"
openspeech serve --port 8000
Endpoint Purpose
POST /v1/plan Turn a request into a plan. Instant, no model time.
POST /v1/jobs Submit audio + a request. Returns a job id immediately.
GET /v1/jobs/{id} Status, quality verdicts, per-step reports.
GET /v1/jobs/{id}/events Server-sent progress events.
GET /v1/jobs/{id}/result The finished WAV.
POST /v1/measure Objective metrics for a file.
GET /v1/ops The full operation catalogue, generated from the registry.
GET /healthz Health, and which backend is live.

Full reference in docs/api.md.


Studio: the application

Post-production for a real episode — several tracks, each needing its own cleanup, assembled in order, delivered at a specified loudness.

{
  "name": "episode-12",
  "target_lufs": -16.0,
  "gap_seconds": 0.4,
  "tracks": [
    {"path": "host.wav",  "label": "host",  "request": "remove the noise and reverb, drop the breaths"},
    {"path": "guest.wav", "label": "guest", "request": "clean it up and take the phone-line sound off"}
  ],
  "output": "episode-12.wav"
}
openspeech studio episode.json -o build/
episode-12: 2 tracks -> build/episode-12.wav
  host                 42.0s ->   42.0s   -28.4 ->  -16.0 LUFS  [pass]
  guest                36.0s ->   36.0s   -31.2 ->  -16.0 LUFS  [pass]
  master             -16.00 LUFS, true peak -1.00 dBTP, 78.4s

Details in docs/studio.md.


Documentation

Architecture How the layers fit, and why the core has one dependency
Planning Clause splitting, matchers, ordering, the LLM planner
Long-form Budget maths, cut placement, stitching, level drift
Quality control Metrics, gates, retry policy, exact repair
Operations All 20 operations, parameters, instruction templates
HTTP API Endpoints, payloads, job lifecycle
Studio Recipes and production reports
Cookbook Task-oriented recipes
Contributing Setup, tests, adding an operation

Honest limits

  • The mock backend is not AuK. It approximates operations with DSP so the surrounding system can be built and tested. Audio from it is not model output, and every surface says so.
  • The rule planner is patterns, not understanding. It covers a wide range of English and Chinese phrasings and reports what it cannot match — but it will miss things. Use --llm for open-ended phrasing, or pass an explicit plan.
  • Chunking cannot rescue every operation. Word-level edits and order-based speaker separation are limited by the context window; OpenSpeech tells you instead of guessing.
  • Quality gates are objective proxies. They catch silence, clipping, wrong duration, unapplied gain, wrong pitch and gross timbre substitution. They are not a judgement of whether the result sounds good.
  • Not all tasks have bilingual instructions upstream. Where AuK ships wording in one language only, OpenSpeech uses that wording rather than inventing a translation the model never saw.

Contributing

Bug reports, new intent patterns, and operations are all welcome. pip install -e ".[dev]", then pytest and ruff check src tests. See docs/CONTRIBUTING.md.

License and credit

OpenSpeech is MIT licensed. It is an independent project and is not affiliated with or endorsed by Tencent.

AuK is MIT licensed, by Tencent Hunyuan and the AuK authors — repository, project page, technical report. See NOTICE for exactly what OpenSpeech derives from it.

@misc{ma2026auktechnicalreportopensource,
  title         = {AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing},
  author        = {Ziyang Ma and Zhikang Niu and Wenming Tu and Tianrui Wang and others},
  year          = {2026},
  eprint        = {2609.08936},
  archivePrefix = {arXiv},
  primaryClass  = {cs.SD},
  url           = {https://arxiv.org/abs/2609.08936}
}

About

Orchestration layer for AuK speech editing: multi-intent planning, context-budget-aware chunking with seamless stitching, and BS.1770 quality gates that verify each edit. Pure-numpy core, no GPU needed to run the tests.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages