AuK returns audio, not a confidence. If a generation collapses to silence, clips, drifts off the requested duration, or simply ignores the instruction, nothing upstream notices. On a 40-minute recording split into 170 model calls, "nothing notices" is not an acceptable position.
Every result is measured against what the operation promised to do.
openspeech.qc.measure() reports, with no reference needed:
| Metric | What it catches |
|---|---|
duration, sample_rate |
Truncation, wrong-length generation |
peak_dbfs, true_peak_dbfs |
Clipping, inter-sample overshoot |
rms_dbfs, lufs |
Level and perceived loudness |
dc_offset |
A bias the model sometimes introduces |
clipped_ratio |
How much of the file is pinned at full scale |
nonfinite_samples |
NaN/Inf from a diverged sample |
speech_ratio, silence_ratio |
A generation that collapsed to silence |
spectral_flatness |
Noise-like output where speech was expected |
f0_median_hz |
The pitch actually produced |
compare() adds the reference-relative half: duration_ratio, level_delta_db, lufs_delta, timbre_similarity (cosine of mean log-mel spectra), pitch_delta_semitones, flatness_delta.
estimate_f0 runs autocorrelation over voiced frames with a normalised peak test and parabolic interpolation. On synthetic harmonic signals it lands within 0.1 Hz of the true f0, and it measures a known two-semitone shift as 2.00 semitones.
That accuracy is the point: without it, "raise the pitch by 2 semitones" is an instruction you send and hope about.
from openspeech.qc import evaluate
report = evaluate(op, result_audio, source_audio, expected_seconds=7.0)
report.verdict # Verdict.PASS | WARN | FAILIntegrity checks run on everything: finite samples, non-empty, not silent, no clipping, no DC offset.
Intent checks depend on the operation, and this is where the real value is:
| Operation | Check | Tolerance |
|---|---|---|
volume_edit |
Measured level change matches the requested dB | ±4 dB, error |
pitch_edit |
Measured f0 change matches the requested semitones | ±1.2 st, warning |
speed_edit |
Duration ratio matches 1/multiplier | ±15%, error |
enhance_speech, improve_quality, extract_vocals |
Output is not noisier than the input | flatness ≤ +0.08, warning |
| Voice-preserving operations | Spectral similarity to the source | ≥0.55, warning |
So a volume edit that did not land is caught, by measurement:
[error] volume_applied: level moved +4.0 dB, asked for +10.0 dB
Tolerances are deliberately loose. The model is generative, and a gate that fires on normal variation is a gate people switch off.
Intent checks are skipped when an integrity check already failed — there is no point asking whether silence has the right pitch.
What happens next depends on the kind of fault. Getting this distinction right saves real money.
A collapse to silence, a diverged sample, a wildly wrong duration: that is a bad draw. Retry with a different seed and more sampling steps.
retry_settings(report, attempt=1, policy)
# {'seed': 7919, 'nfe': 48}A gain that landed 6 dB short is not a bad draw, it is a number. Re-rolling the model costs GPU time and usually lands somewhere else wrong. OpenSpeech closes the gap exactly:
repair: applied the residual +6.0 dB the model left on the table
(asked +10.0 dB, got +4.0 dB)
retry_settings returns None when every failure is one DSP can fix, so those attempts are never spent.
Repairs currently applied:
| Fault | Repair |
|---|---|
| DC offset | Subtract the mean. Exact. |
| Clipping / true-peak overshoot | Iterative true-peak limiter to the ceiling |
| Volume shortfall | Apply the residual gain |
| Length overrun/shortfall ≤5% | Trim or pad — only for operations that promised equal length |
That last restriction matters: trimming a content edit would cut off the words it was told to add. The condition is on the operation's declared duration policy, not on how close the number looks.
A model that substituted the voice, smeared the spectrum, or ignored the instruction cannot be fixed downstream. Those are reported, the step is marked, and the run's verdict reflects it. The gate's job is to be honest, not to make every result look green.
from openspeech.qc import RepairPolicy
policy = RepairPolicy(
max_attempts=3, # total model attempts, including the first
escalate_nfe=True, # more ODE steps on each retry
vary_seed=True, # a different draw, not the same one again
dsp_corrections=True, # apply the exact fixes
accept_warnings=True, # do not burn a retry on a warning
)accept_warnings=False makes the pipeline stricter at the cost of more model calls.
result = process(audio, "boost the volume by 10 dB", engine=engine)
result.verdict # worst verdict across all steps
for step in result.steps:
print(step.op, step.verdict, step.attempts, step.engine_calls)
for repair in step.repairs:
print(" ", repair)
for entry in step.quality:
for check in entry["checks"]:
if not check["passed"]:
print(" ", check["severity"], check["name"], check["detail"])--report out.json on the CLI writes the whole structure, including per-chunk quality entries, to disk. It is JSON-serialisable by design so it can be stored, diffed across runs, or shipped to a monitoring system.